User Simulation Leaderboard
Published results for systems that model a person, one table per benchmark, with the source of every number. Within a table, systems are comparable. Across tables they are not, and nothing is averaged.
Updated 2026-09-13144 results113 systems11 benchmarksMethodologyDataSubmit
tau-USI (User-Sim Index)
Simulation30 systems · individual · mixed: real human transcripts and human outcome judgmentsSource ↗31 simulators against 451 humans on 165 tau-bench tasks. Four behavioral components are Sorensen-Dice overlap with human-annotated behaviors, Eval is (1 - MAE) x 100 on outcome judgments, USI is the composite on 0 to 100. Numbers are mean over three annotation batches.
| # | System | Org | Kind | USI | D1 Comm. | D2 Info. | D3 Clarif. | D4 React. | Eval | ECE ↓ |
|---|---|---|---|---|---|---|---|---|---|---|
| Human (inter-annotator) | reference | human | 92.7 | 87.4 | 97.9 | 88.0 | 93.5 | 97.4 | 0.0810 | |
| 1 | DeepSeek-V3.1 | DeepSeek | prompted | 76.0 | 45.1 | 86.6 | 74.5 | 87.6 | 74.3 | 0.1190 |
| 2 | Kimi K2.5 | Moonshot AI | prompted | 76.0 | 51.4 | 80.5 | 71.1 | 79.3 | 74.7 | 0.1130 |
| 3 | Gemini 2.0 Flash | prompted | 74.7 | 51.6 | 88.9 | 68.2 | 76.9 | 73.7 | 0.1110 | |
| 4 | Llama 4 Maverick | Meta | prompted | 73.7 | 48.8 | 82.6 | 78.3 | 66.6 | 76.7 | 0.1070 |
| 5 | GPT-5.1 | OpenAI | prompted | 73.5 | 47.3 | 77.4 | 73.3 | 88.1 | 72.1 | 0.1720 |
| 6 | GPT-5 | OpenAI | prompted | 72.4 | 49.7 | 73.7 | 73.2 | 73.4 | 74.5 | 0.1020 |
| 7 | GPT-4o-mini | OpenAI | prompted | 72.1 | 40.6 | 84.7 | 70.2 | 73.7 | 75.7 | 0.1230 |
| 8 | MiniMax-M2.5 | MiniMax | prompted | 71.7 | 48.1 | 72.0 | 73.9 | 61.8 | 74.5 | 0.1400 |
| 9 | GPT-5-mini | OpenAI | prompted | 71.5 | 39.4 | 74.4 | 83.1 | 68.7 | 73.5 | 0.1020 |
| 10 | GPT-4o | OpenAI | prompted | 71.2 | 31.5 | 84.4 | 74.6 | 72.4 | 73.7 | 0.0960 |
| 11 | Qwen3-235B | Alibaba | prompted | 71.1 | 60.8 | 75.3 | 71.5 | 56.3 | 74.6 | 0.1170 |
| 12 | Gemini 2.5 Flash-Lite | prompted | 69.6 | 47.9 | 86.0 | 73.2 | 59.5 | 69.6 | 0.1840 | |
| 13 | Qwen2.5-7B | Alibaba | prompted | 69.5 | 35.2 | 70.9 | 75.4 | 74.8 | 73.3 | 0.1250 |
| 14 | Qwen3-Next-80B | Alibaba | prompted | 68.4 | 38.5 | 73.2 | 68.4 | 67.9 | 71.0 | 0.0870 |
| 15 | GPT-oss-120B | OpenAI | prompted | 67.8 | 41.8 | 63.9 | 65.1 | 77.0 | 74.4 | 0.1550 |
| 16 | Gemini 3 Pro | prompted | 67.6 | 40.1 | 81.4 | 65.5 | 57.6 | 73.8 | 0.1250 | |
| 17 | CoSER-8B | CoSER | trained | 67.2 | 37.8 | 71.5 | 71.6 | 69.9 | 63.3 | 0.1090 |
| 18 | Claude 3.5 Sonnet | Anthropic | prompted | 66.0 | 39.2 | 76.9 | 59.6 | 59.3 | 74.1 | 0.1290 |
| 19 | Claude Haiku 4.5 | Anthropic | prompted | 63.2 | 25.9 | 73.4 | 55.6 | 59.0 | 75.4 | 0.1040 |
| 20 | Gemini 3 Flash | prompted | 62.4 | 37.7 | 77.4 | 56.5 | 43.9 | 71.7 | 0.1290 | |
| 21 | GPT-3.5-turbo | OpenAI | prompted | 62.1 | 39.9 | 74.1 | 58.5 | 59.9 | 73.9 | 0.3390 |
| 22 | UserLM-8B | Microsoft Research | trained | 62.0 | 30.8 | 50.8 | 56.8 | 80.0 | 67.4 | 0.1400 |
| 23 | Claude Sonnet 4 | Anthropic | prompted | 62.0 | 48.9 | 68.0 | 47.0 | 43.3 | 76.1 | 0.1140 |
| 24 | Gemini 2.5 Flash | prompted | 61.9 | 38.4 | 73.5 | 56.0 | 43.6 | 68.8 | 0.0860 | |
| 25 | Claude 3 Haiku | Anthropic | prompted | 61.8 | 22.1 | 55.7 | 72.1 | 56.9 | 78.3 | 0.1430 |
| 26 | Gemini 3.1 Pro | prompted | 61.7 | 44.1 | 67.1 | 48.9 | 45.3 | 75.1 | 0.1010 | |
| 27 | Claude 3.7 Sonnet | Anthropic | prompted | 60.3 | 26.4 | 71.2 | 50.8 | 48.9 | 72.7 | 0.0810 |
| 28 | HumanLike-7B | HumanLike | trained | 59.8 | 35.7 | 55.0 | 51.6 | 65.9 | 72.8 | 0.2200 |
| 29 | Claude Opus 4 | Anthropic | prompted | 59.2 | 32.6 | 71.9 | 46.6 | 44.9 | 73.4 | 0.1390 |
| 30 | HumanLM-opinion | Stanford (Zou group) | trained | 46.9 | 30.1 | 19.5 | 38.5 | 50.7 | 61.6 | 0.1920 |
Mind the Sim2Real Gap, arXiv 2603.11245v2, Table 1 (mean of three annotation batches; 17 of 31 rows are in Table 1, the other 14 in Table 4, Appendix A.10). Reported by benchmark authors (third party to every system), 2026-03. Tool-transcribed, checked against the source table.
Persimmon launch: USI dimensions, provider-run
Simulation6 systems · individual · mixed: as tau-USI, on the provider's harnessSource ↗The same four USI behavioral dimensions, run by the provider on its own harness (two runs under manya-serve against three under verbatim for the comparison models, different episode coverage, run-level sample SD as error bars). Not the paper's harness, so it sits beside tau-USI and not inside it; the provider itself rules out an overall ranking from these numbers.
| # | System | Org | Kind | Communication style | Information patterns | Clarification behavior | Error reaction |
|---|---|---|---|---|---|---|---|
| 1 | Persimmon | humans& | trained | 66.0 | 91.8 | 74.2 | 79.3 |
| 2 | Grok 4.6 | xAI | prompted | 58.5 | 88.0 | 82.9 | 52.1 |
| 3 | GPT-5.6 | OpenAI | prompted | 57.4 | 82.7 | 65.7 | 71.2 |
| 4 | GPT-6 (Astra) | OpenAI | prompted | 46.4 | 73.5 | 64.8 | 48.0 |
| 5 | Claude Fable 5.1 | Anthropic | prompted | 44.7 | 78.6 | 57.8 | 45.0 |
| 6 | Claude Opus 5 | Anthropic | prompted | 38.2 | 90.3 | 77.7 | 54.5 |
Persimmon launch post, 2026-09-10 (provider-run; means with run-level sample SD). Reported by the provider (Persimmon's makers), for every row, 2026-09-10. Tool-transcribed, checked against the source table.
Persimmon launch: multi-user Turing test
Simulation5 systems · individual · human rating (a judge's guess)Source ↗Mean rate at which a judge is fooled into taking the simulator for the human; 50 percent is human parity. Provider-run, provider-defined.
| # | System | Org | Kind | Judge fooled, % |
|---|---|---|---|---|
| 1 | Persimmon | humans& | trained | 19.8 |
| 2 | Nemotron Ultra Base | NVIDIA | prompted | 1.2 |
| 3 | GPT-6 (Astra) | OpenAI | prompted | 0.2 |
| 4 | Claude Fable 5 | Anthropic | prompted | 0.2 |
| 5 | OSim-8B | CMU LTI | trained | 0.1 |
Persimmon launch post, 2026-09-10 (provider-run; means with run-level sample SD). Reported by the provider, for every row, 2026-09-10. Tool-transcribed, checked against the source table.
Persimmon launch: trickle test
Simulation43 systems · individual · revealed behaviour (real human disclosure timing)Source ↗Whether the simulator reveals a fact at the turn where the real human revealed it. Precision is the share of revealed facts that the human also revealed at that turn; recall the share of the human's reveals the simulator matched. Provider-run.
| # | System | Org | Kind | Recall, % | Precision, % |
|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | prompted | 93.2 | 79.2 | |
| 2 | Claude Fable 5 | Anthropic | prompted | 92.5 | 86.1 |
| 3 | Gemini 3.8 Flash | prompted | 92.2 | 82.2 | |
| 4 | Claude Opus 4.7 | Anthropic | prompted | 90.1 | 86.2 |
| 5 | Claude Sonnet 4.5 | Anthropic | prompted | 90.0 | 83.0 |
| 6 | GLM 5.2 | Zhipu | prompted | 89.6 | 81.2 |
| 7 | GLM 5.3 | Zhipu | prompted | 89.4 | 81.1 |
| 8 | GPT-6 (Astra) | OpenAI | prompted | 89.0 | 84.2 |
| 9 | Kimi K3 | Moonshot AI | prompted | 88.8 | 84.4 |
| 10 | Qwen 3.6 Plus | Alibaba | prompted | 88.8 | 72.9 |
| 11 | Claude Haiku 4.5 | Anthropic | prompted | 88.7 | 77.0 |
| 12 | Llama 4 Maverick | Meta | prompted | 88.5 | 81.1 |
| 13 | Llama 3.3 70B | Meta | prompted | 88.4 | 78.4 |
| 14 | Qwen 3.5 Plus | Alibaba | prompted | 88.3 | 72.0 |
| 15 | GLM 5.1 | Zhipu | prompted | 88.3 | 81.4 |
| 16 | Claude Sonnet 4.6 | Anthropic | prompted | 88.3 | 84.0 |
| 17 | GPT-5.5 | OpenAI | prompted | 88.3 | 83.8 |
| 18 | Claude Sonnet 5 | Anthropic | prompted | 88.2 | 85.9 |
| 19 | GLM 5 | Zhipu | prompted | 88.0 | 82.8 |
| 20 | Claude Opus 4.8 | Anthropic | prompted | 87.7 | 84.7 |
| 21 | GPT-5.6 (Terra) | OpenAI | prompted | 87.1 | 84.6 |
| 22 | Nemotron 3 Ultra 550B | NVIDIA | prompted | 86.9 | 79.3 |
| 23 | Kimi K2.5 | Moonshot AI | prompted | 86.8 | 78.8 |
| 24 | GPT-5.6 (Sol) | OpenAI | prompted | 86.6 | 85.8 |
| 25 | Grok 4.6 | xAI | prompted | 86.6 | 80.5 |
| 26 | GPT-5.6 (Luna) | OpenAI | prompted | 86.4 | 81.0 |
| 27 | Qwen 3 Max | Alibaba | prompted | 86.2 | 77.1 |
| 28 | DeepSeek V4 Pro | DeepSeek | prompted | 86.1 | 82.1 |
| 29 | Qwen 3.8 Max | Alibaba | prompted | 86.0 | 78.1 |
| 30 | DeepSeek V3.2 | DeepSeek | prompted | 86.0 | 84.6 |
| 31 | Kimi K2.6 | Moonshot AI | prompted | 85.8 | 78.1 |
| 32 | GPT-4.1 | OpenAI | prompted | 85.8 | 81.0 |
| 33 | MiniMax M3 | MiniMax | prompted | 85.2 | 82.0 |
| 34 | Gemma 4 31B | prompted | 84.7 | 82.6 | |
| 35 | Qwen 3.8 27B | Alibaba | prompted | 83.7 | 79.0 |
| 36 | GLM 4.7 | Zhipu | prompted | 83.4 | 82.4 |
| 37 | GPT-4o | OpenAI | prompted | 83.0 | 82.9 |
| 38 | GPT-5.4 Mini | OpenAI | prompted | 82.5 | 81.8 |
| 39 | Persimmon | humans& | trained | 77.0 | 88.5 |
| 40 | GPT-oss-120B | OpenAI | prompted | 74.7 | 62.6 |
| 41 | Nemotron 3.5 Lightning | NVIDIA | prompted | 74.2 | 74.2 |
| 42 | OSim-8B | CMU LTI | trained | 68.4 | 79.9 |
| 43 | Nemotron 3 Nano 30B | NVIDIA | prompted | 64.5 | 73.2 |
Persimmon launch post, 2026-09-10 (provider-run; means with run-level sample SD). Reported by the provider, for every row, 2026-09-10. Tool-transcribed, checked against the source table.
SOUL-Index (OdysSim)
Simulation8 systems · individual · mixed, task by task; mostly LLM judgeSource ↗23 human-simulation tasks across five axes (CONV discourse, SS social skills, COG theory of mind, ROLE persona and role-play, EVAL judgment), each normalized to 0 to 100 and averaged unweighted. 14 of the 23 are judge-scored; a handful score real replies, survey answers or preference judgments. No human row. From Table 2 of the paper; the 'best specialized baseline' column is not a single system and is omitted.
| # | System | Org | Kind | Average (23 tasks) | # best of 23 | UserLLM | MirrorBench | Humanual-Chat | SimArena-Doc | Sotopia-Hard | Fantom | Hitom | Paratomi | Social-R1 | Coser | Lifechoices | Twinvoice | BehaviorChain | SimArena-Math | Mistakes | Humanual-Email | Humanual-News | Humanual-Politics | AlignX | Humanllm | Socsci210 | Humanual-Book | Humanual-Opinion |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.7 | Anthropic | prompted | 65.5 | 7.00 | 57.6 | 63.7 | 22.6 | 83.5 | 32.4 | 80.0 | 93.0 | 90.0 | 67.0 | 66.5 | 92.0 | 83.0 | 96.0 | 68.7 | 74.0 | 50.4 | 41.3 | 43.5 | 71.6 | 44.2 | 77.2 | 61.4 | 46.2 |
| 2 | GPT-5.5 | OpenAI | prompted | 65.2 | 3.00 | 65.3 | 56.7 | 28.2 | 83.4 | 31.9 | 93.0 | 82.0 | 99.0 | 69.0 | 66.2 | 91.0 | 74.0 | 95.0 | 68.5 | 72.0 | 50.1 | 40.2 | 42.0 | 71.2 | 45.7 | 77.2 | 57.6 | 39.8 |
| 3 | Gemini 3.1 Pro | prompted | 64.8 | 7.00 | 67.7 | 48.3 | 21.0 | 83.0 | 27.8 | 93.0 | 86.0 | 97.0 | 79.0 | 62.1 | 84.0 | 86.0 | 92.0 | 71.5 | 73.0 | 46.9 | 42.3 | 32.5 | 73.4 | 46.9 | 78.0 | 62.4 | 36.0 | |
| 4 | OSim-8B | CMU LTI | trained | 64.6 | 8.00 | 90.1 | 68.3 | 28.2 | 84.1 | 49.2 | 80.0 | 79.0 | 83.0 | 60.0 | 62.6 | 82.0 | 68.0 | 94.0 | 70.7 | 59.0 | 51.4 | 42.7 | 41.9 | 72.6 | 39.1 | 75.1 | 63.2 | 42.0 |
| 5 | Qwen 3.6 Plus | Alibaba | prompted | 61.1 | · | 72.1 | 48.0 | 22.2 | 82.4 | 28.3 | 89.0 | 73.0 | 94.0 | 67.0 | 55.9 | 79.0 | 71.0 | 85.0 | 70.9 | 67.0 | 47.9 | 41.8 | 31.6 | 69.8 | 42.7 | 74.5 | 58.4 | 34.2 |
| 6 | Qwen3-8B | Alibaba | prompted | 48.3 | · | 46.0 | 54.0 | 24.7 | 83.6 | 27.7 | 23.0 | 62.0 | 67.0 | 54.0 | 43.5 | 70.0 | 42.0 | 41.0 | 68.9 | 27.0 | 43.7 | 32.5 | 33.2 | 68.6 | 34.1 | 73.6 | 53.6 | 37.2 |
| 7 | OSim-8B-Mid | CMU LTI | trained | 41.1 | · | 49.5 | 49.1 | 7.8 | 80.3 | 45.6 | 62.0 | 54.0 | 72.0 | 42.0 | 24.8 | 58.0 | 25.0 | 42.0 | 68.1 | 18.0 | 22.3 | 15.1 | 15.4 | 53.6 | 16.5 | 68.1 | 38.8 | 17.0 |
| 8 | Base 8B (OdysSim) | CMU LTI | baseline | 26.9 | · | 31.0 | 13.9 | 12.0 | 79.6 | 21.4 | 23.0 | 12.0 | 19.0 | 37.0 | 6.1 | 32.0 | 19.0 | 18.0 | 66.2 | 24.0 | 26.4 | 12.7 | 17.8 | 49.0 | 12.1 | 46.6 | 21.5 | 18.2 |
OdysSim, arXiv 2606.14199v2, Table 2 (0 to 100, higher is better). Reported by benchmark authors, who are also OSim-8B's makers, 2026-06. Transcribed by hand.
SimBench
Simulation16 systems · population · stated answerSource ↗20 human-behavior datasets unified; the SimBench score is the improvement of a model's predicted answer distribution over a uniform baseline, relative to the human distribution, by total variation distance: 100 is perfect alignment, 0 is random. Population level, stated answers. Table 1 of the paper.
| # | System | Org | Kind | SimBench score |
|---|---|---|---|---|
| 1 | Claude 3.7 Sonnet | Anthropic | prompted | 40.8 |
| 2 | Claude 3.7 Sonnet (4000) | Anthropic | prompted | 39.5 |
| 3 | GPT-4.1 | OpenAI | prompted | 34.5 |
| 4 | DeepSeek-R1 | DeepSeek | prompted | 34.5 |
| 5 | o4-mini-high | OpenAI | prompted | 29.0 |
| 6 | Llama 3.1 405B Instruct | Meta | prompted | 28.4 |
| 7 | Qwen2.5-72B-Instruct | Alibaba | prompted | 27.6 |
| 8 | Qwen2.5-32B-Instruct | Alibaba | prompted | 23.8 |
| 9 | OLMo-2-32B-DPO | Ai2 | prompted | 19.8 |
| 10 | OLMo-2-32B | Ai2 | prompted | 15.9 |
| 11 | OLMo-2-13B | Ai2 | prompted | 13.8 |
| 12 | Qwen2.5-72B | Alibaba | prompted | 13.3 |
| 13 | Qwen2.5-32B | Alibaba | prompted | 12.3 |
| 14 | Gemma 3 4B PT | prompted | -0.7 | |
| 15 | Qwen2.5-3B-Instruct | Alibaba | prompted | -12.0 |
| 16 | OLMo-2-7B-Instruct | Ai2 | prompted | -21.4 |
SimBench, arXiv 2510.17516v4, Table 1. Reported by benchmark authors (third party to every system), 2025-10. Tool-transcribed, checked against the source table.
HUMANUAL (HumanLM)
Simulation7 systems · individual · revealed behaviour (real recorded text)Source ↗Six datasets of real recorded responses (news comments, book reviews, opinions, political text, chat, email) from about 23,000 users; response alignment score, higher is better. Table 1 of the HumanLM paper; the paper's own system is among the rows.
| # | System | Org | Kind | Average | News | Book | Opinion | Politics | Chat | |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | HumanLM | Stanford (Zou group) | trained | 13.2 | 9.55 | 18.5 | 25.6 | 12.6 | 6.08 | 6.71 |
| 2 | GRPO-think | Stanford (Zou group) | trained | 10.4 | 7.04 | 12.8 | 23.8 | 10.6 | 3.16 | 4.78 |
| 3 | GRPO | Stanford (Zou group) | trained | 10.3 | 7.92 | 13.3 | 18.2 | 10.9 | 5.83 | 5.90 |
| 4 | Qwen3-8B | Alibaba | prompted | 9.5 | 5.68 | 13.6 | 18.7 | 10.1 | 3.90 | 4.76 |
| 5 | SFT-think | Stanford (Zou group) | trained | 8.6 | 6.00 | 13.4 | 16.7 | 9.2 | 2.50 | 3.94 |
| 6 | Qwen3-8B-think | Alibaba | prompted | 8.4 | 4.83 | 12.8 | 20.4 | 7.0 | 2.16 | 3.22 |
| 7 | SFT | Stanford (Zou group) | trained | 6.5 | 3.10 | 9.3 | 11.3 | 6.3 | 4.57 | 4.30 |
HumanLM, arXiv 2603.03303, Table 1 (response alignment, higher is better). Reported by benchmark authors, who are also HumanLM's makers, 2026-03. Tool-transcribed, checked against the source table.
OmniBehavior
Simulation11 systems · individual · mixed: revealed behaviour for actions, LLM judge for textSource ↗200 users from a large short-video platform, three months of real logs, about 8,000 actions each across five scenarios and 22 action types. Binary behaviors by F1, continuous by normalized error, textual by a judge. Overall is the paper's composite. Table 1. Sits here for its framing as user simulation; its data is a platform behavior log, so it neighbours the RecSys track.
| # | System | Org | Kind | Overall | Video | Live | Ads | E-commerce (binary) | Shop (binary) | Textual (judge) |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.5 | Anthropic | prompted | 44.5 | 33.0 | 64.2 | 31.7 | 51.2 | 30.0 | 57.2 |
| 2 | GLM-4.7 | Zhipu | prompted | 41.5 | 26.9 | 64.4 | 29.0 | 40.3 | 32.9 | 55.3 |
| 3 | Claude Sonnet 4.5 | Anthropic | prompted | 40.5 | 18.9 | 66.0 | 25.0 | 42.8 | 36.1 | 54.3 |
| 4 | GPT-5.2 | OpenAI | prompted | 39.1 | 31.5 | 65.0 | 28.6 | 33.6 | 29.3 | 46.3 |
| 5 | Kimi K2 Instruct | Moonshot AI | prompted | 37.6 | 23.3 | 64.8 | 28.6 | 31.2 | 29.9 | 47.8 |
| 6 | DeepSeek-V3 | DeepSeek | prompted | 37.4 | 21.4 | 64.0 | 27.9 | 25.7 | 33.3 | 52.1 |
| 7 | Claude Sonnet 4 | Anthropic | prompted | 36.9 | 25.3 | 64.6 | 28.9 | 36.8 | 16.5 | 49.1 |
| 8 | Claude Haiku 4.5 | Anthropic | prompted | 36.5 | 22.8 | 63.3 | 26.1 | 30.0 | 26.4 | 50.3 |
| 9 | GPT-4o | OpenAI | prompted | 36.3 | 27.9 | 62.8 | 28.1 | 25.2 | 28.7 | 44.9 |
| 10 | Gemini 3 Flash | prompted | 32.6 | 22.1 | 53.8 | 25.6 | 24.6 | 19.6 | 49.8 | |
| 11 | Qwen3-235B | Alibaba | prompted | 32.1 | 18.3 | 62.4 | 23.8 | 23.2 | 19.2 | 45.7 |
OmniBehavior, arXiv 2604.08362v1, Table 1. Reported by benchmark authors (third party to every system), 2026-04. Tool-transcribed, checked against the source table.
LongNAP / NAPsack
HCI6 systems · individual · LLM judge against a real logged next actionSource ↗360,000 labeled actions over one month of continuous phone use by 20 users. Score is a judge model's similarity between predicted and true next action on 0 to 1; split is temporal within each user (weeks 1 and 2 train, 3 validate, 4 test). Table 3, mean over the 20 users. The paper's own system is among the rows.
| # | System | Org | Kind | Judge similarity (mean of 10 users) |
|---|---|---|---|---|
| 1 | LongNAP | Stanford (GUM group) | trained | 0.3800 |
| 2 | Gemini few-shot RAG | prompted | 0.2700 | |
| 3 | Gemini zero-shot | prompted | 0.2600 | |
| 4 | Qwen SFT | Stanford (GUM group) | trained | 0.2100 |
| 5 | Qwen few-shot RAG | Alibaba | prompted | 0.2000 |
| 6 | Qwen zero-shot | Alibaba | prompted | 0.1800 |
LongNAP, arXiv 2603.05923, Table 3, mean over the 20 users (the last row). Reported by benchmark authors, who are also LongNAP's makers, 2026-03. Tool-transcribed, checked against the source table.
OpenOneRec Amazon transfer benchmark (10 categories)
RecSys4 systems · individual · revealed behaviourSource ↗Recall@10 on ten Amazon categories, as reported in the OpenOneRec repository README (cross-domain transferability table). 'Ours' is the OneRec foundation model.
| # | System | Org | Kind | Baby | Beauty | Cell Phones | Grocery | Health | Home | Pet Supplies | Sports | Tools | Toys |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | OneRec (OpenOneRec-Foundation) | Kuaishou | trained | 0.0513 | 0.0924 | 0.1036 | 0.1029 | 0.0768 | 0.0390 | 0.0834 | 0.0547 | 0.0593 | 0.0953 |
| 2 | SASRec | baseline (reported by OpenOneRec) | baseline | 0.0381 | 0.0639 | 0.0782 | 0.0789 | 0.0506 | 0.0212 | 0.0607 | 0.0389 | 0.0437 | 0.0658 |
| 3 | LC-Rec | baseline (reported by OpenOneRec) | trained | 0.0344 | 0.0764 | 0.0883 | 0.0790 | 0.0616 | 0.0293 | 0.0612 | 0.0418 | 0.0438 | 0.0549 |
| 4 | TIGER | baseline (reported by OpenOneRec) | trained | 0.0318 | 0.0628 | 0.0786 | 0.0691 | 0.0534 | 0.0216 | 0.0542 | 0.0331 | 0.0344 | 0.0527 |
OpenOneRec repository README, RecIF-Bench results table (cross-domain transferability table, Recall@10). Reported by benchmark authors, who are also OneRec's makers, 2026-01. Transcribed by hand.
RecIF-Bench
RecSys7 systems · individual · mixed: revealed behaviour for recommendation and label tasks, LLM judge for understanding and explanationSource ↗Eight tasks in a four-layer hierarchy over 100M interactions from 200k users across short video, ads and product. Recall@32 for recommendation tasks, AUC for label prediction, a judge score for item understanding and explanation. From the OpenOneRec README results table.
| # | System | Org | Kind | Short video rec (R@32) | Ad rec (R@32) | Product rec (R@32) | Label-cond. rec (R@32) | Label pred. (AUC) | Interactive rec (R@32) | Item understanding (judge) | Rec. explanation (judge) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | OneRec-8B-Pro | Kuaishou | trained | 0.0369 | 0.0964 | 0.0538 | 0.0235 | 0.6912 | 0.3458 | 0.3209 | 4.04 |
| 2 | OneRec-8B | Kuaishou | trained | 0.0355 | 0.0877 | 0.0470 | 0.0228 | 0.6615 | 0.3032 | 0.3202 | 3.68 |
| 3 | OneRec-1.7B-Pro | Kuaishou | trained | 0.0274 | 0.0735 | 0.0405 | 0.0182 | 0.6071 | 0.2024 | 0.3133 | 3.51 |
| 4 | OneRec-1.7B | Kuaishou | trained | 0.0272 | 0.0707 | 0.0360 | 0.0184 | 0.6184 | 0.1941 | 0.3175 | 3.35 |
| 5 | LC-Rec | baseline (reported by OpenOneRec) | trained | 0.0180 | 0.0723 | 0.0416 | 0.0170 | 0.6139 | 0.2394 | 0.2517 | 3.94 |
| 6 | TIGER | baseline (reported by OpenOneRec) | trained | 0.0132 | 0.0581 | 0.0283 | 0.0123 | 0.6675 | · | · | · |
| 7 | SASRec | baseline (reported by OpenOneRec) | baseline | 0.0119 | 0.0293 | 0.0175 | 0.0140 | 0.6244 | · | · | · |
OpenOneRec repository README, RecIF-Bench results table. Reported by benchmark authors, who are also OneRec's makers, 2026-01. Transcribed by hand.
In the taxonomy, no cross-system table transcribed yet: OpinionQA, Twin-2K-500, SocSci210, GUM, LOCOMO, BARS, RelBench