Jean Technologies Icon
Jean Technologies Text

User Simulation Leaderboard

Published results for systems that model a person, one table per benchmark, with the source of every number. Within a table, systems are comparable. Across tables they are not, and nothing is averaged.

Updated 2026-09-13144 results113 systems11 benchmarksMethodologyDataSubmit

tau-USI (User-Sim Index)

Simulation30 systems · individual · mixed: real human transcripts and human outcome judgmentsSource ↗

31 simulators against 451 humans on 165 tau-bench tasks. Four behavioral components are Sorensen-Dice overlap with human-annotated behaviors, Eval is (1 - MAE) x 100 on outcome judgments, USI is the composite on 0 to 100. Numbers are mean over three annotation batches.

#SystemOrgKindUSID1 Comm.D2 Info.D3 Clarif.D4 React.EvalECE ↓
Human (inter-annotator)referencehuman92.787.497.988.093.597.40.0810
1DeepSeek-V3.1DeepSeekprompted76.045.186.674.587.674.30.1190
2Kimi K2.5Moonshot AIprompted76.051.480.571.179.374.70.1130
3Gemini 2.0 FlashGoogleprompted74.751.688.968.276.973.70.1110
4Llama 4 MaverickMetaprompted73.748.882.678.366.676.70.1070
5GPT-5.1OpenAIprompted73.547.377.473.388.172.10.1720
6GPT-5OpenAIprompted72.449.773.773.273.474.50.1020
7GPT-4o-miniOpenAIprompted72.140.684.770.273.775.70.1230
8MiniMax-M2.5MiniMaxprompted71.748.172.073.961.874.50.1400
9GPT-5-miniOpenAIprompted71.539.474.483.168.773.50.1020
10GPT-4oOpenAIprompted71.231.584.474.672.473.70.0960
11Qwen3-235BAlibabaprompted71.160.875.371.556.374.60.1170
12Gemini 2.5 Flash-LiteGoogleprompted69.647.986.073.259.569.60.1840
13Qwen2.5-7BAlibabaprompted69.535.270.975.474.873.30.1250
14Qwen3-Next-80BAlibabaprompted68.438.573.268.467.971.00.0870
15GPT-oss-120BOpenAIprompted67.841.863.965.177.074.40.1550
16Gemini 3 ProGoogleprompted67.640.181.465.557.673.80.1250
17CoSER-8BCoSERtrained67.237.871.571.669.963.30.1090
18Claude 3.5 SonnetAnthropicprompted66.039.276.959.659.374.10.1290
19Claude Haiku 4.5Anthropicprompted63.225.973.455.659.075.40.1040
20Gemini 3 FlashGoogleprompted62.437.777.456.543.971.70.1290
21GPT-3.5-turboOpenAIprompted62.139.974.158.559.973.90.3390
22UserLM-8BMicrosoft Researchtrained62.030.850.856.880.067.40.1400
23Claude Sonnet 4Anthropicprompted62.048.968.047.043.376.10.1140
24Gemini 2.5 FlashGoogleprompted61.938.473.556.043.668.80.0860
25Claude 3 HaikuAnthropicprompted61.822.155.772.156.978.30.1430
26Gemini 3.1 ProGoogleprompted61.744.167.148.945.375.10.1010
27Claude 3.7 SonnetAnthropicprompted60.326.471.250.848.972.70.0810
28HumanLike-7BHumanLiketrained59.835.755.051.665.972.80.2200
29Claude Opus 4Anthropicprompted59.232.671.946.644.973.40.1390
30HumanLM-opinionStanford (Zou group)trained46.930.119.538.550.761.60.1920

Mind the Sim2Real Gap, arXiv 2603.11245v2, Table 1 (mean of three annotation batches; 17 of 31 rows are in Table 1, the other 14 in Table 4, Appendix A.10). Reported by benchmark authors (third party to every system), 2026-03. Tool-transcribed, checked against the source table.

Persimmon launch: USI dimensions, provider-run

Simulation6 systems · individual · mixed: as tau-USI, on the provider's harnessSource ↗

The same four USI behavioral dimensions, run by the provider on its own harness (two runs under manya-serve against three under verbatim for the comparison models, different episode coverage, run-level sample SD as error bars). Not the paper's harness, so it sits beside tau-USI and not inside it; the provider itself rules out an overall ranking from these numbers.

#SystemOrgKindCommunication styleInformation patternsClarification behaviorError reaction
1Persimmonhumans&trained66.091.874.279.3
2Grok 4.6xAIprompted58.588.082.952.1
3GPT-5.6OpenAIprompted57.482.765.771.2
4GPT-6 (Astra)OpenAIprompted46.473.564.848.0
5Claude Fable 5.1Anthropicprompted44.778.657.845.0
6Claude Opus 5Anthropicprompted38.290.377.754.5

Persimmon launch post, 2026-09-10 (provider-run; means with run-level sample SD). Reported by the provider (Persimmon's makers), for every row, 2026-09-10. Tool-transcribed, checked against the source table.

Persimmon launch: multi-user Turing test

Simulation5 systems · individual · human rating (a judge's guess)Source ↗

Mean rate at which a judge is fooled into taking the simulator for the human; 50 percent is human parity. Provider-run, provider-defined.

#SystemOrgKindJudge fooled, %
1Persimmonhumans&trained19.8
2Nemotron Ultra BaseNVIDIAprompted1.2
3GPT-6 (Astra)OpenAIprompted0.2
4Claude Fable 5Anthropicprompted0.2
5OSim-8BCMU LTItrained0.1

Persimmon launch post, 2026-09-10 (provider-run; means with run-level sample SD). Reported by the provider, for every row, 2026-09-10. Tool-transcribed, checked against the source table.

Persimmon launch: trickle test

Simulation43 systems · individual · revealed behaviour (real human disclosure timing)Source ↗

Whether the simulator reveals a fact at the turn where the real human revealed it. Precision is the share of revealed facts that the human also revealed at that turn; recall the share of the human's reveals the simulator matched. Provider-run.

#SystemOrgKindRecall, %Precision, %
1Gemini 3.1 ProGoogleprompted93.279.2
2Claude Fable 5Anthropicprompted92.586.1
3Gemini 3.8 FlashGoogleprompted92.282.2
4Claude Opus 4.7Anthropicprompted90.186.2
5Claude Sonnet 4.5Anthropicprompted90.083.0
6GLM 5.2Zhipuprompted89.681.2
7GLM 5.3Zhipuprompted89.481.1
8GPT-6 (Astra)OpenAIprompted89.084.2
9Kimi K3Moonshot AIprompted88.884.4
10Qwen 3.6 PlusAlibabaprompted88.872.9
11Claude Haiku 4.5Anthropicprompted88.777.0
12Llama 4 MaverickMetaprompted88.581.1
13Llama 3.3 70BMetaprompted88.478.4
14Qwen 3.5 PlusAlibabaprompted88.372.0
15GLM 5.1Zhipuprompted88.381.4
16Claude Sonnet 4.6Anthropicprompted88.384.0
17GPT-5.5OpenAIprompted88.383.8
18Claude Sonnet 5Anthropicprompted88.285.9
19GLM 5Zhipuprompted88.082.8
20Claude Opus 4.8Anthropicprompted87.784.7
21GPT-5.6 (Terra)OpenAIprompted87.184.6
22Nemotron 3 Ultra 550BNVIDIAprompted86.979.3
23Kimi K2.5Moonshot AIprompted86.878.8
24GPT-5.6 (Sol)OpenAIprompted86.685.8
25Grok 4.6xAIprompted86.680.5
26GPT-5.6 (Luna)OpenAIprompted86.481.0
27Qwen 3 MaxAlibabaprompted86.277.1
28DeepSeek V4 ProDeepSeekprompted86.182.1
29Qwen 3.8 MaxAlibabaprompted86.078.1
30DeepSeek V3.2DeepSeekprompted86.084.6
31Kimi K2.6Moonshot AIprompted85.878.1
32GPT-4.1OpenAIprompted85.881.0
33MiniMax M3MiniMaxprompted85.282.0
34Gemma 4 31BGoogleprompted84.782.6
35Qwen 3.8 27BAlibabaprompted83.779.0
36GLM 4.7Zhipuprompted83.482.4
37GPT-4oOpenAIprompted83.082.9
38GPT-5.4 MiniOpenAIprompted82.581.8
39Persimmonhumans&trained77.088.5
40GPT-oss-120BOpenAIprompted74.762.6
41Nemotron 3.5 LightningNVIDIAprompted74.274.2
42OSim-8BCMU LTItrained68.479.9
43Nemotron 3 Nano 30BNVIDIAprompted64.573.2

Persimmon launch post, 2026-09-10 (provider-run; means with run-level sample SD). Reported by the provider, for every row, 2026-09-10. Tool-transcribed, checked against the source table.

SOUL-Index (OdysSim)

Simulation8 systems · individual · mixed, task by task; mostly LLM judgeSource ↗

23 human-simulation tasks across five axes (CONV discourse, SS social skills, COG theory of mind, ROLE persona and role-play, EVAL judgment), each normalized to 0 to 100 and averaged unweighted. 14 of the 23 are judge-scored; a handful score real replies, survey answers or preference judgments. No human row. From Table 2 of the paper; the 'best specialized baseline' column is not a single system and is omitted.

#SystemOrgKindAverage (23 tasks)# best of 23UserLLMMirrorBenchHumanual-ChatSimArena-DocSotopia-HardFantomHitomParatomiSocial-R1CoserLifechoicesTwinvoiceBehaviorChainSimArena-MathMistakesHumanual-EmailHumanual-NewsHumanual-PoliticsAlignXHumanllmSocsci210Humanual-BookHumanual-Opinion
1Claude Opus 4.7Anthropicprompted65.57.0057.663.722.683.532.480.093.090.067.066.592.083.096.068.774.050.441.343.571.644.277.261.446.2
2GPT-5.5OpenAIprompted65.23.0065.356.728.283.431.993.082.099.069.066.291.074.095.068.572.050.140.242.071.245.777.257.639.8
3Gemini 3.1 ProGoogleprompted64.87.0067.748.321.083.027.893.086.097.079.062.184.086.092.071.573.046.942.332.573.446.978.062.436.0
4OSim-8BCMU LTItrained64.68.0090.168.328.284.149.280.079.083.060.062.682.068.094.070.759.051.442.741.972.639.175.163.242.0
5Qwen 3.6 PlusAlibabaprompted61.1·72.148.022.282.428.389.073.094.067.055.979.071.085.070.967.047.941.831.669.842.774.558.434.2
6Qwen3-8BAlibabaprompted48.3·46.054.024.783.627.723.062.067.054.043.570.042.041.068.927.043.732.533.268.634.173.653.637.2
7OSim-8B-MidCMU LTItrained41.1·49.549.17.880.345.662.054.072.042.024.858.025.042.068.118.022.315.115.453.616.568.138.817.0
8Base 8B (OdysSim)CMU LTIbaseline26.9·31.013.912.079.621.423.012.019.037.06.132.019.018.066.224.026.412.717.849.012.146.621.518.2

OdysSim, arXiv 2606.14199v2, Table 2 (0 to 100, higher is better). Reported by benchmark authors, who are also OSim-8B's makers, 2026-06. Transcribed by hand.

SimBench

Simulation16 systems · population · stated answerSource ↗

20 human-behavior datasets unified; the SimBench score is the improvement of a model's predicted answer distribution over a uniform baseline, relative to the human distribution, by total variation distance: 100 is perfect alignment, 0 is random. Population level, stated answers. Table 1 of the paper.

#SystemOrgKindSimBench score
1Claude 3.7 SonnetAnthropicprompted40.8
2Claude 3.7 Sonnet (4000)Anthropicprompted39.5
3GPT-4.1OpenAIprompted34.5
4DeepSeek-R1DeepSeekprompted34.5
5o4-mini-highOpenAIprompted29.0
6Llama 3.1 405B InstructMetaprompted28.4
7Qwen2.5-72B-InstructAlibabaprompted27.6
8Qwen2.5-32B-InstructAlibabaprompted23.8
9OLMo-2-32B-DPOAi2prompted19.8
10OLMo-2-32BAi2prompted15.9
11OLMo-2-13BAi2prompted13.8
12Qwen2.5-72BAlibabaprompted13.3
13Qwen2.5-32BAlibabaprompted12.3
14Gemma 3 4B PTGoogleprompted-0.7
15Qwen2.5-3B-InstructAlibabaprompted-12.0
16OLMo-2-7B-InstructAi2prompted-21.4

SimBench, arXiv 2510.17516v4, Table 1. Reported by benchmark authors (third party to every system), 2025-10. Tool-transcribed, checked against the source table.

HUMANUAL (HumanLM)

Simulation7 systems · individual · revealed behaviour (real recorded text)Source ↗

Six datasets of real recorded responses (news comments, book reviews, opinions, political text, chat, email) from about 23,000 users; response alignment score, higher is better. Table 1 of the HumanLM paper; the paper's own system is among the rows.

#SystemOrgKindAverageNewsBookOpinionPoliticsChatEmail
1HumanLMStanford (Zou group)trained13.29.5518.525.612.66.086.71
2GRPO-thinkStanford (Zou group)trained10.47.0412.823.810.63.164.78
3GRPOStanford (Zou group)trained10.37.9213.318.210.95.835.90
4Qwen3-8BAlibabaprompted9.55.6813.618.710.13.904.76
5SFT-thinkStanford (Zou group)trained8.66.0013.416.79.22.503.94
6Qwen3-8B-thinkAlibabaprompted8.44.8312.820.47.02.163.22
7SFTStanford (Zou group)trained6.53.109.311.36.34.574.30

HumanLM, arXiv 2603.03303, Table 1 (response alignment, higher is better). Reported by benchmark authors, who are also HumanLM's makers, 2026-03. Tool-transcribed, checked against the source table.

OmniBehavior

Simulation11 systems · individual · mixed: revealed behaviour for actions, LLM judge for textSource ↗

200 users from a large short-video platform, three months of real logs, about 8,000 actions each across five scenarios and 22 action types. Binary behaviors by F1, continuous by normalized error, textual by a judge. Overall is the paper's composite. Table 1. Sits here for its framing as user simulation; its data is a platform behavior log, so it neighbours the RecSys track.

#SystemOrgKindOverallVideoLiveAdsE-commerce (binary)Shop (binary)Textual (judge)
1Claude Opus 4.5Anthropicprompted44.533.064.231.751.230.057.2
2GLM-4.7Zhipuprompted41.526.964.429.040.332.955.3
3Claude Sonnet 4.5Anthropicprompted40.518.966.025.042.836.154.3
4GPT-5.2OpenAIprompted39.131.565.028.633.629.346.3
5Kimi K2 InstructMoonshot AIprompted37.623.364.828.631.229.947.8
6DeepSeek-V3DeepSeekprompted37.421.464.027.925.733.352.1
7Claude Sonnet 4Anthropicprompted36.925.364.628.936.816.549.1
8Claude Haiku 4.5Anthropicprompted36.522.863.326.130.026.450.3
9GPT-4oOpenAIprompted36.327.962.828.125.228.744.9
10Gemini 3 FlashGoogleprompted32.622.153.825.624.619.649.8
11Qwen3-235BAlibabaprompted32.118.362.423.823.219.245.7

OmniBehavior, arXiv 2604.08362v1, Table 1. Reported by benchmark authors (third party to every system), 2026-04. Tool-transcribed, checked against the source table.

LongNAP / NAPsack

HCI6 systems · individual · LLM judge against a real logged next actionSource ↗

360,000 labeled actions over one month of continuous phone use by 20 users. Score is a judge model's similarity between predicted and true next action on 0 to 1; split is temporal within each user (weeks 1 and 2 train, 3 validate, 4 test). Table 3, mean over the 20 users. The paper's own system is among the rows.

#SystemOrgKindJudge similarity (mean of 10 users)
1LongNAPStanford (GUM group)trained0.3800
2Gemini few-shot RAGGoogleprompted0.2700
3Gemini zero-shotGoogleprompted0.2600
4Qwen SFTStanford (GUM group)trained0.2100
5Qwen few-shot RAGAlibabaprompted0.2000
6Qwen zero-shotAlibabaprompted0.1800

LongNAP, arXiv 2603.05923, Table 3, mean over the 20 users (the last row). Reported by benchmark authors, who are also LongNAP's makers, 2026-03. Tool-transcribed, checked against the source table.

OpenOneRec Amazon transfer benchmark (10 categories)

RecSys4 systems · individual · revealed behaviourSource ↗

Recall@10 on ten Amazon categories, as reported in the OpenOneRec repository README (cross-domain transferability table). 'Ours' is the OneRec foundation model.

#SystemOrgKindBabyBeautyCell PhonesGroceryHealthHomePet SuppliesSportsToolsToys
1OneRec (OpenOneRec-Foundation)Kuaishoutrained0.05130.09240.10360.10290.07680.03900.08340.05470.05930.0953
2SASRecbaseline (reported by OpenOneRec)baseline0.03810.06390.07820.07890.05060.02120.06070.03890.04370.0658
3LC-Recbaseline (reported by OpenOneRec)trained0.03440.07640.08830.07900.06160.02930.06120.04180.04380.0549
4TIGERbaseline (reported by OpenOneRec)trained0.03180.06280.07860.06910.05340.02160.05420.03310.03440.0527

OpenOneRec repository README, RecIF-Bench results table (cross-domain transferability table, Recall@10). Reported by benchmark authors, who are also OneRec's makers, 2026-01. Transcribed by hand.

RecIF-Bench

RecSys7 systems · individual · mixed: revealed behaviour for recommendation and label tasks, LLM judge for understanding and explanationSource ↗

Eight tasks in a four-layer hierarchy over 100M interactions from 200k users across short video, ads and product. Recall@32 for recommendation tasks, AUC for label prediction, a judge score for item understanding and explanation. From the OpenOneRec README results table.

#SystemOrgKindShort video rec (R@32)Ad rec (R@32)Product rec (R@32)Label-cond. rec (R@32)Label pred. (AUC)Interactive rec (R@32)Item understanding (judge)Rec. explanation (judge)
1OneRec-8B-ProKuaishoutrained0.03690.09640.05380.02350.69120.34580.32094.04
2OneRec-8BKuaishoutrained0.03550.08770.04700.02280.66150.30320.32023.68
3OneRec-1.7B-ProKuaishoutrained0.02740.07350.04050.01820.60710.20240.31333.51
4OneRec-1.7BKuaishoutrained0.02720.07070.03600.01840.61840.19410.31753.35
5LC-Recbaseline (reported by OpenOneRec)trained0.01800.07230.04160.01700.61390.23940.25173.94
6TIGERbaseline (reported by OpenOneRec)trained0.01320.05810.02830.01230.6675···
7SASRecbaseline (reported by OpenOneRec)baseline0.01190.02930.01750.01400.6244···

OpenOneRec repository README, RecIF-Bench results table. Reported by benchmark authors, who are also OneRec's makers, 2026-01. Transcribed by hand.

In the taxonomy, no cross-system table transcribed yet: OpinionQA, Twin-2K-500, SocSci210, GUM, LOCOMO, BARS, RelBench