| Model | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1Qwen3.8-Flash-NextAlibaba Cloud | Best: 79.9 | Best: 77.7 | Best: 82.2 | 68.3 | Best: 85.8 | Best: 85.0 | 66.7 | 82.5 | Best: 90.0 | 79.2 | 68.3 | Best: 86.7 | Best: 86.7 | 14.1 | 12/19 | Best: 256K |
| 2MiMo-V2.5Xiaomi | 72.5 | 73.7 | 71.3 | 67.5 | 71.7 | 73.3 | Best: 70.0 | Best: 85.8 | 65.8 | Best: 85.0 | 66.7 | 78.3 | 60.8 | 10.6 | 15/19 | 128K |
| 3GLM-4.7Zhipu AI | 67.6 | 66.2 | 69.0 | 73.3 | 66.7 | 55.0 | Best: 70.0 | 65.8 | 53.3 | 79.2 | Best: 83.3 | 66.7 | 62.5 | 39.8 | Best: 17/19 | 198K |
| 4DeepSeek-V4-Flash-0731DeepSeek | 67.3 | 65.0 | 69.5 | Best: 80.0 | 53.3 | 60.8 | 54.2 | 76.7 | 86.7 | 80.0 | 56.7 | 74.2 | 50.0 | 23.9 | 11/19 | Best: 256K |
| 5Gemma-4-31B-itGoogle DeepMind | 64.3 | 64.5 | 64.0 | 65.0 | 68.3 | 61.7 | 62.5 | 65.0 | 65.0 | 71.7 | 57.5 | 59.2 | 66.7 | Best: 6.3 | 15/19 | Best: 256K |
Scores out of 100. Reply time is the median in seconds (Core Roleplay round), including hidden reasoning. Checks are pass counts out of 19. Scroll sideways for every column.
Roleplay Index
September 2026Overall score out of 100 · Higher is better
Reply speed
Median seconds to a full reply · Lower is better
How to read it
The Roleplay Index is the average of the round scores; each round averages its five categories. Gaps under 1 point are within run-to-run variation, so treat them as ties.
Speed
Fastest median reply: Gemma-4-31B-it, 6.3 s. Measured through a hosted API during testing; real-world speed depends on your provider.
Reuse the data
Every number here is available as JSON and CSV, free to reuse with credit to AI Roleplay Bench.