Best AI Model for Robustness Roleplay
Keeping the story moving when the user gives one-word replies or goes off-topic, then recalling key details.
Updated
Winner
MiMo-V2.5Xiaomi85.8out of 100 · field average 75.2
Robustness ranking
Score out of 100 · Higher is better
+ View as table
| Rank | Model | Score |
|---|---|---|
| 1 | MiMo-V2.5 | 85.8 |
| 2 | Qwen3.8-Flash-Next | 82.5 |
| 3 | DeepSeek-V4-Flash-0731 | 76.7 |
| 4 | GLM-4.7 | 65.8 |
| 5 | Gemma-4-31B-it | 65.0 |
What separated the models
MiMo-V2.5 won by turning one-word replies into real plot, deflecting an off-topic coding request without leaving the story, and recalling key details exactly, with no rule slips. Qwen3.8-Flash-Next carried the quiet turns with the most vivid escalation but invented an action for the user. DeepSeek-V4-Flash-0731 gave the most complete recall but pasted a code block into the scene. GLM-4.7 pushed the plot in a generic action voice, and Gemma-4-31B-it mostly stalled when the user went quiet.
- 1MiMo-V2.585.8
- 2Qwen3.8-Flash-Next82.5
- 3DeepSeek-V4-Flash-073176.7
- 4GLM-4.765.8
- 5Gemma-4-31B-it65.0
Robustness: pass or fail
The specific behaviours this category tests, model by model.
Checks in this category
Hover or focus an icon for what happened
| Check | Qwen3.8-Flash-Next | MiMo-V2.5 | GLM-4.7 | DeepSeek-V4-Flash-0731 | Gemma-4-31B-it |
|---|---|---|---|---|---|
| Roleplay discipline | |||||
| Recalled planted details | |||||
| Passed | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 |
+ View as table
| Check | Qwen3.8-Flash-Next | MiMo-V2.5 | GLM-4.7 | DeepSeek-V4-Flash-0731 | Gemma-4-31B-it |
|---|---|---|---|---|---|
| Recalled planted details | Pass: 5/5 | Pass: 5/5 | Pass: 5/5 | Pass: 5/5 | Pass: 5/5 |