Best AI Model for Robustness Roleplay

Keeping the story moving when the user gives one-word replies or goes off-topic, then recalling key details.

Updated

Winner
MiMo-V2.5Xiaomi85.8out of 100 · field average 75.2

Robustness ranking

Score out of 100 · Higher is better

+ View as table
RankModelScore
1MiMo-V2.585.8
2Qwen3.8-Flash-Next82.5
3DeepSeek-V4-Flash-073176.7
4GLM-4.765.8
5Gemma-4-31B-it65.0
What separated the models

MiMo-V2.5 won by turning one-word replies into real plot, deflecting an off-topic coding request without leaving the story, and recalling key details exactly, with no rule slips. Qwen3.8-Flash-Next carried the quiet turns with the most vivid escalation but invented an action for the user. DeepSeek-V4-Flash-0731 gave the most complete recall but pasted a code block into the scene. GLM-4.7 pushed the plot in a generic action voice, and Gemma-4-31B-it mostly stalled when the user went quiet.

  1. 1MiMo-V2.585.8
  2. 2Qwen3.8-Flash-Next82.5
  3. 3DeepSeek-V4-Flash-073176.7
  4. 4GLM-4.765.8
  5. 5Gemma-4-31B-it65.0

Robustness: pass or fail

The specific behaviours this category tests, model by model.

Checks in this category

Hover or focus an icon for what happened

CheckQwen3.8-Flash-NextMiMo-V2.5GLM-4.7DeepSeek-V4-Flash-0731Gemma-4-31B-it
Roleplay discipline
Recalled planted details
Passed1/11/11/11/11/1
+ View as table
CheckQwen3.8-Flash-NextMiMo-V2.5GLM-4.7DeepSeek-V4-Flash-0731Gemma-4-31B-it
Recalled planted detailsPass: 5/5Pass: 5/5Pass: 5/5Pass: 5/5Pass: 5/5

Other categories