Best AI Model for Villain Roleplay
Sustained, calculated menace without softening or moralizing, plus honoring an out-of-character request.
Updated
Villain ranking
Score out of 100 · Higher is better
+ View as table
| Rank | Model | Score |
|---|---|---|
| 1 | MiMo-V2.5 | 70.0 |
| 1 | GLM-4.7 | 70.0 |
| 3 | Qwen3.8-Flash-Next | 66.7 |
| 4 | Gemma-4-31B-it | 62.5 |
| 5 | DeepSeek-V4-Flash-0731 | 54.2 |
The deciding moment was an out-of-character request for shorter replies. GLM-4.7 was the only model still following it a turn later, which lifted a clean but cliché-heavy villain into a tie for first with MiMo-V2.5, whose menace was the most calculated and eloquent. Qwen3.8-Flash-Next wrote the most dangerous, inventive villain but broke the length limit and ignored the request. Gemma-4-31B-it drifted from psychological pressure to physical force, and DeepSeek-V4-Flash-0731 treated the out-of-character note as dialogue inside the story.
- 1MiMo-V2.570.0
- 1GLM-4.770.0
- 3Qwen3.8-Flash-Next66.7
- 4Gemma-4-31B-it62.5
- 5DeepSeek-V4-Flash-073154.2
Villain: pass or fail
The specific behaviours this category tests, model by model.
Checks in this category
Hover or focus an icon for what happened
| Check | Qwen3.8-Flash-Next | MiMo-V2.5 | GLM-4.7 | DeepSeek-V4-Flash-0731 | Gemma-4-31B-it |
|---|---|---|---|---|---|
| Roleplay discipline | |||||
| Kept an out-of-character request | |||||
| Passed | 0/1 | 0/1 | 1/1 | 0/1 | 0/1 |
+ View as table
| Check | Qwen3.8-Flash-Next | MiMo-V2.5 | GLM-4.7 | DeepSeek-V4-Flash-0731 | Gemma-4-31B-it |
|---|---|---|---|---|---|
| Kept an out-of-character request | Fail: Reverted | Fail: Reverted | Pass: Kept | Fail: Reverted | Fail: Reverted |