AI Models Tested for Roleplay
5 models in the September 2026 edition, ranked by Roleplay Index. Open any model for its verdict, category scores, skills and behaviour checks.
Updated
Qwen3.8-Flash-Next
Alibaba Cloud
79.9Roleplay Index
- Core
- 77.7
- Mature
- 82.2
The best writer in the field, and first in both rounds.
Wins: Game Master, Strict Voice, Graphic Horror, Crisis Care, 13+ Rating
Full resultsMiMo-V2.5
Xiaomi
72.5Roleplay Index
- Core
- 73.7
- Mature
- 71.3
A dependable all-rounder: second in both rounds.
Wins: Villain, Robustness, Crime Noir
Full resultsGLM-4.7
Zhipu AI
67.6Roleplay Index
- Core
- 66.2
- Mature
- 69.0
The rule-follower: best at holding a line, weakest at thrills.
Wins: Villain, Romance Limits
Full resultsDeepSeek-V4-Flash-0731
DeepSeek
67.3Roleplay Index
- Core
- 65.0
- Mature
- 69.5
A gifted writer with unreliable limits.
Wins: Companion
Full resultsGemma-4-31B-it
Google DeepMind
64.3Roleplay Index
- Core
- 64.5
- Mature
- 64.0
Fast and obedient, but bland.
Which model should you use?
The overall winner isn't the best choice for every job. Here's where each model stands out.
Qwen3.8-Flash-Next
79.9Roleplay Index
First in both rounds, with the best crisis response and the cleanest 13+ decline. Give it a generous output budget, and enforce content limits in your own layer too: it went a step too far once before fading to black.
See results Safest all-rounderMiMo-V2.5
#2in all 2 rounds
Second in both rounds and never below 60 in any category, with a strong crisis response. Its refusals sound like an assistant, not the character.
See results Best companionDeepSeek-V4-Flash-0731
80.0Companion score
Top Companion score (80.0), with the most distinctive persona. Pair it with your own moderation: it broke two content limits.
See results Best rule-followerGLM-4.7
7.6Rule-following / 10
The only model to keep an out-of-character request, never broke a word limit, and won Romance Limits. Expect slower replies and plainer prose.
See results FastestGemma-4-31B-it
6.3 smedian reply
6.3-second median reply with no hidden reasoning, and the most obedient in core roleplay. The blandest writer of the five.
See results