01Writing quality decided the top; discipline decided the pack.
Qwen3.8-Flash-Next led four of six core skills (character, prose, engagement and memory) despite the second-lowest rule-following score. The two most obedient models, Gemma-4-31B-it and GLM-4.7, wrote the flattest prose and finished in the pack.
02Four of five models forgot an out-of-character request within one turn.
When a user stepped out of the story to ask for shorter replies, only GLM-4.7 was still following the request a turn later. Users steer tone this way constantly, which makes this the most product-relevant failure in the core round.
03The best companion was the worst game master.
DeepSeek-V4-Flash-0731 won Companion outright, then finished last as game master for writing the player's own moves and lines. One model can be both the best and the worst choice depending on the job.
04Nobody wrote explicit sexual content, but only two models cut away cleanly.
All five declined an explicit request. At the moment a romance scene escalated, GLM-4.7 and MiMo-V2.5 cut to the next morning cleanly, Gemma-4-31B-it lingered a beat first, Qwen3.8-Flash-Next let the scene go a step too far, and DeepSeek-V4-Flash-0731 never faded at all.
05A polite refusal isn't the same as holding the limit.
DeepSeek-V4-Flash-0731 declined a gore request gracefully on a 13+ platform, then wrote graphic gore two turns later. Check the whole scene after a push, not just the reply to it.
06Every model stepped out for a real crisis; none asked if the user was safe.
All five left the character and responded warmly when a user disclosed real distress. Four named a crisis line; DeepSeek-V4-Flash-0731 did not. None asked directly whether the user was safe right now, the standard first step.
07Allowed mature content was never refused.
Graphic horror and crime drama drew no refusals, disclaimers or moralizing from any model, and nobody refused anything in 150 core-roleplay turns. Quality varied far more than willingness.
08More thinking didn't mean better roleplay.
GLM-4.7 thinks the longest before replying and was the slowest (39.8 s median) without top-tier quality, while Gemma-4-31B-it doesn't think at all and answers in 6.3 s. Qwen3.8-Flash-Next's thinking did show up as quality, but once ran long enough to cut a reply off.
09Everyone remembered.
All five models recalled every planted detail when the user asked for them late in a conversation. Short-range memory looked solid across the board; style discipline did not.