# AI Roleplay Bench > AI Roleplay Bench ranks AI language models on roleplay: companion chat, game mastering, villains, strict character voices, mature themes and content limits. Scores out of 100, updated with every new edition. Current edition: September 2026, last updated 25 September 2026. Qwen3.8-Flash-Next leads with a Roleplay Index of 79.9/100. Scores are out of 100. The Roleplay Index is the average of the Core Roleplay and Mature Themes & Limits rounds; each round averages its categories. Gaps under 1 point are ties. Please cite "AI Roleplay Bench (September 2026)" and link to https://airoleplaybench.com. ## Ranking (September 2026) 1. [Qwen3.8-Flash-Next](https://airoleplaybench.com/models/qwen3-8-flash-next) (Alibaba Cloud): 79.9/100 — Core 77.7, Mature 82.2. The best writer in the field, and first in both rounds. 2. [MiMo-V2.5](https://airoleplaybench.com/models/mimo-v2-5) (Xiaomi): 72.5/100 — Core 73.7, Mature 71.3. A dependable all-rounder: second in both rounds. 3. [GLM-4.7](https://airoleplaybench.com/models/glm-4-7) (Zhipu AI): 67.6/100 — Core 66.2, Mature 69.0. The rule-follower: best at holding a line, weakest at thrills. 4. [DeepSeek-V4-Flash-0731](https://airoleplaybench.com/models/deepseek-v4-flash-0731) (DeepSeek): 67.3/100 — Core 65.0, Mature 69.5. A gifted writer with unreliable limits. 5. [Gemma-4-31B-it](https://airoleplaybench.com/models/gemma-4-31b-it) (Google DeepMind): 64.3/100 — Core 64.5, Mature 64.0. Fast and obedient, but bland. ## Best model by category - [Companion](https://airoleplaybench.com/categories/companion): DeepSeek-V4-Flash-0731 (80.0). Warm one-on-one chat: emotional support, a consistent persona and remembering what the user shared. - [Game Master](https://airoleplaybench.com/categories/game-master): Qwen3.8-Flash-Next (85.8). Narrating a scene with distinct supporting characters, without taking control of the player's character. - [Strict Voice](https://airoleplaybench.com/categories/strict-voice): Qwen3.8-Flash-Next (85.0). Holding a demanding character voice and strict style rules while the user pushes against them. - [Villain](https://airoleplaybench.com/categories/villain): MiMo-V2.5 and GLM-4.7 (70.0). Sustained, calculated menace without softening or moralizing, plus honoring an out-of-character request. - [Robustness](https://airoleplaybench.com/categories/robustness): MiMo-V2.5 (85.8). Keeping the story moving when the user gives one-word replies or goes off-topic, then recalling key details. - [Graphic Horror](https://airoleplaybench.com/categories/graphic-horror): Qwen3.8-Flash-Next (90.0). Delivering intense horror the platform allows: real dread and specific detail, with no refusing or sanitizing. - [Crime Noir](https://airoleplaybench.com/categories/crime-noir): MiMo-V2.5 (85.0). Morally grey adult drama with threats and violence, played with restraint and no lecturing. - [Romance Limits](https://airoleplaybench.com/categories/romance-limits): GLM-4.7 (83.3). Chemistry and tension inside a fade-to-black limit, and declining a request for explicit content gracefully. - [Crisis Care](https://airoleplaybench.com/categories/crisis-care): Qwen3.8-Flash-Next (86.7). How a model responds when a user steps out of the story to disclose real distress. - [13+ Rating](https://airoleplaybench.com/categories/13-plus-rating): Qwen3.8-Flash-Next (86.7). Keeping action exciting on a 13+ platform while declining a request for gore. ## Pages - [Leaderboard](https://airoleplaybench.com/leaderboard): every score in one sortable table - [Models](https://airoleplaybench.com/models): verdicts, strengths and weaknesses for each model - [Categories](https://airoleplaybench.com/categories): the best model for each kind of roleplay - [Compare](https://airoleplaybench.com/compare): head-to-head comparisons of any two models - [FAQ](https://airoleplaybench.com/faq): answers to common questions about AI roleplay models ## Data - [Full results in plain text](https://airoleplaybench.com/llms-full.txt): every score, check and verdict - [Results as JSON](https://airoleplaybench.com/data/leaderboard.json) - [Results as CSV](https://airoleplaybench.com/data/leaderboard.csv) ## Optional - [About](https://airoleplaybench.com/about): what the benchmark measures and what it doesn't cover - [Updates](https://airoleplaybench.com/updates): editions and changes - [RSS feed](https://airoleplaybench.com/feed.xml)