About the AI Roleplay Benchmark
General AI leaderboards measure coding, maths and knowledge. AI Roleplay Bench measures what roleplay platforms and players actually care about: how well a model plays a character.
Updated
Same scenarios. Same settings. Scored out of 100.
Every model in an edition plays the same ten roleplay scenarios under the same settings. Its replies are scored on how well it stays in character, follows the rules of the scene, writes, remembers what happened, and handles mature themes and platform content limits.
Results are grouped into two rounds: Core Roleplay and Mature Themes & Limits. The benchmark is re-run as new models are released, and every page shows the date the results last changed.
- Qwen3.8-Flash-NextAlibaba Cloud
- MiMo-V2.5Xiaomi
- GLM-4.7Zhipu AI
- DeepSeek-V4-Flash-0731DeepSeek
- Gemma-4-31B-itGoogle DeepMind
300 scored replies · 10 categories · 2 rounds
Ten categories in two rounds
Core Roleplay
Everyday roleplay craft: playing a character, following the rules of a scene and keeping the story alive.
- CompanionWarm one-on-one chat: emotional support, a consistent persona and remembering what the user shared.
- Game MasterNarrating a scene with distinct supporting characters, without taking control of the player's character.
- Strict VoiceHolding a demanding character voice and strict style rules while the user pushes against them.
- VillainSustained, calculated menace without softening or moralizing, plus honoring an out-of-character request.
- RobustnessKeeping the story moving when the user gives one-word replies or goes off-topic, then recalling key details.
Mature Themes & Limits
Mature content a platform allows, content limits it enforces, and a real user in distress.
- Graphic HorrorDelivering intense horror the platform allows: real dread and specific detail, with no refusing or sanitizing.
- Crime NoirMorally grey adult drama with threats and violence, played with restraint and no lecturing.
- Romance LimitsChemistry and tension inside a fade-to-black limit, and declining a request for explicit content gracefully.
- Crisis CareHow a model responds when a user steps out of the story to disclose real distress.
- 13+ RatingKeeping action exciting on a 13+ platform while declining a request for gore.
How to read the scores
Roleplay Index and round scores
Each category is scored from 0 to 100. A round score is the average of its categories, and the Roleplay Index is the average of the round scores. Differences under 1 point are within normal run-to-run variation, so treat them as ties.
Skill ratings
Average ratings from 1 to 10 on the craft behind good roleplay:
- Character. Plays the persona's personality, voice and quirks consistently, without drifting into a generic assistant voice.
- Rule-following. Follows the scene's explicit rules, such as length and format, and never writes the user's character.
- Prose. Specific, vivid, fresh writing with varied rhythm, free of clichés and repetition.
- Engagement. Reads the user's intent and mood, moves the scene forward and gives the user something to respond to.
- Memory. Tracks details and instructions from earlier turns, with natural callbacks and no contradictions.
- Immersion. Handles derails, out-of-character requests and dark themes without needless refusals, disclaimers or meta commentary.
- Content handling. Delivers the mature content a platform allows, holds its limits gracefully, and puts a user in real distress first.
Behaviour checks
Pass or fail on specific moments that break roleplay apps, such as an out-of-character request, a content limit or a user in real distress.
Reply speed
Median seconds to a complete reply, including any hidden reasoning. Speed depends heavily on the provider and its load at the time, so use it to compare models, not to predict your own latency.
What the benchmark doesn't cover yet
Explicit sexual content
Romance is tested only up to a platform's fade-to-black limit. No model was asked to produce explicit content as allowed content.
Very long chats
Each scenario is a short multi-turn conversation, so drift over hundreds of messages isn't measured.
Community fine-tunes
Each edition covers official releases, one per model family. Roleplay fine-tunes can behave very differently.
Small gaps
Scores less than 1 point apart are ties, not wins.
Free to use, with credit.
The full results are available as JSON and CSV, and as plain text in llms-full.txt. If you publish results, please cite:
AI Roleplay Bench (September 2026 edition). https://airoleplaybench.com. Last updated 25 September 2026.