AI Roleplay Leaderboard

Every model and every score in one table. Tap a column to sort; the best value in each column is marked with a dot.

Updated

Model
1Qwen3.8-Flash-NextAlibaba CloudBest: 79.9Best: 77.7Best: 82.268.3Best: 85.8Best: 85.066.782.5Best: 90.079.268.3Best: 86.7Best: 86.714.112/19Best: 256K
2MiMo-V2.5Xiaomi72.573.771.367.571.773.3Best: 70.0Best: 85.865.8Best: 85.066.778.360.810.615/19128K
3GLM-4.7Zhipu AI67.666.269.073.366.755.0Best: 70.065.853.379.2Best: 83.366.762.539.8Best: 17/19198K
4DeepSeek-V4-Flash-0731DeepSeek67.365.069.5Best: 80.053.360.854.276.786.780.056.774.250.023.911/19Best: 256K
5Gemma-4-31B-itGoogle DeepMind64.364.564.065.068.361.762.565.065.071.757.559.266.7Best: 6.315/19Best: 256K

Scores out of 100. Reply time is the median in seconds (Core Roleplay round), including hidden reasoning. Checks are pass counts out of 19. Scroll sideways for every column.

Roleplay Index

September 2026

Overall score out of 100 · Higher is better

Roleplay Index: The average of the Core Roleplay and Mature Themes & Limits rounds.

Within 1 point, so treat as ties: GLM-4.7 and DeepSeek-V4-Flash-0731.

Reply speed

Median seconds to a full reply · Lower is better

How to read it

The Roleplay Index is the average of the round scores; each round averages its five categories. Gaps under 1 point are within run-to-run variation, so treat them as ties.

Speed

Fastest median reply: Gemma-4-31B-it, 6.3 s. Measured through a hosted API during testing; real-world speed depends on your provider.

Reuse the data

Every number here is available as JSON and CSV, free to reuse with credit to AI Roleplay Bench.