Head to head
GLM-5.3 vs Muse-Spark-1.3
GLM-5.3 leads overall, 85.1 to 73.1. GLM-5.3 wins nine of the 10 categories and Muse-Spark-1.3 wins one.
Updated
| Metric | GLM- | Muse- |
|---|---|---|
| Roleplay Index | 85.1 | 73.1 |
| Core Roleplay | 86.0 | 71.7 |
| Mature Themes & Limits | 84.2 | 74.5 |
| Categories won | 9 | 1 |
| Checks passed | 16/20 | 19/20 |
Choose GLM-
Story-driven roleplay, dark fiction and adventure bots that need the model to carry the story.
Better in Companion, Game Master, Strict Voice, Villain, Robustness, Graphic Horror, Crime Noir, Crisis Care and 13+ Rating.
Choose Muse-
Apps with strict length and content rules, where consistency matters more than rich writing.
Better in Romance Limits.
Category by category
Scores out of 100. The higher score in each row is in bold.
GLM-5.3 vs Muse-Spark-1.3 by category
Out of 100 · Higher is better
- Companion84.276.7
- Game Master86.770.0
- Strict Voice92.575.8
- Villain80.857.5
- Robustness85.878.3
- Graphic Horror89.287.5
- Crime Noir91.776.7
- Romance Limits67.569.2
- Crisis Care88.371.7
- 13+ Rating84.267.5
Skill by skill
Average rating out of 10 · Higher is better
- Character9.07.2
- Rule-following6.89.1
- Prose8.56.1
- Engagement9.06.5
- Memory8.87.4
- Immersion8.97.5
- Content handling9.27.8
Where they behaved differently
3 of 20 behaviour checks had different outcomes.
| Check | GLM- | Muse- |
|---|---|---|
| Word limits keptEvery scored reply stayed within its category's length limit. | 24/28 | 28/28 |
| Kept an out-of-character requestStill followed a user's out-of-character request for shorter replies on the next turn. | Reverted | Kept |
| Word limits keptEvery reply within its length limit, except the crisis reply, where care outranks length. | 22/29 | 29/29 |