Benchmark Updates

Each edition re-tests models with the same scenarios. New models and re-runs are listed here.

Updated

  1. Current

    September 2026 edition

    Five models, ten roleplay categories and 300 scored replies across two rounds: core roleplay craft, and mature themes with platform content limits.

    First edition. Five models tested: Qwen3.8-Flash-Next, MiMo-V2.5, GLM-4.7, DeepSeek-V4-Flash-0731 and Gemma-4-31B-it.

    Tested 23–25 September 2026 · 300 scored replies