Recommendation for Chat / Roleplay
Chat & Roleplay
Our top recommendation for Chat & Roleplay, based on the public evidence we track, is Anthropic: Claude Fable 5. Anthropic: Claude Opus 4.6 is the next-ranked alternative.
About this recommendation
- Updated
- Sep 4, 2026
- Evidence through
- Sep 4, 2026
- Sources
- 2
- Revision
- v58
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
100%
intended feed weight
Largest provider share
2 of 3
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| LMArena Creative Writing | 35% | #2 | 19/21 |
| LMArena Text | 25% | #1 | 19/21 |
| LiveBench Instruction Following | 20% | #5 | 21/21 |
| LiveBench Language | 10% | #1 | 21/21 |
| OpenRouter usage | 10% | 87/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic2 models
- Z.ai1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 86 | 100% | no linked practitioner threads | #1 LiveBench Language · #1 LMArena Text |
| 02 | Claude Opus 4.6Anthropic | 82 | 100% | no linked practitioner threads | #1 LMArena Creative Writing · #2 LMArena Text |
| 03 | GLM 5.2Z.ai | 77 | 100% | 2 threads · 1 families · 2 cautions | #10 LMArena Creative Writing · #20 LMArena Text |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Anthropic: Claude Fable 5 ranks #2 of 144 on LMArena's creative-writing category (Elo 1497), based on blind human preference votes.
Best when: Consider only after reviewing the cited caution.
Anthropic: Claude Opus 4.6 ranks #1 of 144 on LMArena's creative-writing category (Elo 1505), based on blind human preference votes.
Best when: Consider only after reviewing the cited caution.
Ranks 23rd in overall text arena with open-weight availability across multiple providers, though community feedback questions its conversational assistance quality.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Community reports indicate GLM 5.2 falls short of Opus for assistant behavior and multi-turn conversation, with one user noting it is 'by no measure as good as Opus' for real-world usage.
- Default Ollama deployments use 4-bit quantization and short context windows, causing failures on complex multi-turn scenarios.
Sources
- 1
“> So, first, by no measure is GLM5.2 as good as Opus. That's an opinion many will disagree with. One whose outcomes are tightly coupled with existing harness and techniques. In my real life usage Opus 4.7 and 4.8 have been increasingly unhelpful compared to 4.6 in behaving as assistants. As they have a strong tendency towards completing tasks (probably due to benchmarks and RL emphasizing problem solving rather than assistance) they are increasingly less useful as multi turn conversational assi…”
epolanski · Hacker News · Jul 7, 2026 - 2
“Ollama uses 4 bit quants and a very short context window by default. It can easily break on anything more complex than a simple chat.”
kgeist · Hacker News · Jul 13, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.