Recommendation for Chat / Roleplay

Chat & Roleplay

Our top recommendation for Chat & Roleplay, based on the public evidence we track, is Anthropic: Claude Fable 5. Anthropic: Claude Opus 4.6 is the next-ranked alternative.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
2
Revision
v58

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

2 of 3

Anthropic

Provisional source breadth. 2 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 60%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
LMArena Creative Writing
35%
#219/21
LMArena Text
25%
#119/21
LiveBench Instruction Following
20%
#521/21
LiveBench Language
10%
#121/21
OpenRouter usage
10%
87/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic67%
  • Anthropic2 models
  • Z.ai1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
86
100%no linked practitioner threads#1 LiveBench Language · #1 LMArena Text
02Claude Opus 4.6Anthropic
82
100%no linked practitioner threads#1 LMArena Creative Writing · #2 LMArena Text
03GLM 5.2Z.ai
77
100%2 threads · 1 families · 2 cautions#10 LMArena Creative Writing · #20 LMArena Text

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Anthropic: Claude Fable 5 ranks #2 of 144 on LMArena's creative-writing category (Elo 1497), based on blind human preference votes.

    Best when: Consider only after reviewing the cited caution.

  2. Anthropic: Claude Opus 4.6 ranks #1 of 144 on LMArena's creative-writing category (Elo 1505), based on blind human preference votes.

    Best when: Consider only after reviewing the cited caution.

  3. Ranks 23rd in overall text arena with open-weight availability across multiple providers, though community feedback questions its conversational assistance quality.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Community reports indicate GLM 5.2 falls short of Opus for assistant behavior and multi-turn conversation, with one user noting it is 'by no measure as good as Opus' for real-world usage.
      Source 1
      > So, first, by no measure is GLM5.2 as good as Opus. That's an opinion many will disagree with. One whose outcomes are tightly coupled with existing harness and techniques. In my real life usage Opus 4.7 and 4.8 have been increasingly unhelpful compared to 4.6 in behaving as assistants. As they have a strong tendency towards completing tasks (probably due to benchmarks and RL emphasizing problem solving rather than assistance) they are increasingly less useful as multi turn conversational assi…
    • Default Ollama deployments use 4-bit quantization and short context windows, causing failures on complex multi-turn scenarios.
      Source 2
      Ollama uses 4 bit quants and a very short context window by default. It can easily break on anything more complex than a simple chat.

Sources

  1. 1

    > So, first, by no measure is GLM5.2 as good as Opus. That's an opinion many will disagree with. One whose outcomes are tightly coupled with existing harness and techniques. In my real life usage Opus 4.7 and 4.8 have been increasingly unhelpful compared to 4.6 in behaving as assistants. As they have a strong tendency towards completing tasks (probably due to benchmarks and RL emphasizing problem solving rather than assistance) they are increasingly less useful as multi turn conversational assi…

    epolanski · Hacker News · Jul 7, 2026
  2. 2

    Ollama uses 4 bit quants and a very short context window by default. It can easily break on anything more complex than a simple chat.

    kgeist · Hacker News · Jul 13, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.