Recommendation for Writing

Writing

Our top recommendation for Writing, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2][3] Deploy for long-form drafting and editing tasks where LiveBench shows strong performance on paraphrasing, simplifying, and summarization. Watch out: Watch for iterative refinement loops that may underperform, as one user reported disappointing results on repeated improvement requests. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Use for paraphrasing and simplifying tasks where benchmark coverage exists, though human preference lags behind top Claude and Gemini variants.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
13
Revision
v58

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

22

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

2 of 8

Anthropic

Provisional source breadth. 3 citation families and 1 practitioner families support the top result; 1 cautionary thread is retained. The largest citation family contributes 39%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
LMArena Creative Writing
30%
#220/22
LiveBench Instruction Following
25%
#522/22
LMArena Text
20%
#120/22
LiveBench Language
15%
#122/22
OpenRouter usage
10%
87/10022/22

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic25%
  • Anthropic2 models
  • Google1 model
  • Meta1 model
  • OpenAI1 model
  • Qwen1 model
  • xAI1 model
  • Z.ai1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
86
100%2 threads · 1 families · 1 cautions#1 LiveBench Language · #1 LMArena Text
02GPT-5.6 SolOpenAI
80
100%no linked practitioner threads#5 LiveBench Language · #10 LMArena Text
03Gemini 3.6 FlashGoogle
80
100%1 threads · 1 families · 0 cautions#7 LiveBench Instruction Following · #10 LiveBench Language
04Claude Opus 4.6Anthropic
80
100%1 threads · 1 families · 0 cautions#1 LMArena Creative Writing · #2 LMArena Text
05Muse Spark 1.2Meta
77
100%no linked practitioner threads#4 LMArena Text · #8 LiveBench Instruction Following
06Grok 4.6xAI
77
100%no linked practitioner threads#11 LiveBench Language · #15 LiveBench Instruction Following
07Qwen3.7 MaxQwen
75
100%no linked practitioner threads#10 LiveBench Instruction Following · #16 LMArena Text
08GLM 5.2Z.ai
74
100%5 threads · 1 families · 4 cautions#10 LMArena Creative Writing · #20 LMArena Text

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 ranks second in creative-writing human preference and scores well on instruction-following benchmarks covering story generation and summarization.

    Best when: Deploy for long-form drafting and editing tasks where LiveBench shows strong performance on paraphrasing, simplifying, and summarization.

    Tips

    • Deploy for long-form drafting and editing tasks where LiveBench shows strong performance on paraphrasing, simplifying, and summarization.
      Source 1
      Scores 75.77% on LiveBench Instruction Following (#5 of 51), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗
    • Use when you need creative writing quality near the top of human preference rankings without Opus pricing.
      Source 4
      Ranks #2 of 144 on LMArena's creative-writing category (Elo 1497), based on blind human preference votes.
      LMArena creative-writing categoryOpen original ↗

    Watch out for

    • Watch for iterative refinement loops that may underperform, as one user reported disappointing results on repeated improvement requests.
      Source 2
      I haven't tried this in a few months, but last time I tried a loop that rendered the pelican and asked for improvements the results were actually quite disappointing. Be interesting to try that again against GPT-5.6 at Claude Fable 5 though.
  2. GPT-5.6 Sol ranks thirteenth in creative-writing preference and sixteenth on instruction-following benchmarks.

    Best when: Use for paraphrasing and simplifying tasks where benchmark coverage exists, though human preference lags behind top Claude and Gemini variants.

    Tips

    • Use for paraphrasing and simplifying tasks where benchmark coverage exists, though human preference lags behind top Claude and Gemini variants.
      Source 3
      Scores 71.85% on LiveBench Instruction Following (#16 of 51), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗
  3. Gemini 3.6 Flash places in the top ten for creative-writing preference and scores competitively on instruction-following benchmarks.

    Best when: Use for marketing copy and text tasks where one tester found it outperformed DeepSeek v4 Flash Pro.

    Tips

    • Use for marketing copy and text tasks where one tester found it outperformed DeepSeek v4 Flash Pro.
      Source 5
      From my own testing, Gemini 3.5 3.6 Flash is better than DS v4 Flash Pro on text ability.
      jklmnopqrstuvwOpen original ↗
    • Deploy for instruction-following workflows including story generation, where it ranks seventh on LiveBench.
      Source 6
      Scores 75.37% on LiveBench Instruction Following (#7 of 51), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗
  4. Claude Opus 4.6 leads the LMArena creative-writing category with the highest Elo score, indicating strong human preference for its prose quality.

    Best when: Use for premium creative writing where human judges consistently preferred its output over 143 other models.

    Tips

    • Use for premium creative writing where human judges consistently preferred its output over 143 other models.
      Source 7
      Ranks #1 of 144 on LMArena's creative-writing category (Elo 1505), based on blind human preference votes.
      LMArena creative-writing categoryOpen original ↗

    Watch out for

    • Budget for $6.25 to $30.00 per 10,000 emails at typical length, making it costly for high-volume drafting workflows.
      Source 8
      Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."
  5. Muse Spark 1.2 ranks twenty-first in creative-writing preference and shows mid-tier language manipulation scores.

    Best when: Use for language manipulation tasks where LiveBench shows above-median performance.

    Tips

    • Use for language manipulation tasks where LiveBench shows above-median performance.
      Source 9
      Scores 78.57% on LiveBench Language (#27 of 51), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗
  6. Grok 4.6 ranks nineteenth in creative-writing preference and fifteenth on instruction-following benchmarks.

    Best when: Consider for story generation and summarization workflows where LiveBench shows moderate competence.

    Tips

    • Consider for story generation and summarization workflows where LiveBench shows moderate competence.
      Source 10
      Scores 71.87% on LiveBench Instruction Following (#15 of 51), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗
  7. Qwen3.7 Max sits in the upper third of creative-writing preference rankings and scores tenth on LiveBench instruction-following.

    Best when: Use for paraphrasing and simplifying tasks where benchmark scores indicate solid capability.

    Tips

    • Use for paraphrasing and simplifying tasks where benchmark scores indicate solid capability.
      Source 11
      Scores 74.04% on LiveBench Instruction Following (#10 of 51), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗
  8. GLM 5.2 is an open-weight model with reported gaps between benchmark performance and practical writing quality.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Expect to manually correct subtle errors in rewritten articles, as users report it misses mistakes that Claude Sonnet 4.6 catches.
      Source 12
      I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…
    • Verify implementation details carefully, as one user found it produced outdated library versions in a coding task compared to Claude Fable's detailed guides.
      Source 13
      I don't think the writer has used top tier models very much. I have subscriptions to basically every provider, the difference between glm5.2 and opus is not even close, the gap is huge. raw benchmarks glm is impressive , but in practice these models are lacking so much. I had fable create a detailed implementation guide that explained how to implement everything in immense detail, it included all the libraries to use and versions. I then had deepseek v4 pro execute and it used old versions , di…
      AgentMasterRaceOpen original ↗

Frequently asked

What is the top-ranked model for Writing?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Deploy for long-form drafting and editing tasks where LiveBench shows strong performance on paraphrasing, simplifying, and summarization.[1]
What should I watch out for with Anthropic: Claude Fable 5?
Watch for iterative refinement loops that may underperform, as one user reported disappointing results on repeated improvement requests.[2]
What is an alternative to Anthropic: Claude Fable 5?
OpenAI: GPT-5.6 Sol is the next-ranked option. Use for paraphrasing and simplifying tasks where benchmark coverage exists, though human preference lags behind top Claude and Gemini variants.[3]

Sources

  1. 1

    Scores 75.77% on LiveBench Instruction Following (#5 of 51), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  2. 2

    I haven't tried this in a few months, but last time I tried a loop that rendered the pelican and asked for improvements the results were actually quite disappointing. Be interesting to try that again against GPT-5.6 at Claude Fable 5 though.

    simonw · Hacker News · Jul 9, 2026
  3. 3

    Scores 71.85% on LiveBench Instruction Following (#16 of 51), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  4. 4

    Ranks #2 of 144 on LMArena's creative-writing category (Elo 1497), based on blind human preference votes.

    LMArena creative-writing category · Benchmark · Sep 1, 2026
  5. 5

    From my own testing, Gemini 3.5 3.6 Flash is better than DS v4 Flash Pro on text ability.

    jklmnopqrstuvw · Hacker News · Aug 13, 2026
  6. 6

    Scores 75.37% on LiveBench Instruction Following (#7 of 51), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  7. 7

    Ranks #1 of 144 on LMArena's creative-writing category (Elo 1505), based on blind human preference votes.

    LMArena creative-writing category · Benchmark · Sep 1, 2026
  8. 8

    Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."

    johndhi · Hacker News · Jun 26, 2026
  9. 9

    Scores 78.57% on LiveBench Language (#27 of 51), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026
  10. 10

    Scores 71.87% on LiveBench Instruction Following (#15 of 51), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  11. 11

    Scores 74.04% on LiveBench Instruction Following (#10 of 51), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  12. 12

    I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…

    sixtyj · Hacker News · Jun 30, 2026
  13. 13

    I don't think the writer has used top tier models very much. I have subscriptions to basically every provider, the difference between glm5.2 and opus is not even close, the gap is huge. raw benchmarks glm is impressive , but in practice these models are lacking so much. I had fable create a detailed implementation guide that explained how to implement everything in immense detail, it included all the libraries to use and versions. I then had deepseek v4 pro execute and it used old versions , di…

    AgentMasterRace · Hacker News · Jul 7, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.