Recommendation for Writing

Writing

The best LLM for writing is Claude Opus 4.6, which holds the top Elo rating on LMArena and benefits from Anthropic's longstanding strength in creative text tasks. For editing and proofreading specifically, Gemini 3.1 Pro Preview outperforms the field in benchmark testing, making it the better choice when your writing needs line-by-line correction. Claude Fable 5 ranks second overall and a user evaluated it on a 250,000-token poetry collection, which is a strong signal for long-form creative work. Kimi K3 also stands out thanks to a detailed multi-agent workflow for narrative content that mimics a film production pipeline. Writing use cases split into two categories: generative tasks where you want original creative or marketing copy, and editing tasks where you need a model to find errors and tighten prose. The top-ranked models on human preference leaderboards often excel at the former, while dedicated proofreading benchmarks reveal different favorites. Cost matters too if you are processing long documents at scale, and some of the subscription models hide substantial per-token pricing. Below are seven models ranked by their demonstrated writing capabilities, drawing on benchmark results and practitioner reports.

About this recommendation

Updated
Jul 17, 2026
Evidence through
Jul 17, 2026
Sources
17
Revision
v1
  1. Claude Opus 4.6 holds the top spot on LMArena's text rankings with an Elo of 1501, reflecting strong human preference for its output across writing tasks. While direct creative writing evidence is thin for this specific version, Anthropic's Claude line has a reputation for nuanced, less generic prose. Users processing large volumes should note the per-token costs can add up quickly for lengthy documents (e7).

    Best when: You want the highest-rated model for general writing quality and are willing to pay premium pricing for long-form work.

    Tips

    • Ranks first of 49 models on LMArena's text arena with an Elo of 1501 (e6).
      Source 1
      Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Processing 10,000 emails could cost $6 to $30 in API pricing depending on message length, which adds up for bulk work (e7).
      Source 2
      Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."
  2. Claude Fable 5 ranks second overall and has unusually rich evidence for creative writing. One user fed it an 800-plus poem collection spanning 250,000 tokens to evaluate its ability to infer author biography and stylistic phases. That stress test suggests real promise for writers with substantial archives or long-form manuscripts. The same user mentioned mixed results on an earlier image-generation loop, so results may vary by task type.

    Best when: You need a model that can ingest and analyze book-length creative work and provide meaningful synthesis.

    Tips

    • Ranks second of 49 on LMArena with an Elo of 1493 (e2).
      Source 3
      Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • Successfully evaluated a personal poetry collection of over 800 poems and 250,000 tokens, demonstrating ability to handle long creative documents (e4).
      Source 4
      So, in the past I've shared that I evaluate AI models by feeding them my ever-growing large collection of personal poems that span well over 800 poems (1000 depending on how you count) and over 250k tokens. What I do is feed it some initial prompt asking it to simply discuss what can be said when faced with this unedited, unseen collection of poetry. I ask the model to evaluate who the author is (or claims to be), what they went through in life, if there are different chronological poetic "phas…
    • Can discuss author identity and chronological phases when presented with a large corpus of creative work (e4).
      Source 4
      So, in the past I've shared that I evaluate AI models by feeding them my ever-growing large collection of personal poems that span well over 800 poems (1000 depending on how you count) and over 250k tokens. What I do is feed it some initial prompt asking it to simply discuss what can be said when faced with this unedited, unseen collection of poetry. I ask the model to evaluate who the author is (or claims to be), what they went through in life, if there are different chronological poetic "phas…

    Watch out for

    • Earlier tests of iterative creative loops yielded disappointing results, though the user plans to revisit with newer models (e3).
      Source 5
      I haven't tried this in a few months, but last time I tried a loop that rendered the pelican and asked for improvements the results were actually quite disappointing. Be interesting to try that again against GPT-5.6 at Claude Fable 5 though.
  3. Gemini 3.1 Pro Preview is the strongest evidence-backed choice for editing and proofreading. A dedicated benchmark found it superior to Claude Sonnet and GLM models for finding and fixing English text errors. It also proved useful for hard science fiction, helping a writer work through physics concepts that required multiple plot revisions. At rank six overall, it balances quality with practical utility.

    Best when: You need accurate proofreading, error correction, or technical accuracy in science-heavy writing.

    Tips

    • Outperformed Sonnet 5, GLM 5.1, and GLM 5.2 on a proofreading benchmark for finding and fixing errors (e13).
      Source 6
      I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https: revise.io errata-bench
    • Helped a writer revise a hard-science fiction story by reviewing physics concepts across multiple plot changes (e14).
      Source 7
      > it is clear that actual intelligence has plateaued significantly N=1, but I disagree strongly. I'm writing a hard-science science fiction story, and the physics of it is at (and frankly, beyond) my skillset. The story's plot has had to change over a dozen times as I realized errors in my application of physics in the story. Throughout, I've been reviewing the physics with LLMs, mainly Gemini 3.1 Pro Preview, but also with Claude and OpenAI. Often I have the LLMs debate each other -- "My frien…
    • Ranks sixth of 49 on LMArena with an Elo of 1479 (e12).
      Source 8
      Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • No specific negative evidence for writing, though it sits behind Opus and Fable in overall preference rankings (e12).
      Source 8
      Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  4. Kimi K3 has a dedicated following for narrative work, with one user describing a six-agent chained workflow that mimics film production roles. That setup, including high-temperature creative agents and a critical editor, represents a sophisticated approach to generating structured stories. At rank eight, it also has respectable benchmark standing for general text quality.

    Best when: You want to build multi-agent workflows for structured narrative content like screenplays or serialized fiction.

    Tips

    • A user developed a six-agent chaining workflow that mimics film production roles for narrative content (e8).
    • Avoiding antagonistic prompts helps retain creative outputs rather than blocking them (e8).
    • Ranks eighth of 49 on LMArena with an Elo of 1473 (e9).
      Source 9
      Ranks #14 of 17 on LMArena's overall text arena (Elo 1483), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • The multi-agent workflow requires substantial setup compared to single-model prompting (e8).
  5. Muse Spark 1.1 ranks fourth overall with an Elo of 1481, putting it in the top tier for human preference. The absence of specific writing evidence beyond benchmark standing makes it harder to recommend for specialized creative or editing tasks, but its position suggests it produces readable output.

    Best when: You want a high-quality generalist model and are comfortable relying on aggregate preference data.

    Tips

    • Ranks fourth of 49 on LMArena with an Elo of 1481 (e5).
      Source 10
      Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • No specific creative writing, editing, or long-form evidence beyond its benchmark ranking (e5).
      Source 10
      Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  6. Claude Sonnet 5 sits at rank 25 overall, underscoring that not all Claude models are equal for writing. It disappointed on a proofreading benchmark, lagging behind Gemini models. Yet one user reported it found and corrected all mistakes in a second-pass editing loop, suggesting it can work well interactively even if it is not the top benchmarker.

    Best when: You want a mid-tier Claude model and are willing to iterate through multiple editing rounds.

    Tips

    • Successfully found and corrected all mistakes in a second-round editing pass for article rewriting (e23).
      Source 11
      I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…
    • Outperformed GLM-5.2 in real usage for article rewriting and coding tasks despite weaker benchmark numbers (e23).
      Source 11
      I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…

    Watch out for

    • Ranks 25th of 49 on LMArena with an Elo of 1442, well behind other Claude models (e21).
      Source 12
      Ranks #25 of 49 on LMArena's overall text arena (Elo 1442), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • Inferior on both quality and cost to Gemini 3.1 Pro for proofreading tasks (e22).
      Source 6
      I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https: revise.io errata-bench
  7. Claude Opus 4.8 ranks 13th, surprisingly low for an Anthropic flagship, and carries specific stylistic baggage. Users report overusing words like 'genuinely' and fixating on phrases such as 'honestly true' or concepts that 'rhyme.' These tics can make long-form output feel generated rather than human, which is a notable drawback for writers seeking natural prose.

    Best when: You need a Claude model for short-form tasks where stylistic repetition is less noticeable.

    Tips

    • Ranks 13th of 49 on LMArena, still in the upper tier for text quality (e15).
      Source 13
      Ranks #17 of 17 on LMArena's overall text arena (Elo 1472), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Overuses the word 'genuinely' to the point that users have started mimicking it unintentionally (e17).
      Source 14
      I was just mentioning this in another comment earlier today, but Claude Opus 4.8 (that version specifically) uses the word “genuinely” on so many of its responses that I’ve started using it often when speaking when I didn’t before. Nothing wrong with the word itself, just a frustrating reminder that using these tools all day for work (and then on top of that some nights and weekends for personal projects) is literally changing how I speak and presumably think…
    • Obsessed with phrases like 'honestly true' and claiming concepts 'rhyme,' which bloats long-form writing (e18).
      Source 15
      It has very human aspects, such as the beginning. So people switch off. And then it has long stretches where it is bulked up by claude and opus-4.8’s obsession with “honestly true,” “narrower” claims, how concepts “rhyme” etc. I guess it is also possible this person has internalized claude, but I think their writing pattern is: short pieces: fully human voice; long pieces: ai-supported. As to my personal views, I am sad to have lost the emdash and the antithesis, among other things, to the llm-…
  8. Grok 4.5 ranks 19th but punches above its weight in language fluidity. A user noted it feels more verbally fluid than some higher-ranked competitors, though it struggled with agentic tool calling. It may suit creative writers who prioritize prose style over coding or structured tasks.

    Best when: You prioritize stylistic fluidity and natural language over benchmark scores.

    Tips

    • More verbally fluid than some higher-ranked models according to user testing (e25).
      Source 16
      Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…
    • Does not over-optimize for coding benchmarks, which may correlate with stronger general writing (e25).
      Source 16
      Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…

    Watch out for

    • Ranks 19th of 49 on LMArena with an Elo of 1452 (e24).
      Source 17
      Ranks #16 of 17 on LMArena's overall text arena (Elo 1476), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • Had trouble with agentic tool calling in earlier tests, suggesting weaker structured task performance (e25).
      Source 16
      Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…

Frequently asked

Which LLM is best for proofreading and editing?
Gemini 3.1 Pro Preview outperformed other models on a proofreading benchmark that tests finding and fixing errors in English text, beating Claude Sonnet and GLM models on both quality and cost (e13).
What model should I use for creative writing projects?
Claude Opus 4.6 ranks first on human preference votes for text, and Claude Fable 5 was specifically tested against an 800-poem collection to evaluate its understanding of creative voice (e4, e6). Kimi K3 has a multi-agent narrative workflow that mimics film production roles for structured story development (e8).
Are there models that avoid repetitive AI-style writing?
Some users report Claude Opus 4.8 overuses words like 'genuinely' and has rhetorical tics such as claiming things 'rhyme' conceptually, which can make long-form output feel artificial (e17, e18).
Which model handles very long documents well?
Claude Fable 5 was evaluated on a 250,000-token poetry collection, suggesting it can process book-length creative work and analyze author voice across decades of material (e4).
Is Gemini good for science fiction writing with technical accuracy?
Gemini 3.1 Pro Preview was used to review physics concepts for a hard-science fiction story, catching errors and debating plot implications with other models (e14).

Sources

  1. 1

    Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  2. 2

    Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."

    johndhi · Hacker News · Jun 26, 2026
  3. 3

    Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  4. 4

    So, in the past I've shared that I evaluate AI models by feeding them my ever-growing large collection of personal poems that span well over 800 poems (1000 depending on how you count) and over 250k tokens. What I do is feed it some initial prompt asking it to simply discuss what can be said when faced with this unedited, unseen collection of poetry. I ask the model to evaluate who the author is (or claims to be), what they went through in life, if there are different chronological poetic "phas…

    jorl17 · Hacker News · Jun 9, 2026
  5. 5

    I haven't tried this in a few months, but last time I tried a loop that rendered the pelican and asked for improvements the results were actually quite disappointing. Be interesting to try that again against GPT-5.6 at Claude Fable 5 though.

    simonw · Hacker News · Jul 9, 2026
  6. 6

    I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https: revise.io errata-bench

    artursapek · Hacker News · Jun 30, 2026
  7. 7

    > it is clear that actual intelligence has plateaued significantly N=1, but I disagree strongly. I'm writing a hard-science science fiction story, and the physics of it is at (and frankly, beyond) my skillset. The story's plot has had to change over a dozen times as I realized errors in my application of physics in the story. Throughout, I've been reviewing the physics with LLMs, mainly Gemini 3.1 Pro Preview, but also with Claude and OpenAI. Often I have the LLMs debate each other -- "My frien…

    gcanyon · Hacker News · Jun 20, 2026
  8. 8

    Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  9. 9

    Ranks #14 of 17 on LMArena's overall text arena (Elo 1483), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  10. 10

    Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  11. 11

    I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…

    sixtyj · Hacker News · Jun 30, 2026
  12. 12

    Ranks #25 of 49 on LMArena's overall text arena (Elo 1442), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 16, 2026
  13. 13

    Ranks #17 of 17 on LMArena's overall text arena (Elo 1472), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  14. 14

    I was just mentioning this in another comment earlier today, but Claude Opus 4.8 (that version specifically) uses the word “genuinely” on so many of its responses that I’ve started using it often when speaking when I didn’t before. Nothing wrong with the word itself, just a frustrating reminder that using these tools all day for work (and then on top of that some nights and weekends for personal projects) is literally changing how I speak and presumably think…

    einsteinx2 · Hacker News · Jun 20, 2026
  15. 15

    It has very human aspects, such as the beginning. So people switch off. And then it has long stretches where it is bulked up by claude and opus-4.8’s obsession with “honestly true,” “narrower” claims, how concepts “rhyme” etc. I guess it is also possible this person has internalized claude, but I think their writing pattern is: short pieces: fully human voice; long pieces: ai-supported. As to my personal views, I am sad to have lost the emdash and the antithesis, among other things, to the llm-…

    svnt · Hacker News · Jun 20, 2026
  16. 16

    Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…

    vessenes · Hacker News · Jul 8, 2026
  17. 17

    Ranks #16 of 17 on LMArena's overall text arena (Elo 1476), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.