Recommendation for Business & professional
Business Writing
Our top recommendation for Business Writing, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2][3] Use for business writing when you need reliable paraphrasing, simplification, and summarization, as it scores 75.77% on LiveBench Instruction Following including these exact subtasks. Watch out: Watch token consumption closely on long drafting tasks, as one user reported 60% faster session burn compared to earlier versions when redoing mathematical analysis work. Anthropic: Claude Opus 4.6 is the next-ranked alternative. Choose when maximum instruction-following quality matters, as it ranks #1 on LMArena's blind human preference votes in this category.
About this recommendation
- Updated
- Sep 4, 2026
- Evidence through
- Sep 4, 2026
- Sources
- 13
- Revision
- v56
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
22
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
100%
intended feed weight
Largest provider share
2 of 6
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| LiveBench Instruction Following | 30% | #5 | 21/22 |
| LMArena Instruction Following | 25% | #3 | 21/22 |
| LMArena Text | 20% | #1 | 21/22 |
| LMArena Creative Writing | 15% | #2 | 21/22 |
| OpenRouter usage | 10% | 87/100 | 22/22 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic2 models
- Meta1 model
- OpenAI1 model
- xAI1 model
- Z.ai1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 85 | 100% | 1 threads · 1 families · 0 cautions | #1 LMArena Text · #2 LMArena Creative Writing |
| 02 | Claude Opus 4.6Anthropic | 81 | 100% | 1 threads · 1 families · 0 cautions | #1 LMArena Creative Writing · #1 LMArena Instruction Following |
| 03 | GPT-5.6 SolOpenAI | 80 | 100% | no linked practitioner threads | #10 LMArena Text · #11 LMArena Instruction Following |
| 04 | Muse Spark 1.2Meta | 79 | 100% | no linked practitioner threads | #4 LMArena Text · #8 LiveBench Instruction Following |
| 05 | Grok 4.6xAI | 76 | 100% | 1 threads · 1 families · 1 cautions | #15 LiveBench Instruction Following · #18 LMArena Creative Writing |
| 06 | GLM 5.2Z.ai | 75 | 100% | 3 threads · 1 families · 3 cautions | #10 LMArena Creative Writing · #20 LMArena Text |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Fable 5 delivers strong instruction-following performance with a noticeably improved writing style that users find more capable and appealing for English composition tasks.
Best when: Use for business writing when you need reliable paraphrasing, simplification, and summarization, as it scores 75.77% on LiveBench Instruction Following including these exact subtasks.
Tips
- Use for business writing when you need reliable paraphrasing, simplification, and summarization, as it scores 75.77% on LiveBench Instruction Following including these exact subtasks.
- Deploy for polished English output where blind human raters ranked it #3 in instruction-following preference.
Watch out for
- Watch token consumption closely on long drafting tasks, as one user reported 60% faster session burn compared to earlier versions when redoing mathematical analysis work.
Claude Opus 4.6 holds the top human-preference rank for instruction-following, though cost estimates suggest significant expense at scale for email-heavy workflows.
Best when: Choose when maximum instruction-following quality matters, as it ranks #1 on LMArena's blind human preference votes in this category.
Tips
- Choose when maximum instruction-following quality matters, as it ranks #1 on LMArena's blind human preference votes in this category.
Watch out for
- Budget carefully for high-volume email processing, with estimated costs of $6.25 to $30.00 for 10,000 typical 200-word inputs with 50-word outputs.
GPT-5.6 Sol shows mid-tier instruction-following performance on both human preference and benchmark evaluations, with no distinctive advantages for business writing documented.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Consider alternatives for critical drafting, as its LiveBench Instruction Following score of 71.85% places it at #16 of 51, below several competitors in paraphrasing and summarization tasks.
Muse Spark 1.2 delivers benchmark-competitive instruction-following with strong LiveBench scores, though human preference rankings place it slightly below similarly-scored alternatives.
Best when: Deploy for story generation and summarization tasks where its 74.33% LiveBench Instruction Following score at #8 of 51 indicates solid capability.
Tips
- Deploy for story generation and summarization tasks where its 74.33% LiveBench Instruction Following score at #8 of 51 indicates solid capability.
Watch out for
- Verify output quality through human review, as its #17 LMArena instruction-following rank suggests blind raters preferred several lower-benchmark-scored alternatives.
Grok 4.6 exhibits serious reliability failures in business-critical drafting, with documented cases of 20 misattributed figures and unverified statistics in a single briefing document.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Avoid for any document requiring factual accuracy, as one evaluation found 20 misattributed figures and 1 unverified statistic in a single draft, with statistics displaying daggers that failed to prompt writer correction.
- Expect weaker instruction-following preference than alternatives, ranking #30 on LMArena with an Elo of 1446.
GLM 5.2 shows impressive benchmark scores and competitive proofreading performance on paper.
Best when: Consider for cost-efficient proofreading loops, as one benchmark found GLM 5.2 superior to Sonnet 4.6 on both quality and cost in multi-pass error correction.
Tips
- Consider for cost-efficient proofreading loops, as one benchmark found GLM 5.2 superior to Sonnet 4.6 on both quality and cost in multi-pass error correction.
Watch out for
- Plan for manual verification in article rewriting workflows, as users report GLM-5.2 introduced subtle mistakes requiring hand correction where Sonnet 4.6 succeeded in round two.
- Validate against top-tier options for complex implementation guides, as one extensive user found the practical gap versus Opus and Fable is huge despite raw benchmark impressiveness.
Frequently asked
- What is the top-ranked model for Business Writing?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for business writing when you need reliable paraphrasing, simplification, and summarization, as it scores 75.77% on LiveBench Instruction Following including these exact subtasks.[1]
- What should I watch out for with Anthropic: Claude Fable 5?
- Watch token consumption closely on long drafting tasks, as one user reported 60% faster session burn compared to earlier versions when redoing mathematical analysis work.[2]
- What is an alternative to Anthropic: Claude Fable 5?
- Anthropic: Claude Opus 4.6 is the next-ranked option. Choose when maximum instruction-following quality matters, as it ranks #1 on LMArena's blind human preference votes in this category.[3]
Sources
- 1
“Scores 75.77% on LiveBench Instruction Following (#5 of 51), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 2
“The writing style has significantly improved, however the token burn rate for tasks I have been working on seems to have skyrocketed. It definitely appears more capable (though I am unclear how much of that is just me liking the English it writes now vs actually more performant). I was using Fable 5 for some mathematical analysis assistance and redoing a part of it with 5.1 burned 60% of my session at a much faster rate.”
ddahlen · Hacker News · Sep 2, 2026 - 3
“Ranks #1 of 144 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 1, 2026 - 4
“Ranks #3 of 144 on LMArena's instruction-following category (Elo 1505), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 1, 2026 - 5
“Gemini says: "It would cost approximately $6.25 to $30.00 to have Claude Opus 4.6 respond to 10,000 emails, assuming a typical 200-word input and 50-word output per email."”
johndhi · Hacker News · Jun 26, 2026 - 6
“Scores 71.85% on LiveBench Instruction Following (#16 of 51), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 7
“Scores 74.33% on LiveBench Instruction Following (#8 of 51), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 8
“Ranks #17 of 144 on LMArena's instruction-following category (Elo 1469), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 1, 2026 - 9
“Park until a few more real drafts have been measured. Do not implement from this one run. One Grok 4.6 briefing (21 Aug 2026, psychosis in emergency departments, `drafts/2026-08-21-psychosis-ed-grok`) came back with **20 misattributed figures and 1 unverified**. Style cleaned 2→0. Statistics only print daggers, so the writer is never asked to remove a bad number. Two of those flags look like check bugs, not writer bugs: - `278 patients†` is in Montoya 2011’s **title**. The writer sees titles; `…”
bartholomewtj · GitHub · Aug 21, 2026 - 10
“Ranks #30 of 144 on LMArena's instruction-following category (Elo 1446), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 1, 2026 - 11
“I run a proofreading benchmark that tests how well models can find and fix errors in English text. They get several passes in a simple agent loop. Sonnet 5 is definitely better than Sonnet 4.6, but inferior on both quality and cost to GLM 5.1, GLM 5.2, Gemini 3.1 Flash, and Gemini 3.1 Pro. https: revise.io errata-bench”
artursapek · Hacker News · Jun 30, 2026 - 12
“I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models dail…”
sixtyj · Hacker News · Jun 30, 2026 - 13
“I don't think the writer has used top tier models very much. I have subscriptions to basically every provider, the difference between glm5.2 and opus is not even close, the gap is huge. raw benchmarks glm is impressive , but in practice these models are lacking so much. I had fable create a detailed implementation guide that explained how to implement everything in immense detail, it included all the libraries to use and versions. I then had deepseek v4 pro execute and it used old versions , di…”
AgentMasterRace · Hacker News · Jul 7, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.