Recommendation for Translation
Translation
Our top recommendation for Translation, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use when you need translations that human evaluators prefer over 143 other models, as it ranks #1 on LMArena's blind text arena. Google: Gemini 3.7 Flash is the next-ranked alternative. Consider for translation workflows where LiveBench Language scores above 85% are acceptable, as it hits 85.46% at rank #8.
About this recommendation
- Updated
- Sep 4, 2026
- Evidence through
- Sep 4, 2026
- Sources
- 12
- Revision
- v59
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
44%
intended feed weight
Largest provider share
2 of 8
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| WMT translationunavailable | 45% | feed unavailable | 0/21 |
| LiveBench Language | 20% | #1 | 18/21 |
| LMArena Text | 15% | #1 | 19/21 |
| LiveBench Instruction Following | 10% | #5 | 18/21 |
| OpenRouter usage | 10% | 87/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic2 models
- Google2 models
- Moonshot AI1 model
- OpenAI1 model
- xAI1 model
- Z.ai1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 59 | 44% | no linked practitioner threads | #1 LiveBench Language · #1 LMArena Text |
| 02 | Gemini 3.7 FlashGoogle | 58 | 44% | no linked practitioner threads | #2 LiveBench Instruction Following · #7 LiveBench Language |
| 03 | GPT-5.6 SolOpenAI | 57 | 44% | no linked practitioner threads | #5 LiveBench Language · #10 LMArena Text |
| 04 | Kimi K3Moonshot AI | 56 | 44% | 1 threads · 1 families · 0 cautions | #6 LiveBench Language · #8 LMArena Text |
| 05 | GLM 5.3Z.ai | 53 | 44% | no linked practitioner threads | #10 LMArena Text · #19 LiveBench Language |
| 06 | Grok 4.6xAI | 53 | 44% | no linked practitioner threads | #11 LiveBench Language · #15 LiveBench Instruction Following |
| 07 | Claude Opus 4.6Anthropic | 53 | 44% | no linked practitioner threads | #2 LMArena Text · #12 LiveBench Language |
| 08 | Gemini 3.8 FlashGoogle | 51 | 25% | no linked practitioner threads | #1 LiveBench Instruction Following · #4 LiveBench Language |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Leads on both human preference and language manipulation benchmarks, suggesting strong translation fluency and natural output quality.
Best when: Use when you need translations that human evaluators prefer over 143 other models, as it ranks #1 on LMArena's blind text arena.
Tips
- Use when you need translations that human evaluators prefer over 143 other models, as it ranks #1 on LMArena's blind text arena.
- Deploy for complex language manipulation tasks where it scores 90.68% on LiveBench Language, the highest of any candidate.
Strong mid-tier performer with solid language manipulation scores and competitive human preference rankings.
Best when: Consider for translation workflows where LiveBench Language scores above 85% are acceptable, as it hits 85.46% at rank #8.
Tips
- Consider for translation workflows where LiveBench Language scores above 85% are acceptable, as it hits 85.46% at rank #8.
Balances strong language manipulation performance with moderate human preference placement.
Best when: Use for translation tasks requiring high LiveBench Language scores, where its 87.68% ranks #5 among all candidates.
Tips
- Use for translation tasks requiring high LiveBench Language scores, where its 87.68% ranks #5 among all candidates.
Open-weight option with competitive benchmarks and explicit Mandarin translation focus, plus cost advantages.
Best when: Choose for Chinese-English translation workflows, as the model's benchmarks are published in Mandarin with translation resources available.
Tips
- Choose for Chinese-English translation workflows, as the model's benchmarks are published in Mandarin with translation resources available.
- Deploy when cost matters, as evidence indicates it runs cheaper than GPT 5.6 Sol.
- Use for general translation where 85.53% on LiveBench Language provides sufficient quality.
Open-weight model with modest language manipulation scores, filling coverage requirements for open-weight deployment routes.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Expect lower translation quality than top performers, scoring 79.86% on LiveBench Language at rank #21 of 51.
Mid-tier closed-weight option with moderate benchmark scores and weaker human preference showing.
Best when: Consider for translation workflows where LiveBench Language scores around 83.7% are sufficient, ranking #12 of 51.
Tips
- Consider for translation workflows where LiveBench Language scores around 83.7% are sufficient, ranking #12 of 51.
Watch out for
- Note significant human preference gap, ranking #29 on LMArena with Elo 1461, far below top-tier models.
Near-top human preference performer with surprisingly modest language manipulation benchmark scores.
Best when: Use when human evaluators' blind preferences matter most, as it ranks #2 on LMArena with Elo 1505.
Tips
- Use when human evaluators' blind preferences matter most, as it ranks #2 on LMArena with Elo 1505.
Watch out for
- Verify suitability for your specific translation task, as its LiveBench Language score of 83.27% ranks only #13, below several cheaper alternatives.
Strong language manipulation performer with missing human preference data, suggesting possible new release status.
Best when: Use for translation tasks where LiveBench Language scores near 88% are needed, as it ranks #4 with 87.79%.
Tips
- Use for translation tasks where LiveBench Language scores near 88% are needed, as it ranks #4 with 87.79%.
Frequently asked
- What is the top-ranked model for Translation?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use when you need translations that human evaluators prefer over 143 other models, as it ranks #1 on LMArena's blind text arena.[1]
- What is an alternative to Anthropic: Claude Fable 5?
- Google: Gemini 3.7 Flash is the next-ranked option. Consider for translation workflows where LiveBench Language scores above 85% are acceptable, as it hits 85.46% at rank #8.[2]
Sources
- 1
“Ranks #1 of 144 on LMArena's overall text arena (Elo 1507), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 1, 2026 - 2
“Scores 85.46% on LiveBench Language (#8 of 51), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 3
“Scores 90.68% on LiveBench Language (#1 of 51), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 4
“Scores 87.68% on LiveBench Language (#5 of 51), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 5
“Full benchmarks in Mandarin: https: mp.weixin.qq.com s V4xhEIy8xDXSMDPrPkmUAQ Translation: https: mp-weixin-qq-com.translate.goog s V4xhEIy8xDXSMDPrPk... Cheaper then GPT 5.6 Sol (according to their results) ...”
benjiro29 · Hacker News · Jul 16, 2026 - 6
“Scores 85.53% on LiveBench Language (#7 of 51), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 7
“Scores 79.86% on LiveBench Language (#21 of 51), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 8
“Scores 83.7% on LiveBench Language (#12 of 51), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 9
“Ranks #29 of 144 on LMArena's overall text arena (Elo 1461), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 1, 2026 - 10
“Ranks #2 of 144 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 1, 2026 - 11
“Scores 83.27% on LiveBench Language (#13 of 51), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026 - 12
“Scores 87.79% on LiveBench Language (#4 of 51), an objective evaluation of language manipulation tasks.”
LiveBench Language · Benchmark · Jun 25, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.