Recommendation for Translation

Translation

Our top recommendation for Translation, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use when you need translations that human evaluators prefer over 143 other models, as it ranks #1 on LMArena's blind text arena. Google: Gemini 3.7 Flash is the next-ranked alternative. Consider for translation workflows where LiveBench Language scores above 85% are acceptable, as it hits 85.46% at rank #8.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
12
Revision
v59

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

44%

intended feed weight

Largest provider share

2 of 8

Anthropic

Provisional source breadth. 3 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 53%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
WMT translationunavailable
45%
feed unavailable0/21
LiveBench Language
20%
#118/21
LMArena Text
15%
#119/21
LiveBench Instruction Following
10%
#518/21
OpenRouter usage
10%
87/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic25%
  • Anthropic2 models
  • Google2 models
  • Moonshot AI1 model
  • OpenAI1 model
  • xAI1 model
  • Z.ai1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
59
44%no linked practitioner threads#1 LiveBench Language · #1 LMArena Text
02Gemini 3.7 FlashGoogle
58
44%no linked practitioner threads#2 LiveBench Instruction Following · #7 LiveBench Language
03GPT-5.6 SolOpenAI
57
44%no linked practitioner threads#5 LiveBench Language · #10 LMArena Text
04Kimi K3Moonshot AI
56
44%1 threads · 1 families · 0 cautions#6 LiveBench Language · #8 LMArena Text
05GLM 5.3Z.ai
53
44%no linked practitioner threads#10 LMArena Text · #19 LiveBench Language
06Grok 4.6xAI
53
44%no linked practitioner threads#11 LiveBench Language · #15 LiveBench Instruction Following
07Claude Opus 4.6Anthropic
53
44%no linked practitioner threads#2 LMArena Text · #12 LiveBench Language
08Gemini 3.8 FlashGoogle
51
25%no linked practitioner threads#1 LiveBench Instruction Following · #4 LiveBench Language

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Leads on both human preference and language manipulation benchmarks, suggesting strong translation fluency and natural output quality.

    Best when: Use when you need translations that human evaluators prefer over 143 other models, as it ranks #1 on LMArena's blind text arena.

    Tips

    • Use when you need translations that human evaluators prefer over 143 other models, as it ranks #1 on LMArena's blind text arena.
      Source 1
      Ranks #1 of 144 on LMArena's overall text arena (Elo 1507), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • Deploy for complex language manipulation tasks where it scores 90.68% on LiveBench Language, the highest of any candidate.
      Source 3
      Scores 90.68% on LiveBench Language (#1 of 51), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗
  2. Strong mid-tier performer with solid language manipulation scores and competitive human preference rankings.

    Best when: Consider for translation workflows where LiveBench Language scores above 85% are acceptable, as it hits 85.46% at rank #8.

    Tips

    • Consider for translation workflows where LiveBench Language scores above 85% are acceptable, as it hits 85.46% at rank #8.
      Source 2
      Scores 85.46% on LiveBench Language (#8 of 51), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗
  3. Balances strong language manipulation performance with moderate human preference placement.

    Best when: Use for translation tasks requiring high LiveBench Language scores, where its 87.68% ranks #5 among all candidates.

    Tips

    • Use for translation tasks requiring high LiveBench Language scores, where its 87.68% ranks #5 among all candidates.
      Source 4
      Scores 87.68% on LiveBench Language (#5 of 51), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗
  4. Open-weight option with competitive benchmarks and explicit Mandarin translation focus, plus cost advantages.

    Best when: Choose for Chinese-English translation workflows, as the model's benchmarks are published in Mandarin with translation resources available.

    Tips

    • Choose for Chinese-English translation workflows, as the model's benchmarks are published in Mandarin with translation resources available.
      Source 5
      Full benchmarks in Mandarin: https: mp.weixin.qq.com s V4xhEIy8xDXSMDPrPkmUAQ Translation: https: mp-weixin-qq-com.translate.goog s V4xhEIy8xDXSMDPrPk... Cheaper then GPT 5.6 Sol (according to their results) ...
    • Deploy when cost matters, as evidence indicates it runs cheaper than GPT 5.6 Sol.
      Source 5
      Full benchmarks in Mandarin: https: mp.weixin.qq.com s V4xhEIy8xDXSMDPrPkmUAQ Translation: https: mp-weixin-qq-com.translate.goog s V4xhEIy8xDXSMDPrPk... Cheaper then GPT 5.6 Sol (according to their results) ...
    • Use for general translation where 85.53% on LiveBench Language provides sufficient quality.
      Source 6
      Scores 85.53% on LiveBench Language (#7 of 51), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗
  5. Open-weight model with modest language manipulation scores, filling coverage requirements for open-weight deployment routes.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Expect lower translation quality than top performers, scoring 79.86% on LiveBench Language at rank #21 of 51.
      Source 7
      Scores 79.86% on LiveBench Language (#21 of 51), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗
  6. Mid-tier closed-weight option with moderate benchmark scores and weaker human preference showing.

    Best when: Consider for translation workflows where LiveBench Language scores around 83.7% are sufficient, ranking #12 of 51.

    Tips

    • Consider for translation workflows where LiveBench Language scores around 83.7% are sufficient, ranking #12 of 51.
      Source 8
      Scores 83.7% on LiveBench Language (#12 of 51), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗

    Watch out for

    • Note significant human preference gap, ranking #29 on LMArena with Elo 1461, far below top-tier models.
      Source 9
      Ranks #29 of 144 on LMArena's overall text arena (Elo 1461), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  7. Near-top human preference performer with surprisingly modest language manipulation benchmark scores.

    Best when: Use when human evaluators' blind preferences matter most, as it ranks #2 on LMArena with Elo 1505.

    Tips

    • Use when human evaluators' blind preferences matter most, as it ranks #2 on LMArena with Elo 1505.
      Source 10
      Ranks #2 of 144 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Verify suitability for your specific translation task, as its LiveBench Language score of 83.27% ranks only #13, below several cheaper alternatives.
      Source 11
      Scores 83.27% on LiveBench Language (#13 of 51), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗
  8. Strong language manipulation performer with missing human preference data, suggesting possible new release status.

    Best when: Use for translation tasks where LiveBench Language scores near 88% are needed, as it ranks #4 with 87.79%.

    Tips

    • Use for translation tasks where LiveBench Language scores near 88% are needed, as it ranks #4 with 87.79%.
      Source 12
      Scores 87.79% on LiveBench Language (#4 of 51), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗

Frequently asked

What is the top-ranked model for Translation?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use when you need translations that human evaluators prefer over 143 other models, as it ranks #1 on LMArena's blind text arena.[1]
What is an alternative to Anthropic: Claude Fable 5?
Google: Gemini 3.7 Flash is the next-ranked option. Consider for translation workflows where LiveBench Language scores above 85% are acceptable, as it hits 85.46% at rank #8.[2]

Sources

  1. 1

    Ranks #1 of 144 on LMArena's overall text arena (Elo 1507), based on blind human preference votes.

    LMArena text arena · Benchmark · Sep 1, 2026
  2. 2

    Scores 85.46% on LiveBench Language (#8 of 51), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026
  3. 3

    Scores 90.68% on LiveBench Language (#1 of 51), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026
  4. 4

    Scores 87.68% on LiveBench Language (#5 of 51), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026
  5. 5

    Full benchmarks in Mandarin: https: mp.weixin.qq.com s V4xhEIy8xDXSMDPrPkmUAQ Translation: https: mp-weixin-qq-com.translate.goog s V4xhEIy8xDXSMDPrPk... Cheaper then GPT 5.6 Sol (according to their results) ...

    benjiro29 · Hacker News · Jul 16, 2026
  6. 6

    Scores 85.53% on LiveBench Language (#7 of 51), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026
  7. 7

    Scores 79.86% on LiveBench Language (#21 of 51), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026
  8. 8

    Scores 83.7% on LiveBench Language (#12 of 51), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026
  9. 9

    Ranks #29 of 144 on LMArena's overall text arena (Elo 1461), based on blind human preference votes.

    LMArena text arena · Benchmark · Sep 1, 2026
  10. 10

    Ranks #2 of 144 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.

    LMArena text arena · Benchmark · Sep 1, 2026
  11. 11

    Scores 83.27% on LiveBench Language (#13 of 51), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026
  12. 12

    Scores 87.79% on LiveBench Language (#4 of 51), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.