Recommendation for Literary & nuanced

Literary Translation

Our top recommendation for Literary Translation, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use when translating classical literature with meter or rhyme, as it can identify and continue existing translation styles from minimal context. DeepSeek: DeepSeek V4 Flash 0423 is the next-ranked alternative. Its currently supported evidence is cautionary: Watch for whitespace and markdown corruption in streaming responses, which could damage formatted literary text.

About this recommendation

Updated
Sep 6, 2026
Evidence through
Sep 6, 2026
Sources
5
Revision
v57

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

44%

intended feed weight

Largest provider share

3 of 5

Anthropic

Provisional source breadth. 4 citation families and 1 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 38%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
WMT translationunavailable
45%
feed unavailable0/21
LiveBench Language
20%
#119/21
LMArena Text
15%
#120/21
LiveBench Instruction Following
10%
#519/21
OpenRouter usage
10%
86/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic60%
  • Anthropic3 models
  • deepseek1 model
  • OpenAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
59
44%1 threads · 1 families · 0 cautions#1 LiveBench Language · #1 LMArena Text
02DeepSeek V4 Flash 0423deepseek
53
44%1 threads · 1 families · 0 cautions#37 LiveBench Instruction Following · #46 LiveBench Language
03Claude Opus 4.6Anthropic
53
44%no linked practitioner threads#2 LMArena Text · #13 LiveBench Language
04Claude Opus 4.8Anthropic
52
44%no linked practitioner threads#13 LMArena Text · #15 LiveBench Instruction Following
05GPT-6 AstraOpenAI
48
25%no linked practitioner threads#3 LiveBench Language · #7 LiveBench Instruction Following

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Leads on both human preference and language manipulation benchmarks, with demonstrated ability to continue classical English verse translations when given source text anchors.

    Best when: Use when translating classical literature with meter or rhyme, as it can identify and continue existing translation styles from minimal context.

    Tips

    • Use when translating classical literature with meter or rhyme, as it can identify and continue existing translation styles from minimal context.
      Source 1
      > The only thing Fable was given was a clean of copy of the ancient Greek. This is not a bastardization of other translations FWIW I gave GPT-5.5 two lines of William Cowper's translation (long out of copyright to avoid any filters) without web search and not only was it able to identify it, it was able to continue it verbatim for roughly 30 lines at which point it went off course, but only by skipping like two books ahead. If I had supplied the original text as an anchor it presumably could ha…
    • Deploy for high-stakes literary work where human judges consistently prefer its output over 143 competing models.
      Source 3
      Ranks #1 of 144 on LMArena's overall text arena (Elo 1507), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  2. The only open-weight candidate with evidence, though the sole documentation describes streaming infrastructure bugs rather than translation quality.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Watch for whitespace and markdown corruption in streaming responses, which could damage formatted literary text.
      Source 2
      ### OmniRoute Version 3.8.49 (docker image `diegosouzapw/omniroute:latest`) ### Installation Method docker ### Operating System Linux (Proxmox LXC) ### Provider(s) Involved opencode-go, opencode-zen, openrouter (any streaming provider) ### Model(s) Involved deepseek-v4-flash, ox-alpha-free (multiple models — not model-specific) ### Client Tool Claude Code CLI (Anthropic `/v1/messages` streaming) ### Description Streaming responses through OmniRoute get their whitespace/markdown structure corrup…
  3. Anthropic: Claude Opus 4.6 ranks #2 of 144 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.

    Best when: Consider only after reviewing the cited caution.

  4. Falls to 26th place on language manipulation despite moderate human preference rankings, suggesting a disconnect between perceived and measured quality.

    Best when: Consider when you need Anthropic's policy framework but cannot access Fable 5 or Opus 4.6, as it still places 13th in human votes.

    Tips

    • Consider when you need Anthropic's policy framework but cannot access Fable 5 or Opus 4.6, as it still places 13th in human votes.
      Source 4
      Ranks #13 of 144 on LMArena's overall text arena (Elo 1482), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  5. Achieves third-place language manipulation scores without any human preference ranking data, making it a benchmark-strong but unvalidated option for literary work.

    Best when: Consider for experimental translation pipelines where objective language task performance is the primary selection criterion.

    Tips

    • Consider for experimental translation pipelines where objective language task performance is the primary selection criterion.
      Source 5
      Scores 89.43% on LiveBench Language (#3 of 52), an objective evaluation of language manipulation tasks.
      LiveBench LanguageOpen original ↗

Frequently asked

What is the top-ranked model for Literary Translation?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use when translating classical literature with meter or rhyme, as it can identify and continue existing translation styles from minimal context.[1]

Sources

  1. 1

    > The only thing Fable was given was a clean of copy of the ancient Greek. This is not a bastardization of other translations FWIW I gave GPT-5.5 two lines of William Cowper's translation (long out of copyright to avoid any filters) without web search and not only was it able to identify it, it was able to continue it verbatim for roughly 30 lines at which point it went off course, but only by skipping like two books ahead. If I had supplied the original text as an anchor it presumably could ha…

    magicalist · Hacker News · Aug 7, 2026
  2. 2

    ### OmniRoute Version 3.8.49 (docker image `diegosouzapw/omniroute:latest`) ### Installation Method docker ### Operating System Linux (Proxmox LXC) ### Provider(s) Involved opencode-go, opencode-zen, openrouter (any streaming provider) ### Model(s) Involved deepseek-v4-flash, ox-alpha-free (multiple models — not model-specific) ### Client Tool Claude Code CLI (Anthropic `/v1/messages` streaming) ### Description Streaming responses through OmniRoute get their whitespace/markdown structure corrup…

    vinnyduke · GitHub · Aug 26, 2026
  3. 3

    Ranks #1 of 144 on LMArena's overall text arena (Elo 1507), based on blind human preference votes.

    LMArena text arena · Benchmark · Sep 2, 2026
  4. 4

    Ranks #13 of 144 on LMArena's overall text arena (Elo 1482), based on blind human preference votes.

    LMArena text arena · Benchmark · Sep 2, 2026
  5. 5

    Scores 89.43% on LiveBench Language (#3 of 52), an objective evaluation of language manipulation tasks.

    LiveBench Language · Benchmark · Jun 25, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.