Recommendation for Fast & cheap

Fast Summaries

Our top recommendation for Fast Summaries, based on the public evidence we track, is Google: Gemini 3.8 Flash.[1][2][3] Use for high-volume TL;DR pipelines where benchmark-topping instruction following (81.41%) translates to reliable summary formatting. Watch out: Expect slightly lower human preference ranking (#6 LMArena) than top-tier models if output polish matters beyond raw accuracy. Google: Gemini 3.7 Flash is the next-ranked alternative. Deploy when 3.8 Flash is unavailable, as it still ranks #2 on LiveBench Instruction Following (79.93%) with nearly identical latency profile.

About this recommendation

Updated
Sep 6, 2026
Evidence through
Sep 6, 2026
Sources
6
Revision
v58

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

20

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

70%

intended feed weight

Largest provider share

3 of 6

Google

Provisional source breadth. 3 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 45%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Gemini 3.8 Flash
Evaluation feedWeightWinner resultField measured
price weight
30%
100/10020/20
LiveBench Instruction Following
25%
#118/20
Route reliabilityunavailable
25%
feed unavailable0/20
LMArena Text
10%
#619/20
OpenRouter usage
10%
94/10020/20

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Google50%
  • Google3 models
  • Anthropic2 models
  • OpenAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Gemini 3.8 FlashGoogle
76
70%no linked practitioner threads#1 LiveBench Instruction Following · #6 LMArena Text
02Gemini 3.7 FlashGoogle
74
70%no linked practitioner threads#2 LiveBench Instruction Following · #9 LMArena Text
03Gemini 3.6 FlashGoogle
70
70%no linked practitioner threads#8 LiveBench Instruction Following · #15 LMArena Text
04GPT-5.6 SolOpenAI
70
70%1 threads · 1 families · 0 cautions#12 LMArena Text · #17 LiveBench Instruction Following
05Claude Fable 5Anthropic
69
70%no linked practitioner threads#1 LMArena Text · #5 LiveBench Instruction Following
06Claude Opus 4.6Anthropic
65
70%no linked practitioner threads#2 LMArena Text · #35 LiveBench Instruction Following

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Gemini 3.8 Flash leads the Flash lineup with the highest LiveBench Instruction Following score among fast summarization candidates, ranking #1 on that benchmark.

    Best when: Use for high-volume TL;DR pipelines where benchmark-topping instruction following (81.41%) translates to reliable summary formatting.

    Tips

    • Use for high-volume TL;DR pipelines where benchmark-topping instruction following (81.41%) translates to reliable summary formatting.
      Source 1
      Scores 81.41% on LiveBench Instruction Following (#1 of 52), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗

    Watch out for

    • Expect slightly lower human preference ranking (#6 LMArena) than top-tier models if output polish matters beyond raw accuracy.
      Source 2
      Ranks #6 of 144 on LMArena's overall text arena (Elo 1494), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  2. Gemini 3.7 Flash sits just behind its successor on LiveBench Instruction Following while maintaining a top-10 LMArena position.

    Best when: Deploy when 3.8 Flash is unavailable, as it still ranks #2 on LiveBench Instruction Following (79.93%) with nearly identical latency profile.

    Tips

    • Deploy when 3.8 Flash is unavailable, as it still ranks #2 on LiveBench Instruction Following (79.93%) with nearly identical latency profile.
      Source 3
      Scores 79.93% on LiveBench Instruction Following (#2 of 52), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗
  3. Gemini 3.6 Flash offers the lowest cost entry in Google's Flash series with a still-respectable #8 ranking on LiveBench Instruction Following.

    Best when: Choose for the cheapest Google-hosted summarization when 75.37% LiveBench Instruction Following suffices for your TL;DR quality bar.

    Tips

    • Choose for the cheapest Google-hosted summarization when 75.37% LiveBench Instruction Following suffices for your TL;DR quality bar.
      Source 4
      Scores 75.37% on LiveBench Instruction Following (#8 of 52), including paraphrasing, simplifying, story generation, and summarization.
      LiveBench Instruction FollowingOpen original ↗
  4. GPT-5.6 Sol shows middling LiveBench results (#17) with community reports favoring older 5.4 models for speed, raising questions about its cost-latency tradeoff.

    Best when: Test against 5.4-mini if you need OpenAI ecosystem access, as one user reports 5.4 remains faster with comparable results for quick summaries.

    Tips

    • Test against 5.4-mini if you need OpenAI ecosystem access, as one user reports 5.4 remains faster with comparable results for quick summaries.
      Source 5
      I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?

    Watch out for

    • Question the upgrade path from 5.4/5.4-mini, since community evidence suggests 5.6 may not deliver the speed gains expected for high-volume TL;DRs.
      Source 5
      I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?
  5. Claude Fable 5 tops LMArena overall preferences yet lands fifth on LiveBench Instruction Following, suggesting stronger general chat than tight summarization.

    Best when: Select when human preference rankings predict user satisfaction better than instruction-following benchmarks for your summary consumers.

    Tips

    • Select when human preference rankings predict user satisfaction better than instruction-following benchmarks for your summary consumers.
      Source 6
      Ranks #1 of 144 on LMArena's overall text arena (Elo 1507), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  6. Anthropic: Claude Opus 4.6 ranks #2 of 144 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.

    Best when: Consider only after reviewing the cited caution.

Frequently asked

What is the top-ranked model for Fast Summaries?
Google: Gemini 3.8 Flash ranks first in the current evidence-weighted comparison. Use for high-volume TL;DR pipelines where benchmark-topping instruction following (81.41%) translates to reliable summary formatting.[1]
What should I watch out for with Google: Gemini 3.8 Flash?
Expect slightly lower human preference ranking (#6 LMArena) than top-tier models if output polish matters beyond raw accuracy.[2]
What is an alternative to Google: Gemini 3.8 Flash?
Google: Gemini 3.7 Flash is the next-ranked option. Deploy when 3.8 Flash is unavailable, as it still ranks #2 on LiveBench Instruction Following (79.93%) with nearly identical latency profile.[3]

Sources

  1. 1

    Scores 81.41% on LiveBench Instruction Following (#1 of 52), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  2. 2

    Ranks #6 of 144 on LMArena's overall text arena (Elo 1494), based on blind human preference votes.

    LMArena text arena · Benchmark · Sep 2, 2026
  3. 3

    Scores 79.93% on LiveBench Instruction Following (#2 of 52), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  4. 4

    Scores 75.37% on LiveBench Instruction Following (#8 of 52), including paraphrasing, simplifying, story generation, and summarization.

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  5. 5

    I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?

    thomas_witt · Hacker News · Jul 10, 2026
  6. 6

    Ranks #1 of 144 on LMArena's overall text arena (Elo 1507), based on blind human preference votes.

    LMArena text arena · Benchmark · Sep 2, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.