Recommendation for Fast & cheap
Fast Summaries
Our top recommendation for Fast Summaries, based on the public evidence we track, is Google: Gemini 3.8 Flash.[1][2][3] Use for high-volume TL;DR pipelines where benchmark-topping instruction following (81.41%) translates to reliable summary formatting. Watch out: Expect slightly lower human preference ranking (#6 LMArena) than top-tier models if output polish matters beyond raw accuracy. Google: Gemini 3.7 Flash is the next-ranked alternative. Deploy when 3.8 Flash is unavailable, as it still ranks #2 on LiveBench Instruction Following (79.93%) with nearly identical latency profile.
About this recommendation
- Updated
- Sep 6, 2026
- Evidence through
- Sep 6, 2026
- Sources
- 6
- Revision
- v58
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
70%
intended feed weight
Largest provider share
3 of 6
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| price weight | 30% | 100/100 | 20/20 |
| LiveBench Instruction Following | 25% | #1 | 18/20 |
| Route reliabilityunavailable | 25% | feed unavailable | 0/20 |
| LMArena Text | 10% | #6 | 19/20 |
| OpenRouter usage | 10% | 94/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Google3 models
- Anthropic2 models
- OpenAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Gemini 3.8 FlashGoogle | 76 | 70% | no linked practitioner threads | #1 LiveBench Instruction Following · #6 LMArena Text |
| 02 | Gemini 3.7 FlashGoogle | 74 | 70% | no linked practitioner threads | #2 LiveBench Instruction Following · #9 LMArena Text |
| 03 | Gemini 3.6 FlashGoogle | 70 | 70% | no linked practitioner threads | #8 LiveBench Instruction Following · #15 LMArena Text |
| 04 | GPT-5.6 SolOpenAI | 70 | 70% | 1 threads · 1 families · 0 cautions | #12 LMArena Text · #17 LiveBench Instruction Following |
| 05 | Claude Fable 5Anthropic | 69 | 70% | no linked practitioner threads | #1 LMArena Text · #5 LiveBench Instruction Following |
| 06 | Claude Opus 4.6Anthropic | 65 | 70% | no linked practitioner threads | #2 LMArena Text · #35 LiveBench Instruction Following |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Gemini 3.8 Flash leads the Flash lineup with the highest LiveBench Instruction Following score among fast summarization candidates, ranking #1 on that benchmark.
Best when: Use for high-volume TL;DR pipelines where benchmark-topping instruction following (81.41%) translates to reliable summary formatting.
Tips
- Use for high-volume TL;DR pipelines where benchmark-topping instruction following (81.41%) translates to reliable summary formatting.
Watch out for
- Expect slightly lower human preference ranking (#6 LMArena) than top-tier models if output polish matters beyond raw accuracy.
Gemini 3.7 Flash sits just behind its successor on LiveBench Instruction Following while maintaining a top-10 LMArena position.
Best when: Deploy when 3.8 Flash is unavailable, as it still ranks #2 on LiveBench Instruction Following (79.93%) with nearly identical latency profile.
Tips
- Deploy when 3.8 Flash is unavailable, as it still ranks #2 on LiveBench Instruction Following (79.93%) with nearly identical latency profile.
Gemini 3.6 Flash offers the lowest cost entry in Google's Flash series with a still-respectable #8 ranking on LiveBench Instruction Following.
Best when: Choose for the cheapest Google-hosted summarization when 75.37% LiveBench Instruction Following suffices for your TL;DR quality bar.
Tips
- Choose for the cheapest Google-hosted summarization when 75.37% LiveBench Instruction Following suffices for your TL;DR quality bar.
GPT-5.6 Sol shows middling LiveBench results (#17) with community reports favoring older 5.4 models for speed, raising questions about its cost-latency tradeoff.
Best when: Test against 5.4-mini if you need OpenAI ecosystem access, as one user reports 5.4 remains faster with comparable results for quick summaries.
Tips
- Test against 5.4-mini if you need OpenAI ecosystem access, as one user reports 5.4 remains faster with comparable results for quick summaries.
Watch out for
- Question the upgrade path from 5.4/5.4-mini, since community evidence suggests 5.6 may not deliver the speed gains expected for high-volume TL;DRs.
Claude Fable 5 tops LMArena overall preferences yet lands fifth on LiveBench Instruction Following, suggesting stronger general chat than tight summarization.
Best when: Select when human preference rankings predict user satisfaction better than instruction-following benchmarks for your summary consumers.
Tips
- Select when human preference rankings predict user satisfaction better than instruction-following benchmarks for your summary consumers.
Anthropic: Claude Opus 4.6 ranks #2 of 144 on LMArena's overall text arena (Elo 1505), based on blind human preference votes.
Best when: Consider only after reviewing the cited caution.
Frequently asked
- What is the top-ranked model for Fast Summaries?
- Google: Gemini 3.8 Flash ranks first in the current evidence-weighted comparison. Use for high-volume TL;DR pipelines where benchmark-topping instruction following (81.41%) translates to reliable summary formatting.[1]
- What should I watch out for with Google: Gemini 3.8 Flash?
- Expect slightly lower human preference ranking (#6 LMArena) than top-tier models if output polish matters beyond raw accuracy.[2]
- What is an alternative to Google: Gemini 3.8 Flash?
- Google: Gemini 3.7 Flash is the next-ranked option. Deploy when 3.8 Flash is unavailable, as it still ranks #2 on LiveBench Instruction Following (79.93%) with nearly identical latency profile.[3]
Sources
- 1
“Scores 81.41% on LiveBench Instruction Following (#1 of 52), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 2
“Ranks #6 of 144 on LMArena's overall text arena (Elo 1494), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 2, 2026 - 3
“Scores 79.93% on LiveBench Instruction Following (#2 of 52), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 4
“Scores 75.37% on LiveBench Instruction Following (#8 of 52), including paraphrasing, simplifying, story generation, and summarization.”
LiveBench Instruction Following · Benchmark · Jun 25, 2026 - 5
“I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?”
thomas_witt · Hacker News · Jul 10, 2026 - 6
“Ranks #1 of 144 on LMArena's overall text arena (Elo 1507), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 2, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.