Our methodology
Trust the process.
Then inspect it.
This field guide is automated and non-commercial. That makes the method, limits, and audit trail more important, not less.
end-to-end process
Six gates from raw signal to a published answer
- 01
Collect
Catalog facts, benchmarks, prices, and public practitioner reports.
- 02
Classify
Associate evidence with the exact model, task, sentiment, and source type.
- 03
Score
Apply the use case's fixed weights to current, eligible model data.
- 04
Draft claims
Propose atomic tips and cautions from the supplied evidence.
- 05
Verify
Audit each claim once and discard unsupported output.
- 06
Publish
Replace the complete recommendation atomically and rebuild the site.
Ranking, claim drafting, and verification are separate stages. The drafting model never receives authority to reorder candidates. A separate model audits atomic claims once; unsupported claims are discarded, and deterministic code assembles the published prose.
source observatory
What the pipeline can actually see
25 of 37 configured or intended sources are healthy. Disabled sources stay in task coverage as visible gaps.
| Source | Kind | Status | Rows matched | Last success |
|---|---|---|---|---|
| Aider Polyglot | benchmark | healthy | 41/69 | 2026-09-04 |
| Berkeley Function Calling | benchmark | healthy | 51/109 | 2026-09-04 |
| Configured evaluation gateway | benchmark | degraded | 0/0 | never |
| design arena | benchmark | healthy | 402/447 | 2026-09-04 |
| Frontier-Bench | benchmark | healthy | 13/18 | 2026-09-04 |
| livebench | benchmark | healthy | 350/371 | 2026-09-04 |
| LMArena Agent | benchmark | healthy | 52/56 | 2026-09-04 |
| LMArena WebDev | benchmark | healthy | 98/121 | 2026-09-04 |
| LMArena Creative Writing | benchmark | failing | 0/0 | 2026-09-03 |
| LMArena Document | benchmark | healthy | 34/38 | 2026-09-04 |
| LMArena Instruction Following | benchmark | healthy | 166/398 | 2026-09-04 |
| LMArena Long Query | benchmark | failing | 0/0 | 2026-09-03 |
| LMArena Math | benchmark | failing | 0/0 | 2026-09-03 |
| LMArena Search | benchmark | healthy | 6/34 | 2026-09-04 |
| LMArena Text | benchmark | healthy | 165/398 | 2026-09-04 |
| LMArena Vision | benchmark | failing | 0/0 | 2026-09-03 |
| OpenRouter usage | benchmark | healthy | 400/427 | 2026-09-04 |
| OSWorld Verified | benchmark | healthy | 38/40 | 2026-09-04 |
| SWE-bench Verified | benchmark | healthy | 16/49 | 2026-09-04 |
| SWE-rebench | benchmark | healthy | 93/117 | 2026-09-04 |
| Terminal-Bench 2.1 | benchmark | healthy | 13/18 | 2026-09-04 |
| Hacker News | community | healthy | 4271/4271 | 2026-09-04 |
| Hugging Face discussions | community | healthy | 89/89 | 2026-09-04 |
| CRAG | owned eval | disabled | 0/0 | never |
| FACTS Grounding | owned eval | disabled | 0/0 | never |
| LongBench v2 | owned eval | disabled | 0/0 | never |
| Route reliability | owned eval | disabled | 0/0 | never |
| Structured-output eval | owned eval | disabled | 0/0 | never |
| VLMEvalKit tasks | owned eval | disabled | 0/0 | never |
| WMT translation | owned eval | disabled | 0/0 | never |
| GitHub issues | practitioner | healthy | 501/501 | 2026-09-04 |
| RSS · https://aider.chat/feed.xml | practitioner | healthy | 0/10 | 2026-09-04 |
| RSS · https://huggingface.co/blog/feed.xml | practitioner | healthy | 0/859 | 2026-09-04 |
| RSS · https://metr.org/feed.xml | practitioner | healthy | 0/99 | 2026-09-04 |
| RSS · https://simonwillison.net/atom/everything/ | practitioner | healthy | 5/30 | 2026-09-04 |
| RSS · https://www.interconnects.ai/feed | practitioner | healthy | 0/20 | 2026-09-04 |
| Stack Exchange | practitioner | healthy | 1/1 | 2026-09-04 |
- 01
Catalog, not copy
OpenRouter and models.dev supply structured facts such as pricing, context, modalities, release dates, capabilities, and third-party benchmark identifiers. Newly cataloged models are joined automatically to benchmark rows by permanent slug. We never republish vendor descriptions.
- 02
Evidence has a type
Benchmarks, independent evaluators, practitioner reports, community discussions, operational telemetry, and official release evaluations are kept distinct. LiveBench, BFCL, SWE-rebench, Terminal-Bench, OSWorld, LMArena, Design Arena, Aider, SWE-bench, and configured task suites retain their evaluator and harness lineage. First-party release claims can never count as independent confirmation.
- 03
The question sets the weights
There is no universal best model. Each use case defines its own task-specific benchmark, price, context, adoption, reliability, and ownership signals. Popularity is capped as an adoption signal rather than treated as quality. Only the newest coherent benchmark snapshots participate.
- 04
Coverage has a named status
A model with one excellent result should not automatically equal one tested broadly. Coverage is measured against every intended feed, including a configured feed that is failing or not yet available. Established means at least 60% of the intended weighted benchmark mix across at least two independent evaluator families. One independent result or official release evidence is provisional.
- 05
The writer does not rank
The deterministic scorer owns candidate order. A language model turns supplied evidence into readable verdicts but cannot promote a model over the computed order.
- 06
Every source is checked
Citations are checked for model ownership and source existence, then a separate validation pass tests whether each excerpt supports the prose. One conversation counts once even if several comments were collected. Evaluators, repositories, publications, communities, and negative reports retain separate lineage. Unsupported output fails closed.
- 07
Freshness is explicit
Catalog, every benchmark source, official vendor release discovery, evidence, and synthesis each refresh daily. New model announcements are found through reviewed first-party sitemaps and exact model identity checks. Relevant evidence, material catalog changes, new benchmark snapshots, or the freshness ceiling trigger regeneration for affected use cases.
- 08
Constraints rerank a wider field
Editorial pages stay concise, but synthesis retains a larger pool of citation-audited candidates. Picker constraints such as open weights, price, context, modalities, and deployment route deterministically rerank that pool. A one-sided candidate can appear there only when its supported evidence is shown and the missing side is stated plainly.
- 09
One answer survives failures
Every selectable use case has one explicit top recommendation. If an upstream source or synthesis run fails, the last completely verified answer remains published and the pipeline raises an operational alert. Partial data never replaces it.
known limits
A useful map is still not the territory.
Community evidence leans developer-heavy, public discussion can be promotional or mistaken, benchmark coverage varies, and listed API prices do not capture hosting or latency. Official release evidence closes the launch-day information gap but remains visibly provisional until independent sources reproduce relevant results. Treat the top pick as the strongest model to test first, then validate it against your prompts, latency target, and failure budget.