Our methodology

Trust the process.
Then inspect it.

This field guide is automated and non-commercial. That makes the method, limits, and audit trail more important, not less.

end-to-end process

Six gates from raw signal to a published answer

  1. 01

    Collect

    Catalog facts, benchmarks, prices, and public practitioner reports.

  2. 02

    Classify

    Associate evidence with the exact model, task, sentiment, and source type.

  3. 03

    Score

    Apply the use case's fixed weights to current, eligible model data.

  4. 04

    Draft claims

    Propose atomic tips and cautions from the supplied evidence.

  5. 05

    Verify

    Audit each claim once and discard unsupported output.

  6. 06

    Publish

    Replace the complete recommendation atomically and rebuild the site.

Ranking, claim drafting, and verification are separate stages. The drafting model never receives authority to reorder candidates. A separate model audits atomic claims once; unsupported claims are discarded, and deterministic code assembles the published prose.

source observatory

What the pipeline can actually see

25 of 37 configured or intended sources are healthy. Disabled sources stay in task coverage as visible gaps.

SourceKindStatusRows matchedLast success
Aider Polyglotbenchmarkhealthy41/692026-09-04
Berkeley Function Callingbenchmarkhealthy51/1092026-09-04
Configured evaluation gatewaybenchmarkdegraded0/0never
design arenabenchmarkhealthy402/4472026-09-04
Frontier-Benchbenchmarkhealthy13/182026-09-04
livebenchbenchmarkhealthy350/3712026-09-04
LMArena Agentbenchmarkhealthy52/562026-09-04
LMArena WebDevbenchmarkhealthy98/1212026-09-04
LMArena Creative Writingbenchmarkfailing0/02026-09-03
LMArena Documentbenchmarkhealthy34/382026-09-04
LMArena Instruction Followingbenchmarkhealthy166/3982026-09-04
LMArena Long Querybenchmarkfailing0/02026-09-03
LMArena Mathbenchmarkfailing0/02026-09-03
LMArena Searchbenchmarkhealthy6/342026-09-04
LMArena Textbenchmarkhealthy165/3982026-09-04
LMArena Visionbenchmarkfailing0/02026-09-03
OpenRouter usagebenchmarkhealthy400/4272026-09-04
OSWorld Verifiedbenchmarkhealthy38/402026-09-04
SWE-bench Verifiedbenchmarkhealthy16/492026-09-04
SWE-rebenchbenchmarkhealthy93/1172026-09-04
Terminal-Bench 2.1benchmarkhealthy13/182026-09-04
Hacker Newscommunityhealthy4271/42712026-09-04
Hugging Face discussionscommunityhealthy89/892026-09-04
CRAGowned evaldisabled0/0never
FACTS Groundingowned evaldisabled0/0never
LongBench v2owned evaldisabled0/0never
Route reliabilityowned evaldisabled0/0never
Structured-output evalowned evaldisabled0/0never
VLMEvalKit tasksowned evaldisabled0/0never
WMT translationowned evaldisabled0/0never
GitHub issuespractitionerhealthy501/5012026-09-04
RSS · https://aider.chat/feed.xmlpractitionerhealthy0/102026-09-04
RSS · https://huggingface.co/blog/feed.xmlpractitionerhealthy0/8592026-09-04
RSS · https://metr.org/feed.xmlpractitionerhealthy0/992026-09-04
RSS · https://simonwillison.net/atom/everything/practitionerhealthy5/302026-09-04
RSS · https://www.interconnects.ai/feedpractitionerhealthy0/202026-09-04
Stack Exchangepractitionerhealthy1/12026-09-04
  1. 01

    Catalog, not copy

    OpenRouter and models.dev supply structured facts such as pricing, context, modalities, release dates, capabilities, and third-party benchmark identifiers. Newly cataloged models are joined automatically to benchmark rows by permanent slug. We never republish vendor descriptions.

  2. 02

    Evidence has a type

    Benchmarks, independent evaluators, practitioner reports, community discussions, operational telemetry, and official release evaluations are kept distinct. LiveBench, BFCL, SWE-rebench, Terminal-Bench, OSWorld, LMArena, Design Arena, Aider, SWE-bench, and configured task suites retain their evaluator and harness lineage. First-party release claims can never count as independent confirmation.

  3. 03

    The question sets the weights

    There is no universal best model. Each use case defines its own task-specific benchmark, price, context, adoption, reliability, and ownership signals. Popularity is capped as an adoption signal rather than treated as quality. Only the newest coherent benchmark snapshots participate.

  4. 04

    Coverage has a named status

    A model with one excellent result should not automatically equal one tested broadly. Coverage is measured against every intended feed, including a configured feed that is failing or not yet available. Established means at least 60% of the intended weighted benchmark mix across at least two independent evaluator families. One independent result or official release evidence is provisional.

  5. 05

    The writer does not rank

    The deterministic scorer owns candidate order. A language model turns supplied evidence into readable verdicts but cannot promote a model over the computed order.

  6. 06

    Every source is checked

    Citations are checked for model ownership and source existence, then a separate validation pass tests whether each excerpt supports the prose. One conversation counts once even if several comments were collected. Evaluators, repositories, publications, communities, and negative reports retain separate lineage. Unsupported output fails closed.

  7. 07

    Freshness is explicit

    Catalog, every benchmark source, official vendor release discovery, evidence, and synthesis each refresh daily. New model announcements are found through reviewed first-party sitemaps and exact model identity checks. Relevant evidence, material catalog changes, new benchmark snapshots, or the freshness ceiling trigger regeneration for affected use cases.

  8. 08

    Constraints rerank a wider field

    Editorial pages stay concise, but synthesis retains a larger pool of citation-audited candidates. Picker constraints such as open weights, price, context, modalities, and deployment route deterministically rerank that pool. A one-sided candidate can appear there only when its supported evidence is shown and the missing side is stated plainly.

  9. 09

    One answer survives failures

    Every selectable use case has one explicit top recommendation. If an upstream source or synthesis run fails, the last completely verified answer remains published and the pipeline raises an operational alert. Partial data never replaces it.

known limits

A useful map is still not the territory.

Community evidence leans developer-heavy, public discussion can be promotional or mistaken, benchmark coverage varies, and listed API prices do not capture hosting or latency. Official release evidence closes the launch-day information gap but remains visibly provisional until independent sources reproduce relevant results. Treat the top pick as the strongest model to test first, then validate it against your prompts, latency target, and failure budget.