Our methodology

Trust the process.
Then inspect it.

This field guide is automated and non-commercial. That makes the method, limits, and audit trail more important, not less.

  1. 01

    Catalog, not copy

    OpenRouter and models.dev supply structured facts such as pricing, context, modalities, release dates, and capabilities. We never republish vendor descriptions.

  2. 02

    Evidence has a type

    Benchmarks, independent evaluators, practitioner reports, community discussions, and vendor-community posts are kept distinct. They carry different certainty and independence.

  3. 03

    The question sets the weights

    There is no universal best model. Each use case defines its own benchmark, price, context, popularity, and ownership signals. Only the newest coherent benchmark snapshots participate.

  4. 04

    Coverage affects confidence

    A model with one excellent result should not automatically equal one tested broadly. Scores receive a gentle confidence adjustment based on signal coverage.

  5. 05

    The writer does not rank

    The deterministic scorer owns candidate order. A language model turns supplied evidence into readable verdicts but cannot promote a model over the computed order.

  6. 06

    Every source is checked

    Citations are checked for model ownership and source existence, then a separate validation pass tests whether each excerpt supports the prose. One discussion thread counts as one independent source. Unsupported output fails closed.

  7. 07

    Freshness is explicit

    Old evidence ages out. Relevant evidence, catalog changes, benchmark snapshots, and a freshness ceiling trigger per-use-case regeneration. Publication is atomic.

known limits

A useful map is still not the territory.

Community evidence leans developer-heavy, public discussion can be promotional or mistaken, benchmark coverage varies, and listed API prices do not capture hosting or latency. Source families are capped and vendor-affiliated material is not treated as independent confirmation. Treat the top pick as the strongest model to test first, then validate it against your workload.