How we help you choose a model.

A useful starting point—not an IQ test, a promise of success, or a substitute for trying the model on your work.

Capability score: published evidence, not our own test

Version 1 uses LiveBench, release 2026-06-25. We average the published task results within each of its seven categories, then average those seven category scores equally. The result is shown on a 0–100 scale, rounded to one decimal. We do not blend incompatible benchmark scales or claim to have tested these models ourselves.

The categories are reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following. Scores come from the official public leaderboard data. Refreshing these published results makes no model inference calls and incurs no model API evaluation charges.

Each score belongs to one reviewed model version and the configuration printed beside it. A high-effort result is not evidence for that model at low effort. Names, family membership, parameter count, release date and price are never used to invent a score. Provider availability is catalog information; it does not prove a host reproduces the benchmark configuration or performance.

Task fit: our disclosed interpretation

Task fit is a weighted average of the same seven category scores, rounded to one decimal. These editorial weights are a first version, not validated predictions of success. All weights below are percentages and each row totals 100. General use includes all text models; automation additionally requires catalog-confirmed tool support. Other feature pills are independent requirements.

Task-fit weights in percent
TaskReasoningCodingAgentic codingMathematicsData analysisLanguageInstruction following
General use20%10%5%10%10%20%25%
Writing10%0%0%0%0%45%45%
Research35%0%0%10%15%20%20%
Coding15%40%30%5%0%0%10%
Data analysis20%10%0%20%35%0%15%
Automation20%10%25%0%15%0%30%

Writing uses language and instruction-following results as proxies; it does not measure creativity or your taste. Research measures reasoning over supplied material, not search quality, factual reliability or citations. Automation uses reasoning, coding and instruction results plus tool support, not an end-to-end agent evaluation. Vision support does not imply a measured vision score.

Difficulty and the first rows

Quick tasks allow models within 15 task-fit points of the strongest currently scored model for that task and its hard feature requirements. Balanced allows 8 points; complex tasks allows 3. These are transparent comparison bands, not thresholds proven to solve your job. The slider neither changes inference settings nor measures response speed. Price, provider and privacy filters do not lower this reference score.

Best match is the highest task fit among eligible models in the band. Best value is the lowest eligible rate within 5 points of the reference—or the narrower selected band. Lowest cost is the lowest eligible rate anywhere in the selected band. A model can hold several badges; we show it once rather than pretending there are three different winners. The rest of the band stays listed, and models outside it remain accessible.

Prices are compared using a reference mix of three input tokens to one output token, not a forecast of your monthly bill. Rows show input and output prices separately per million tokens. Your token mix, reasoning, retries, context, caching and service tier can change what you pay. Unknown prices do not exclude a model or count as zero; an explicit price limit requires a comparable price.

Coverage, freshness and exclusions

Every daily model refresh applies the same rules. New models enter the catalog; an exact reviewed mapping is needed before a score appears. Unscored models remain visible, newest first, as long as an eligible provider listing exists. Missing evidence means unknown, not poor performance. We do not borrow another version's score or promote a score based on a model's name.

We fetch the category definitions and results from the same source commit. A score must have all seven categories, no conflicting mapping, a source update within 180 days and a successful retrieval within seven days. A source update is not an individual model's test date; LiveBench does not provide that date in this table. Old or unavailable evidence becomes unscored. A failed refresh preserves the last successful source date and retrieval date.

Providers must have a usable link and a listing checked within three days. Deprecated, stale, unofficial, unverified and quarantined offers do not establish availability here. Unexplained prices, special tiers, quantizations, variants and free tiers may establish that a provider offers the model, but their rates cannot lead or be called the lowest comparable price. Full provider pages retain the broader catalog and its caveats.

SWE-bench and BFCL remain supplementary evidence on model pages. They do not gate this broader finder, and their scores are not mixed into Capability. We deliberately leave gaps rather than fabricate universal coverage. This methodology can be revised, with corresponding code and source changes reviewed together.

Try the model finder → · Browse all models →