Best AI models for coding.

334 available models considered · 42 with comparable evidence

A short function, a debugging session and a change across a repository ask different things of a model. This shortlist emphasizes coding and agentic coding results, with reasoning and instruction following also contributing. Choose a difficulty level around the scope of the change you need.

For a balanced task, the current best measured match is Claude Fable 5.1 at 80.3/100 task fit. Change the difficulty or add requirements in the finder.

Choose task difficulty

Change the range

Quick tasks

Explain a function, write a small test or fix a local syntax or logic error.

Open finder →

Default range

Balanced

Implement a bounded feature, review a change or debug an issue across a few files.

Open finder →

Change the range

Complex tasks

Plan a migration, investigate an unfamiliar repository or coordinate a change across several components.

Open finder →

Balanced task range

Best measured matches

Add requirements in the finder →

Task fit applies the published weights for coding. Price helps choose a provider after capability; it does not raise a weaker model into this range.

  1. Best match

    Claude Fable 5.1

    Anthropic

    Task fit

    80.3 / 100

    Capability 83.4/100

    Lowest eligible listed rate

    $10 in · $50 out / M tokens

    Amazon Bedrock

    Price checked 22 Sept 2026

    Measured as claude-fable-5-1-max-effort · max effort. Source ↗

  2. Best value · Lowest cost

    DeepSeek V4.1 Flash

    DeepSeek · open weights

    Task fit

    79.9 / 100

    Capability 81.1/100

    Lowest eligible listed rate

    $0.0422 in · $0.0844 out / M tokens

    TrustedRouter

    Price checked 22 Sept 2026

    Measured as deepseek-v4.1-flash-max · max effort. Source ↗

  3. Task fit

    78.9 / 100

    Capability 83/100

    Lowest eligible listed rate

    $10 in · $50 out / M tokens

    Amazon Bedrock

    Price checked 22 Sept 2026

    Measured as claude-fable-5-max-effort · max effort. Source ↗

  4. Task fit

    77.7 / 100

    Capability 81.6/100

    Lowest eligible listed rate

    $1.25 in · $0.15 out / M tokens

    NanoGPT

    Price checked 22 Sept 2026

    Measured as muse-spark-1.3-xhigh · xhigh effort. Source ↗

  5. Task fit

    77 / 100

    Capability 80.1/100

    Lowest eligible listed rate

    $5 in · $25 out / M tokens

    Amazon Bedrock

    Price checked 22 Sept 2026

    Measured as claude-opus-5-max-effort · max effort. Source ↗

  6. Task fit

    76.2 / 100

    Capability 81.1/100

    Lowest eligible listed rate

    $2 in · $10 out / M tokens

    Nous Portal

    Price checked 22 Sept 2026

    Measured as gpt-5.6-sol-max · max effort. Source ↗

  7. Kimi K3

    Moonshot AI · open weights

    Task fit

    76.2 / 100

    Capability 79.2/100

    Lowest eligible listed rate

    $2 in · $0.2 out / M tokens

    NanoGPT

    Price checked 22 Sept 2026

    Measured as kimi-k3 · published configuration; effort not stated. Source ↗

  8. Task fit

    75.6 / 100

    Capability 82.2/100

    Lowest eligible listed rate

    $9.5 in · $47.5 out / M tokens

    Jiekou

    Price checked 22 Sept 2026

    Measured as gpt-6-astra-max · max effort. Source ↗

  9. Task fit

    74.9 / 100

    Capability 78.8/100

    Lowest eligible listed rate

    $0.75 in · $3.75 out / M tokens

    DeepInfra

    Price checked 22 Sept 2026

    Measured as gemini-3.7-flash-high · high thinking. Source ↗

  10. Task fit

    74.4 / 100

    Capability 80.2/100

    Lowest eligible listed rate

    $4.5455 in · $27.2727 out / M tokens

    Poe

    Price checked 22 Sept 2026

    Measured as gpt-5.5-xhigh · xhigh effort. Source ↗

How to read this guide

Published coding results depend on the benchmark and evaluation setup. They do not predict success on your repository or guarantee that a provider reproduces the measured configuration.

The recommendations use the same Balanced settings as the interactive finder. See the Capability methodology for sources, exact identity rules, task weights and difficulty ranges.

Other tasksgeneral usewritingresearchdata analysisautomation