Best AI models for coding.
334 available models considered · 42 with comparable evidence
A short function, a debugging session and a change across a repository ask different things of a model. This shortlist emphasizes coding and agentic coding results, with reasoning and instruction following also contributing. Choose a difficulty level around the scope of the change you need.
For a balanced task, the current best measured match is Claude Fable 5.1 at 80.3/100 task fit. Change the difficulty or add requirements in the finder.
Choose task difficulty
Change the range
Quick tasks
Explain a function, write a small test or fix a local syntax or logic error.
Open finder →Default range
Balanced
Implement a bounded feature, review a change or debug an issue across a few files.
Open finder →Change the range
Complex tasks
Plan a migration, investigate an unfamiliar repository or coordinate a change across several components.
Open finder →Balanced task range
Best measured matches
Task fit applies the published weights for coding. Price helps choose a provider after capability; it does not raise a weaker model into this range.
Task fit
80.3 / 100
Capability 83.4/100
Lowest eligible listed rate
$10 in · $50 out / M tokens
Amazon Bedrock ↗Price checked 22 Sept 2026
Measured as claude-fable-5-1-max-effort · max effort. Source ↗
Task fit
79.9 / 100
Capability 81.1/100
Lowest eligible listed rate
$0.0422 in · $0.0844 out / M tokens
TrustedRouter ↗Price checked 22 Sept 2026
Measured as deepseek-v4.1-flash-max · max effort. Source ↗
- Claude Fable 5 ↗
Anthropic
Task fit
78.9 / 100
Capability 83/100
Lowest eligible listed rate
$10 in · $50 out / M tokens
Amazon Bedrock ↗Price checked 22 Sept 2026
Measured as claude-fable-5-max-effort · max effort. Source ↗
- Muse Spark 1.3 ↗
Meta
Task fit
77.7 / 100
Capability 81.6/100
- Claude Opus 5 ↗
Anthropic
Task fit
77 / 100
Capability 80.1/100
Lowest eligible listed rate
$5 in · $25 out / M tokens
Amazon Bedrock ↗Price checked 22 Sept 2026
Measured as claude-opus-5-max-effort · max effort. Source ↗
- GPT-5.6 Sol ↗
OpenAI
Task fit
76.2 / 100
Capability 81.1/100
Lowest eligible listed rate
$2 in · $10 out / M tokens
Nous Portal ↗Price checked 22 Sept 2026
Measured as gpt-5.6-sol-max · max effort. Source ↗
- Kimi K3 ↗
Moonshot AI · open weights
Task fit
76.2 / 100
Capability 79.2/100
- GPT-6 Astra ↗
OpenAI
Task fit
75.6 / 100
Capability 82.2/100
- Gemini 3.7 Flash ↗
Google
Task fit
74.9 / 100
Capability 78.8/100
Lowest eligible listed rate
$0.75 in · $3.75 out / M tokens
DeepInfra ↗Price checked 22 Sept 2026
Measured as gemini-3.7-flash-high · high thinking. Source ↗
- GPT-5.5 ↗
OpenAI
Task fit
74.4 / 100
Capability 80.2/100
How to read this guide
Published coding results depend on the benchmark and evaluation setup. They do not predict success on your repository or guarantee that a provider reproduces the measured configuration.
The recommendations use the same Balanced settings as the interactive finder. See the Capability methodology for sources, exact identity rules, task weights and difficulty ranges.