Best AI models for automation.
286 available models considered · 42 with comparable evidence
An automated workflow asks a model to choose a next step, call a tool and respond to what comes back. This shortlist includes models with confirmed tool support and emphasizes instruction following, agentic coding and reasoning. The surrounding workflow determines which actions are available.
For a balanced task, the current best measured match is DeepSeek V4.1 Flash at 77.6/100 task fit. Change the difficulty or add requirements in the finder.
Choose task difficulty
Change the range
Quick tasks
Choose a tool for a well-defined request and populate its arguments.
Open finder →Default range
Balanced
Chain a few tool calls, use their results and handle a routine missing field.
Open finder →Change the range
Complex tasks
Coordinate several steps with dependencies, recover from tool failures and revise a plan as results arrive.
Open finder →Balanced task range
Best measured matches
Task fit applies the published weights for automation. Price helps choose a provider after capability; it does not raise a weaker model into this range.
Task fit
77.6 / 100
Capability 81.1/100
Lowest eligible listed rate
$0.0422 in · $0.0844 out / M tokens
TrustedRouter ↗Price checked 22 Sept 2026
Measured as deepseek-v4.1-flash-max · max effort. Source ↗
- Claude Fable 5.1 ↗
Anthropic
Task fit
77.4 / 100
Capability 83.4/100
Lowest eligible listed rate
$10 in · $50 out / M tokens
Amazon Bedrock ↗Price checked 22 Sept 2026
Measured as claude-fable-5-1-max-effort · max effort. Source ↗
- Muse Spark 1.3 ↗
Meta
Task fit
77.4 / 100
Capability 81.6/100
- Claude Fable 5 ↗
Anthropic
Task fit
76.9 / 100
Capability 83/100
Lowest eligible listed rate
$10 in · $50 out / M tokens
Amazon Bedrock ↗Price checked 22 Sept 2026
Measured as claude-fable-5-max-effort · max effort. Source ↗
- GPT-6 Astra ↗
OpenAI
Task fit
76 / 100
Capability 82.2/100
- Qwen3.8 Max ↗
Alibaba Cloud
Task fit
75.1 / 100
Capability 78.5/100
- Kimi K3 ↗
Moonshot AI · open weights
Task fit
75 / 100
Capability 79.2/100
- GPT-5.6 Sol ↗
OpenAI
Task fit
74.3 / 100
Capability 81.1/100
Lowest eligible listed rate
$2 in · $10 out / M tokens
Nous Portal ↗Price checked 22 Sept 2026
Measured as gpt-5.6-sol-max · max effort. Source ↗
- Gemini 3.7 Flash ↗
Google
Task fit
74.2 / 100
Capability 78.8/100
Lowest eligible listed rate
$0.75 in · $3.75 out / M tokens
DeepInfra ↗Price checked 22 Sept 2026
Measured as gemini-3.7-flash-high · high thinking. Source ↗
- Muse Spark 1.2 ↗
Meta
Task fit
73.9 / 100
Capability 78/100
How to read this guide
Automation task fit is a weighted capability proxy, not a measured workflow completion rate. Tool support alone does not establish reliable planning or compatibility with your application.
The recommendations use the same Balanced settings as the interactive finder. See the Capability methodology for sources, exact identity rules, task weights and difficulty ranges.