Best AI models for automation.

286 available models considered · 42 with comparable evidence

An automated workflow asks a model to choose a next step, call a tool and respond to what comes back. This shortlist includes models with confirmed tool support and emphasizes instruction following, agentic coding and reasoning. The surrounding workflow determines which actions are available.

For a balanced task, the current best measured match is DeepSeek V4.1 Flash at 77.6/100 task fit. Change the difficulty or add requirements in the finder.

Choose task difficulty

Change the range

Quick tasks

Choose a tool for a well-defined request and populate its arguments.

Open finder →

Default range

Balanced

Chain a few tool calls, use their results and handle a routine missing field.

Open finder →

Change the range

Complex tasks

Coordinate several steps with dependencies, recover from tool failures and revise a plan as results arrive.

Open finder →

Balanced task range

Best measured matches

Add requirements in the finder →

Task fit applies the published weights for automation. Price helps choose a provider after capability; it does not raise a weaker model into this range.

  1. Best match · Best value · Lowest cost

    DeepSeek V4.1 Flash

    DeepSeek · open weights

    Task fit

    77.6 / 100

    Capability 81.1/100

    Lowest eligible listed rate

    $0.0422 in · $0.0844 out / M tokens

    TrustedRouter

    Price checked 22 Sept 2026

    Measured as deepseek-v4.1-flash-max · max effort. Source ↗

  2. Task fit

    77.4 / 100

    Capability 83.4/100

    Lowest eligible listed rate

    $10 in · $50 out / M tokens

    Amazon Bedrock

    Price checked 22 Sept 2026

    Measured as claude-fable-5-1-max-effort · max effort. Source ↗

  3. Task fit

    77.4 / 100

    Capability 81.6/100

    Lowest eligible listed rate

    $1.25 in · $0.15 out / M tokens

    NanoGPT

    Price checked 22 Sept 2026

    Measured as muse-spark-1.3-xhigh · xhigh effort. Source ↗

  4. Task fit

    76.9 / 100

    Capability 83/100

    Lowest eligible listed rate

    $10 in · $50 out / M tokens

    Amazon Bedrock

    Price checked 22 Sept 2026

    Measured as claude-fable-5-max-effort · max effort. Source ↗

  5. Task fit

    76 / 100

    Capability 82.2/100

    Lowest eligible listed rate

    $9.5 in · $47.5 out / M tokens

    Jiekou

    Price checked 22 Sept 2026

    Measured as gpt-6-astra-max · max effort. Source ↗

  6. Qwen3.8 Max

    Alibaba Cloud

    Task fit

    75.1 / 100

    Capability 78.5/100

    Lowest eligible listed rate

    $1.3 in · $3.9 out / M tokens

    Eden AI

    Price checked 22 Sept 2026

    Measured as qwen3.8-max · published configuration; effort not stated. Source ↗

  7. Kimi K3

    Moonshot AI · open weights

    Task fit

    75 / 100

    Capability 79.2/100

    Lowest eligible listed rate

    $2 in · $0.2 out / M tokens

    NanoGPT

    Price checked 22 Sept 2026

    Measured as kimi-k3 · published configuration; effort not stated. Source ↗

  8. Task fit

    74.3 / 100

    Capability 81.1/100

    Lowest eligible listed rate

    $2 in · $10 out / M tokens

    Nous Portal

    Price checked 22 Sept 2026

    Measured as gpt-5.6-sol-max · max effort. Source ↗

  9. Task fit

    74.2 / 100

    Capability 78.8/100

    Lowest eligible listed rate

    $0.75 in · $3.75 out / M tokens

    DeepInfra

    Price checked 22 Sept 2026

    Measured as gemini-3.7-flash-high · high thinking. Source ↗

  10. Task fit

    73.9 / 100

    Capability 78/100

    Lowest eligible listed rate

    $1.25 in · $0.15 out / M tokens

    NanoGPT

    Price checked 22 Sept 2026

    Measured as muse-spark-1.2-xhigh · xhigh effort. Source ↗

How to read this guide

Automation task fit is a weighted capability proxy, not a measured workflow completion rate. Tool support alone does not establish reliable planning or compatibility with your application.

The recommendations use the same Balanced settings as the interactive finder. See the Capability methodology for sources, exact identity rules, task weights and difficulty ranges.

Other tasksgeneral usewritingresearchcodingdata analysis