Ranking IA

The best AI model for each task.

Pick a category. Use “Free only” to see just the models you can use without paying.

Updated 6 October 2026 · 358 of 358 models evaluated · judges: jev-router → claude-sonnet-4.5

Coding & Debugging

Writing, reviewing and fixing code.

  1. 1
    Claude Opus 5.5
    Anthropic
    “Flagship model explicitly built for demanding coding and multi-step changes in large codebases with reasoning support”
    10/10$4 / $20
    per 1M tokens (in / out)
  2. 2
    Gemini 3.1 Pro Preview
    Google
    “Frontier reasoning model with enhanced software engineering performance and improved agentic reliability for complex workflows”
    9/10$2 / $12
    per 1M tokens (in / out)
  3. 3
    Claude Fable 5.1
    Anthropic
    “Biggest gains in agentic coding with improvements in long code refactors and front-end development workflows”
    9/10$10 / $50
    per 1M tokens (in / out)
Also good: KAT-Coder-Pro V2.5 9Kimi K2.7 Code 9

How to use free models →

Method

How we build the ranking.

  1. CatalogueWe read the full OpenRouter catalogue: price, context and capabilities of every model. Routers and “-latest” aliases are left out.
  2. ScreeningA cheap model scores every model, 20 at a time, across ten categories. The best 12 in each category go through.
  3. RankingA stronger model compares the 12 finalists side by side and orders them. Only this pass awards medals.
  4. RulesSome categories require real capabilities: Images needs vision, Long Documents needs 200K context, Agents needs tool calling.
  5. SortingThe judge's score carries most weight. Price only separates near-ties.

Scores are an AI model's judgement based on published descriptions, not benchmark measurements. Use the ranking as a starting point and test with your own work. How to use free models →