The best AI model for each task.
Pick a category. Use “Free only” to see just the models you can use without paying.
Coding & Debugging
Writing, reviewing and fixing code.
- 1Claude Opus 5.5Anthropic“Flagship model explicitly built for demanding coding and multi-step changes in large codebases with reasoning support”
- 2Gemini 3.1 Pro PreviewGoogle“Frontier reasoning model with enhanced software engineering performance and improved agentic reliability for complex workflows”
- 3Claude Fable 5.1Anthropic“Biggest gains in agentic coding with improvements in long code refactors and front-end development workflows”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1Laguna S 2.1freePoolside“Dedicated coding agent model, 70.2% Terminal-Bench 2.1”
- 2Gemma 4 31BfreeGoogle“Competent open model for code generation and multi-file review”
- 3Nemotron 3 UltrafreeNVIDIA“Frontier-scale MoE with strong software engineering capability”
Reasoning & Maths
Multi-step deduction, planning and logic.
- 1GPT-6 Astra ProOpenAI“Pro reasoning mode explicitly optimized for complex tasks, flagship model with extended thinking capability”
- 2Gemini 3.1 Pro PreviewGoogle“Frontier reasoning model with enhanced software engineering and agentic reliability for complex workflows”
- 3GPT-6 AstraOpenAI“Flagship model with strong reasoning capability, scientific work strength, slightly below pro mode”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1Ling 3.1 FlashfreeinclusionAI“Hybrid reasoning MoE with 25B active parameters”
- 2Gemma 4 31BfreeGoogle“Configurable thinking mode elevates analytical and mathematical performance”
- 3Nemotron 3 UltrafreeNVIDIA“Explicitly a frontier-reasoning and orchestration model”
Writing & Editing
Drafting, rewriting and changing tone.
- 1Claude Opus 5.5Anthropic“Flagship model with exceptional prose quality, nuanced tone control, and creative writing capabilities across all formats”
- 2Claude Opus 5Anthropic“Strong general writing with excellent clarity and style, slightly behind 5.5 in creative nuance and polish”
- 3Claude Opus 4.8Anthropic“Proven writing excellence with sophisticated language control, though superseded by newer Opus iterations in refinement”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1Gemma 4 31BfreeGoogle“Clean, natural writing style suitable for professional text creation”
- 2Nemotron 3 UltrafreeNVIDIA“Large model produces polished, controllable prose”
- 3InklingfreeThinking Machines“Solid generalist prose”
Long Documents
Whole books, contracts and transcripts (200K+ tokens).
- 1GPT-5.5 ProOpenAI“Million-plus token context with deep reasoning optimized for complex high-stakes workloads requiring extended analysis”
- 2GPT-5.5OpenAI“Million-plus token context with improved efficiency though lacks explicit reasoning mode for complex analysis”
- 3Gemini 3.1 Pro PreviewGoogle“Million-token context with enhanced software engineering performance and improved agentic reliability across complex workflows”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1Gemma 4 26B A4BfreeGoogle“Handles 256k token context window efficiently due to sparse activation”
- 2Gemma 4 31BfreeGoogle“Solid 256k context window for document collections and logs”
- 3Nemotron 3 UltrafreeNVIDIA“262k window plus hybrid attention for very long documents”
Data Extraction
Pulling structured data from messy text: JSON, tables.
- 1Claude Sonnet 5Anthropic“Frontier Sonnet with adaptive thinking and proven excellence in structured output and professional data work”
- 2Claude Sonnet 4.6Anthropic“Exceptional at complex codebase navigation and iterative work, ideal for large-scale extraction tasks”
- 3Claude Opus 5.5Anthropic“Flagship reasoning model excels at multi-step extraction workflows and handling complex structured data”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1LFM2.5-2.6BfreeLiquidAI“Explicitly targeted at data extraction and RAG pipelines”
- 2Gemma 4 26B A4BfreeGoogle“Efficient field parsing and classification at high throughput”
- 3Nemotron 3 UltrafreeNVIDIA“Strong instruction following yields dependable structured output”
Images & Screenshots
Charts, diagrams, photos and scanned documents.
- 1Claude Opus 5.5Anthropic“Flagship model explicitly strong at visual analysis with massive context and latest generation capabilities”
- 2Claude Opus 5Anthropic“Explicitly highlighted for visual analysis strength, proven flagship performance with full multimodal support”
- 3Gemini 3.1 Pro PreviewGoogle“Latest Gemini with comprehensive video and image support, frontier reasoning enhances visual understanding”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1Gemma 4 26B A4BfreeGoogle“Capable multimodal comprehension of images and video clips”
- 2Gemma 4 31BfreeGoogle“Strong visual perception over diagrams, video, and UI interfaces”
- 3Nemotron 3 Nano OmnifreeNVIDIA“Native image and video input for perception sub-agent duties”
Agents & Tool Use
Tool calling and autonomous workflows.
- 1Claude Opus 4.7Anthropic“Built specifically for long-running asynchronous agents with strongest agentic workflow performance in the Opus line”
- 2Claude Opus 5.5Anthropic“Flagship model with reasoning support excels at multi-step changes and complex agentic tasks”
- 3Claude Fable 5.1Anthropic“Biggest gains in agentic coding and long-running workflows over already strong Fable 5 baseline”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1Ling 3.1 FlashfreeinclusionAI“Large MoE tuned for agentic and tool-driven workflows”
- 2Laguna S 2.1freePoolside“Built for terminal agent loops and tool use”
- 3Gemma 4 31BfreeGoogle“Native function calling support makes it viable for agentic loops”
Fast & Cheap
High volume, where cost matters most.
- 1Gemini 2.5 Flash LiteGoogle“Purpose-built for ultra-low latency and cost efficiency with optimized throughput for high-volume production workloads”
- 2Gemini 3.1 Flash LiteGoogle“GA model explicitly optimized for low-latency high-volume workloads with excellent multimodal support and production stability”
- 3Gemini 3.1 Flash Lite PreviewGoogle“Outperforms 2.5 Flash Lite on quality while maintaining high-volume optimization but preview status reduces production confidence”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1Gemma 4 26B A4BfreeGoogle“MoE architecture activates only 3.8B parameters for rapid low-cost inference”
- 2LFM2.5-2.6BfreeLiquidAI“2.6B compact model, very cheap for high-volume simple tasks”
- 3Nemotron 3.5 LightningfreeNVIDIA“3B active params deliver very high throughput at low cost”
Translation
Many languages, Portuguese included.
- 1Gemini 2.5 ProGoogle“Gemini models have consistently demonstrated superior multilingual capabilities across diverse languages and translation tasks”
- 2Gemini 3.1 Pro PreviewGoogle“Frontier reasoning enhances translation quality and cultural nuance understanding across many languages”
- 3Gemini 2.5 FlashGoogle“Strong multilingual performance with excellent speed-to-quality ratio for translation and localization workflows”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1Ling 3.1 FlashfreeinclusionAI“Broad multilingual coverage, strong in Chinese and English”
- 2Nemotron 3 UltrafreeNVIDIA“Good multilingual ability, primarily English-optimized”
- 3InklingfreeThinking Machines“Generalist multilingual ability”
Summaries
Condensing without losing what matters.
- 1Claude Opus 5.5Anthropic“Flagship reasoning model with 1M context excels at distilling complex information while preserving critical details”
- 2Claude Fable 5.1Anthropic“Mythos-class model optimized for knowledge work with strong long-context comprehension and faithful condensation”
- 3Claude Opus 5Anthropic“Strong reasoning capabilities and visual analysis make it excellent for multi-modal document summarization”
Provisional podium: these results come from the first (screening) pass. The final podium appears after the ranking pass.
- 1Gemma 4 31BfreeGoogle“Reliable summarizer for long papers and meetings”
- 2Gemma 4 26B A4BfreeGoogle“Faithfully summarizes documents across moderate-to-large contexts”
- 3Nemotron 3 UltrafreeNVIDIA“Faithful condensation across long, complex inputs”
How we build the ranking.
- CatalogueWe read the full OpenRouter catalogue: price, context and capabilities of every model. Routers and “-latest” aliases are left out.
- ScreeningA cheap model scores every model, 20 at a time, across ten categories. The best 12 in each category go through.
- RankingA stronger model compares the 12 finalists side by side and orders them. Only this pass awards medals.
- RulesSome categories require real capabilities: Images needs vision, Long Documents needs 200K context, Agents needs tool calling.
- SortingThe judge's score carries most weight. Price only separates near-ties.
Scores are an AI model's judgement based on published descriptions, not benchmark measurements. Use the ranking as a starting point and test with your own work. How to use free models →