Which LLM to Choose in 2026? Latest Models, Benchmarks & Rankings
Choosing the right Large Language Model in 2026 depends on balancing four hard constraints — data privacy, latency, cost, and required intelligence level — against the task itself. This guide walks through a five-step decision process: (1) eliminate candidates that fail your non-negotiables, (2) classify the task by intelligence tier, (3) read the right benchmarks for that tier, (4) shortlist 3–5 models and run your own evaluations, and (5) design a routing strategy rather than betting on a single model. Reviewed September 2026, with ranked benchmark scores, a dated release log covering the latest LLM models, and data for GPT-5.6, Claude Opus 4.8, Claude Sonnet 5, Claude Fable 5, GLM-5.2, and Grok 4.5.
Which LLM should you choose in 2026?
There is no single best LLM — the right choice depends on the task. For frontier reasoning and coding, Claude Opus, GPT-5, and Gemini Ultra lead; for cost-efficient high volume, mid-tier models like Claude Sonnet and GPT mini win; and for private or air-gapped deployments, open models such as Llama and Qwen are best. Choose by eliminating models that fail your non-negotiables (privacy, latency, cost), then match the remaining options to your task's intelligence tier.
LLM Rankings 2026: Top AI Models by Benchmark Score
Claude Fable 5 leads coding at 95.0% on SWE-bench Verified, Gemini 3.1 Pro leads science at 94.3% on GPQA Diamond and abstract reasoning at 77.1% on ARC-AGI-2, and GPT-5.2, Gemini 3.1 Pro and Grok 4 all solve AIME 2025 at 100%. Open-weight MiniMax M2.5 reaches 80.2% on SWE-bench. Daily-refreshed scores for every model in the table sit in the LLM leaderboard.
| Rank & model | Lab | Coding (SWE-bench Verified) | Math (AIME 2025) | Reasoning (GPQA Diamond) | Price in / out per 1M |
|---|---|---|---|---|---|
| 1. Claude Fable 5 | Anthropic | 95.0% | — | 92.6% | $10 / $50 |
| 2. Claude Opus 4.8 | Anthropic | 88.6% | — | 93.6% | $5 / $25 |
| 3. Claude Sonnet 5 | Anthropic | 85.2% | — | — | $2 / $10 (intro) |
| 4. Claude Opus 4.6 | Anthropic | 80.8% | 99.8% | 91.3% | $5 / $25 |
| 5. Gemini 3.1 Pro | Google DeepMind | 80.6% | 100% | 94.3% | $2 / $12 |
| 6. MiniMax M2.5 (open weights) | MiniMax | 80.2% | — | — | Self-hosted |
| 7. GPT-5.2 | OpenAI | 80.0% | 100% | 92.4% | — |
| 8. GPT-5.4 | OpenAI | ~80% | 88% | 92.0% | $2.50 / $15 |
| 9. Claude Sonnet 4.6 | Anthropic | 79.6% | ~95% | 74.1% | $3 / $15 |
| 10. Gemini 3 Flash | Google DeepMind | 78.0% | — | 90.4% | $0.50 / $3.00 |
| 11. GLM-5 (open weights) | Zhipu AI | 77.8% | 92.7% | 86.0% | Self-hosted |
| 12. Kimi K2.5 (open weights) | Moonshot AI | 76.8% | — | 87.6% | Self-hosted |
| 13. Qwen 3.5 (397B) (open weights) | Alibaba | 76.4% | 91.3% | 88.4% | Self-hosted |
| 14. Grok 4 | xAI | — | 100% | — | $3 / $15 |
How to compare LLMs
Compare LLMs in four steps: eliminate models that fail a hard constraint (privacy, deployment mode, context window, latency, budget); pick the benchmark that matches the task; treat scores within two to three points as a tie; then confirm the finalists on 100 to 200 examples from your own workload.
Latest LLM Releases and Model Updates
The releases that changed a selection decision in 2026: GPT-5.6 (July 9), Claude Sonnet 5 (June 30), GLM-5.2 (June 13), plus Claude Opus 4.8, Claude Fable 5 and Grok 4.5. Every entry records the license, context window, price per million tokens, the benchmark delta and whether the weights can be deployed privately.
- July 9, 2026 GPT-5.6 (Luna / Terra / Sol) — OpenAI
- June 30, 2026 Claude Sonnet 5 — Anthropic
- June 13, 2026 GLM-5.2 — Zhipu AI / Z.ai
- 2026 (Anthropic frontier refresh) Claude Opus 4.8 and Claude Fable 5 — Anthropic
- 2026 (xAI coding refresh) Grok 4.5 — xAI
- March 27, 2026 GLM-5.1 — Zhipu AI
- March 18, 2026 MiniMax M2.7 — MiniMax
- LLM Rankings 2026 (Updated September 2026)
- Latest LLM Releases & Model Updates
- 1. Executive Summary
- 2. Understanding LLM Benchmarks
- 3. Current Model Landscape
- Best Open-Source LLMs in 2026
- LLM Context Window Comparison
- 4. Intelligence Level Taxonomy
- 5. Task-to-Model Matching Framework
- 6. The LLM Selection Decision Process
- 7. Model Routing & Cascade Strategies
- 8. Evaluation Methodology After Shortlisting
- 9. Key Leaderboards & Resources
- 10. Open-Source vs. Proprietary Decision Guide
- 11. Quick Reference: Recommendations by Use Case
- 12. Sources
Executive Summary
Selecting the right LLM is not about finding the "best" model -- it is about finding the right model for your specific task, constraints, and budget. The most common mistake organizations make is selecting models based on general reputation or top-line benchmark scores without analyzing their actual requirements.
No single model dominates every task. The optimal architecture in 2026 routes different requests to different models based on task complexity, latency requirements, and cost constraints.
- No single model dominates every task. Claude Opus 4.6 leads on coding (Arena code Elo 1548) and nuanced writing, GPT-5.4 excels at structured reasoning and computer use (75% OSWorld, surpassing human expert baseline), Gemini 3.1 Pro wins on abstract reasoning (ARC-AGI-2), multimodal input, and scientific benchmarks (GPQA 94.3%), Grok 4 leads HLE (50.7%), and new open-source entrants like MiniMax M2.5/M2.7, GLM-5/5.1, and Kimi K2.5 now rival frontier proprietary models on SWE-bench.
- Benchmark scores are necessary but insufficient. Models scoring within 2-3% of each other on MMLU are functionally indistinguishable on that metric -- your specific use case is the real differentiator.
- Effective task definition often matters more than model selection. Well-crafted prompts with a mid-tier model frequently outperform poorly prompted frontier models.
- The optimal architecture in 2026 routes different requests to different models based on task complexity, latency requirements, and cost constraints.
Understanding LLM Benchmarks
What are LLM benchmarks?
An LLM benchmark is a standardized test that scores a language model on one capability: GPQA Diamond for scientific reasoning, SWE-bench Verified for real-world coding, AIME 2025 for math, ARC-AGI-2 for abstract reasoning, BFCL v4 for tool use. MMLU, GSM8K and HumanEval are saturated in 2026 and no longer separate frontier models.
Treat any two scores within 2–3% of each other as a tie, and read the ranked results in the 2026 LLM rankings or the daily-refreshed LLM benchmark leaderboard.
2.1 Major Benchmark Registry
| Benchmark | What It Measures | Format | Questions | Difficulty | Saturation |
|---|---|---|---|---|---|
| MMLU | Broad knowledge across 57 academic subjects (STEM, humanities, social sciences, professional) | Multiple-choice (4 options) | 16,000+ | Undergrad to professional | Saturated -- top models 88-94% |
| MMLU-Pro | Enhanced MMLU with harder questions and 10 answer options | Multiple-choice (10 options) | ~12,000 | Graduate+ | Active -- 16-33% drop vs MMLU |
| GPQA Diamond | PhD-level science reasoning (biology, physics, chemistry) | Multiple-choice | 448 | Expert-level (PhD experts: 65-74%) | Active differentiator |
| HumanEval | Function-level code generation correctness | Code generation (Python) | 164 | Intermediate programming | Saturated -- top at 95-99% |
| HumanEval+ | Extended HumanEval with more test cases and edge cases | Code generation | 164 (more tests) | Intermediate-Advanced | Active |
| SWE-bench Verified | Real-world software engineering (fixing actual GitHub bugs) | Full repo code modification | 500 | Professional engineer | Gold standard for coding |
| LiveCodeBench | Contamination-free code evaluation from new contest problems | Code generation, self-repair | Rolling | Competitive programming | Active -- continuously updated |
| GSM8K | Grade-school math word problems | Free-form numerical | 8,500 | Elementary-Middle school | Saturated -- top >95% |
| MATH | Competition-level mathematics (AMC, AIME) | Free-form proof/answer | 12,500 | Olympiad | Active for non-reasoning models |
| AIME 2025 | Advanced math olympiad problems | Free-form | 30 | Olympiad | Active -- very hard |
| ARC-AGI 2 | Abstract visual pattern reasoning | Visual pattern completion | Varies | Fluid intelligence | Active -- frontier differentiator |
| IFEval | Instruction-following capability with verifiable constraints | Constrained text generation | ~500 | Varies | Active |
| TruthfulQA | Factual accuracy and resistance to common misconceptions | Multiple-choice / generative | 817 | General knowledge | Contaminated -- being replaced |
| HELM | Holistic evaluation: accuracy, calibration, robustness, fairness, bias, toxicity, efficiency | Multiple metrics across 42 scenarios | Varies | Varies | Active framework |
| BFCL v4 | Function/tool calling accuracy (serial, parallel, multi-turn, agentic) | Function call generation | Varies | API/Agent tasks | De facto standard for tool use |
| RULER | Long-context comprehension (multi-needle retrieval, tracing, aggregation) | Various retrieval/reasoning | Varies | Long-context tasks | Active |
| MMMU Pro | Multimodal academic reasoning across 30+ subjects | Visual + text reasoning | Varies | Graduate+ | Active for vision models |
| Arena Elo (LMSYS) | Human preference in open-ended conversation | Pairwise human comparison | Millions of votes | Real-world preference | Most ecologically valid |
| SWE-bench Pro | Multi-language software engineering with standardized scaffold | Full repo code modification | Varies | Professional engineer | Emerging -- less contamination |
| Humanity's Last Exam | Extremely hard questions from domain experts worldwide | Mixed | 2,500 | Beyond PhD | Active -- <51% for all models |
2.2 Benchmark Comparison Table (Top Models, Updated July 2026)
| Model | GPQA Diamond | MMLU / MMLU-Pro | AIME 2025 | SWE-bench Verified | ARC-AGI 2 | HLE | Arena Elo (Text) | Arena Elo (Code) |
|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 92.6% | 91.5% (MMLU-Pro) | -- | 95.0% | -- | 53.3% | ~1665 | ~1653 |
| Claude Opus 4.8 | 93.6% | -- | -- | 88.6% | -- | 49.8% / 57.9% (tools) | ~1510 | ~1582 |
| Claude Opus 4.6 | 91.3% | 91.1% (MMMLU) | 99.8% | 80.8% | 68.8% | 40.0% | 1502 | 1548 |
| Claude Sonnet 5 | -- | -- | -- | 85.2% | -- | 43.2% / 57.4% (tools) | -- | -- |
| Claude Sonnet 4.6 | 74.1% | 89.3% (MMMLU) | ~95% | 79.6% | ~58% | -- | 1438 | ~1530 |
| GPT-5.6 (Sol) | -- | -- | -- | -- | -- | 47.2% | -- | -- |
| GPT-5.2 | 92.4% | -- | 100% | 80.0% | 52.9% | 35.2% | ~1460 | ~1520 |
| GPT-5.4 | 92.0% | -- | 88% | ~80% | 73.3% / 83.3% (Pro) | 36.6-41.6% | ~1463 | -- |
| Gemini 3.1 Pro | 94.3% | 92.6% (MMMLU) | 100% | 80.6% | 77.1% | 44.7% | ~1492 | ~1480 |
| Gemini 3 Flash | 90.4% | -- | -- | 78.0% | -- | 33.7% | -- | -- |
| Grok 4.5 | -- | -- | -- | -- | -- | -- | -- | -- |
| Grok 4 | -- | -- | 100% | -- | -- | 50.7% | ~1493 | -- |
| Grok 4.20 Beta | -- | -- | -- | -- | -- | -- | ~1505-1535 | -- |
| DeepSeek R1 | 71.5% | 90.8% / 84.0% | ~95% | ~72% | -- | -- | ~1430 | ~1450 |
| DeepSeek V3.2 | -- | -- / 85.0% | 89.3% | 67.8% | -- | -- | 1421 | -- |
| Qwen 3.5 (397B) | 88.4% | -- | 91.3% | 76.4% | -- | -- | -- | -- |
| Qwen3-235B | 88.4% | ~85% | 85.7% | ~75% | -- | -- | ~1400 | -- |
| Kimi K2.5 | 87.6% | -- | -- | 76.8% | -- | 31.5% / 51.8% (tools) | -- | -- |
| GLM-5.2 | 91.2% | -- | -- | -- | 22.8% | 40.5% / 54.7% (tools) | -- | -- |
| GLM-5 | 86.0% | -- | 92.7% | 77.8% | -- | -- | 1451 | -- |
| MiniMax M2.5 | -- | -- | -- | 80.2% | -- | -- | -- | -- |
| MiniMax M2.7 | -- | -- | -- | 78.0% | -- | -- | -- | -- |
| Step-3.5-Flash | -- | -- | 99.8% | 74.4% | -- | -- | -- | -- |
| Llama 4 Maverick | -- | 85.5% | -- | -- | -- | -- | ~1380 | -- |
July 2026 update: Claude Opus 4.8, Sonnet 5, and Fable 5 figures are confirmed against Anthropic's official Claude Sonnet 5 System Card (June 30, 2026) where overlapping; Anthropic did not publish GPQA Diamond, MMLU-Pro, AIME 2025, or ARC-AGI-2 for Sonnet 5 at all (shown as ––, not a sourcing gap). Claude Fable 5's SWE-bench Verified (95.0%) is single-sourced and unconfirmed directly by Anthropic; treat as directional. GPT-5.6 and Grok 4.5 did not publish classic academic benchmarks (GPQA/MMLU/AIME/SWE-bench Verified/ARC-AGI-2) at launch, reporting only agentic/coding evals instead — notably Grok 4.5 SWE-bench Pro 64.7% and GPT-5.6 Terminal-Bench 2.1 88.8% (Sol) / 91.9% (Sol Ultra). GLM-5.2's GPQA Diamond (91.2%) has one conflicting lower secondary report (80.3%); ARC-AGI-2 (22.8%) is confirmed via ARC Prize's verified leaderboard.
2.3 Benchmark Limitations & Saturation
Critical understanding: Benchmarks are indicators, not guarantees.
| Issue | Explanation | Impact |
|---|---|---|
| Saturation | Top models on MMLU (88-94%), GSM8K (>95%), HumanEval (>95%) are indistinguishable | Use GPQA, SWE-bench, ARC-AGI 2, HLE instead for frontier differentiation |
| Data Contamination | Training data may include benchmark questions (confirmed for TruthfulQA, suspected for MMLU) | Inflated scores; prefer rolling benchmarks like LiveCodeBench |
| Evaluation DoF | Prompt framing, few-shot count, chain-of-thought, grading method can move scores 5-15% | Always note evaluation conditions when comparing |
| Multiple-Choice Artifacts | MCQ benchmarks reward test-taking heuristics, not deep reasoning | Prefer open-ended generation benchmarks |
| Scaffold Dependence | SWE-bench scores depend heavily on the agentic scaffold (Claude Code, Codex CLI, etc.) | Compare under same conditions or acknowledge differences |
| Real-World Gap | "Models that dominate leaderboards often underperform in production" (LXT, 2026) | Always validate with your own domain-specific evaluation |
2.4 Which Benchmarks Matter for Which Use Cases
| Use Case | Primary Benchmarks | Secondary Benchmarks | Notes |
|---|---|---|---|
| General Q&A / Knowledge | MMLU-Pro, Arena Elo | MMLU, TruthfulQA | MMLU alone is insufficient; prefer MMLU-Pro |
| Code Generation | SWE-bench Verified, SWE-bench Pro, LiveCodeBench | HumanEval+, BFCL, Aider Polyglot | SWE-bench Pro emerging as successor due to contamination concerns on Verified |
| Mathematical Reasoning | AIME 2025, MATH | GSM8K (floor only) | GSM8K is too easy; use for minimum capability check |
| Scientific Reasoning | GPQA Diamond | HLE | GPQA is the best frontier differentiator |
| Creative Writing | Arena Elo (Creative Writing) | -- | No good automated benchmark; human preference is key |
| Instruction Following | IFEval | Arena Elo | IFEval tests verifiable constraint adherence |
| Tool Use / Function Calling | BFCL v4 | -- | Only reliable tool-use benchmark |
| Long Context Processing | RULER, Needle-in-Haystack | LongGenBench | RULER is more rigorous than basic NIAH |
| Multimodal / Vision | MMMU Pro, Arena Vision | MMMU | Composite scoring: MMMU Pro 60% + Arena Vision 40% |
| Multilingual | MMMLU | MLNeedle | MMMLU extends MMLU across languages |
| Agentic Tasks | SWE-bench, BFCL v4 | WebArena, OSWorld | Still-emerging evaluation landscape |
| Safety / Factuality | HalluLens, SimpleQA | TruthfulQA (legacy) | TruthfulQA is contaminated; prefer newer benchmarks |
Current Model Landscape
How do the leading LLMs compare in 2026?
As of July 2026, Claude Fable 5 leads frontier coding (SWE-bench Verified 95.0%) and HLE (53.3%), Gemini 3.1 Pro leads scientific reasoning (GPQA Diamond 94.3%) and abstract reasoning (ARC-AGI-2 77.1%), and GPT-5.4 leads computer use. Open-weight models have closed the coding gap: MiniMax M2.5 hits 80.2% on SWE-bench. Prices span $0.05 to $50 per million tokens, so the practical comparison is capability per dollar for your task tier — not a single leaderboard rank.
3.1 Frontier Model Comparison
Claude (Anthropic)
| Model | Best For | Strengths | Weaknesses |
|---|---|---|---|
| Claude Fable 5 | Frontier coding, agentic workflows, deepest reasoning available | Leads Opus 4.8 across FrontierCode, CursorBench, ArxivMath, and HLE (53.3%); GPQA 92.6%; MMLU-Pro 91.5%; Mythos-class capability made generally available; 1M context | Most expensive Anthropic model ($10/$50); briefly suspended worldwide June 12–July 1, 2026 under a US export-control order |
| Claude Opus 4.8 | Complex coding, nuanced writing, deep reasoning, extended thinking | GPQA 93.6%; SWE-bench Verified 88.6% (up from 87.6% on Opus 4.7); HLE 57.9% w/ tools; direct upgrade to Opus 4.7 at the same price | Trails Fable 5 on most frontier benchmarks; $5/$25 |
| Claude Opus 4.6 | Complex coding, nuanced writing, deep reasoning, extended thinking | Highest Arena coding Elo (1548); SWE-bench 80.8%; best prose quality; ARC-AGI 2 68.8%; 1M context (GA); OSWorld 72.5% | Most expensive Anthropic model ($5/$25); no native audio/video |
| Claude Sonnet 5 | Near-Opus intelligence at Sonnet pricing for coding, agents, and everyday professional work | SWE-bench Verified 85.2% (Anthropic's first model to clear 80% at launch); HLE 57.4% w/ tools; upgrade to Sonnet 4.6 across agentic coding, search, and multimodal reasoning | Anthropic did not publish GPQA/MMLU-Pro/AIME/ARC-AGI-2 for this model; trails Opus/Mythos-class in almost all internal comparisons |
| Claude Sonnet 4.6 | Balanced quality/cost for production workloads | SWE-bench 79.6%; strong coding; OSWorld 72.5%; GDPval-AA leader (1633 Elo); 1M context (GA); math 89% | Slightly below Opus on quality; GPQA gap (74.1% vs 91.3%) |
| Claude Haiku 4.5 | High-volume, cost-sensitive tasks | Fast; cheap ($1/$5); good for classification/extraction | Not suitable for complex reasoning |
OpenAI
| Model | Best For | Strengths | Weaknesses |
|---|---|---|---|
| GPT-5.6 (Sol) | Agentic/terminal coding workflows | Terminal-Bench 2.1 88.8% (91.9% Sol Ultra); HLE 47.2%; three tiers (Luna/Terra/Sol); released July 9, 2026 | OpenAI did not publish GPQA/MMLU-Pro/AIME/SWE-bench Verified/ARC-AGI-2 at launch; independent evaluator METR flagged high eval-gaming rates; Sol $5/$30 |
| GPT-5.4 | Structured reasoning, computer use, agentic tasks | ARC-AGI-2 73.3%/83.3% Pro; GPQA 92.0%; native computer use; OSWorld 75%; 272K/1.05M context; Arena Elo ~1463 | Expensive for extended context (2x over 272K); output $15/M |
| GPT-5.4 Mini | Cost-efficient mid-tier tasks | SWE-bench Pro 54.4%; GPQA 87.5%; $0.75/$4.50; 400K context; near-flagship performance at lower cost | Lower reasoning ceiling than flagship |
| GPT-5.2 | Math, science, coding at frontier level | 100% AIME 2025; GPQA 92.4%; SWE-bench 80.0% | 400K context; being superseded by 5.4 |
| GPT-5 Nano | Ultra-cheap high-volume processing | $0.05/$0.40 per M tokens; 400K context | Limited reasoning depth |
| o3 | Deep mathematical and logical reasoning | Extended thinking; strong on hard math | Slower; 200K context; higher latency |
| o3 Pro | Maximum reasoning capability | Best reasoning available | Very expensive ($150+); slow |
Google DeepMind
| Model | Best For | Strengths | Weaknesses |
|---|---|---|---|
| Gemini 3.1 Pro | Multimodal tasks, abstract reasoning, large document processing | ARC-AGI-2 leader (77.1%); GPQA 94.3%; HLE 44.7%; Arena Elo ~1492; 1M context; native text/image/audio/video; $2/$12 | Less refined prose than Claude; pricing doubles over 200K context |
| Gemini 3 Flash | High-throughput multimodal at moderate cost | MMMU Pro 81.2%; GPQA 90.4%; SWE-bench 78%; 1M context; 64K max output; 3x faster than 2.5 Pro | Reasoning ceiling lower than Pro on non-coding; $0.50/$3.00 |
| Gemini 3.1 Flash Lite | Ultra-cheap high-volume multimodal | 381 tok/s; GPQA 86.9%; 2.5x faster than 2.5 Flash; $0.25/$1.50 | Lower reasoning depth |
| Gemini 2.5 Pro | Proven production workloads | Well-tested; 1M context; $1.25/$10 | Being superseded by 3.x series |
xAI
| Model | Best For | Strengths | Weaknesses |
|---|---|---|---|
| Grok 4.5 | Coding/agentic workflows built on Cursor session data | SWE-bench Pro 64.7%; Terminal-Bench 2.1 83.3%; SWE Marathon pass@1 29.0% (#1); xAI's first model built specifically for coding after its Cursor acquisition; $2/$6 | xAI did not publish GPQA/MMLU-Pro/AIME/SWE-bench Verified/ARC-AGI-2/HLE at launch; 500K context (smaller than Grok 4's 2M) |
| Grok 4 | Long-context reasoning, hard math | 260K/2M context; HLE leader 50.7%; USAMO'25 leader (61.9%); AIME 100%; Arena Elo ~1493; $3/$15 | Smaller ecosystem; less battle-tested; expensive for extended context |
| Grok 4.20 Beta | Multi-agent collaboration | 2M context; 4-agent parallel debate architecture; lowest hallucination rate (22%); Arena Elo ~1505-1535; $2/$6 | Beta; newer, less proven |
3.2 Specialized & Domain-Specific Models
| Domain | Specialized Models | General Models That Excel | Key Consideration |
|---|---|---|---|
| Medical/Clinical | Med-PaLM 2, Med-Gemini, PMC-LLaMA, GatorTronGPT, BioMistral | Claude Opus (low hallucination), Gemini Pro | Regulatory compliance (FDA, HIPAA); 85% of healthcare leaders exploring GenAI (McKinsey 2025) |
| Legal | LegalBERT, Harvey AI (proprietary), SaulLM | Claude Opus (document analysis), GPT-5 | Accuracy paramount; hallucination is liability |
| Finance | BloombergGPT, FinGPT, InvestLM | GPT-5 (SEC filing analysis), Gemini Pro | Real-time data needs; regulatory compliance |
| Code | Qwen3-Coder, MiniMax M2.5/M2.7, GLM-5/5.1, DeepSeek-Coder V3, StarCoder2, Codestral | Claude Opus 4.6, GPT-5.4 | SWE-bench Verified / SWE-bench Pro are the benchmarks to watch; MiniMax M2.5 now matches proprietary frontier; Gemini 3 Flash (78%) beats many larger models |
| Multilingual | NLLB, SeamlessM4T | Qwen3 (201 languages), Mistral Large 3 (80+) | Test in YOUR target languages specifically |
3.3 Pricing Comparison (Updated July 2026, per 1M tokens)
| Tier | Model | Input | Output | Cost Rating |
|---|---|---|---|---|
| Free/Near-Free | Llama 4 Scout/Maverick (self-hosted) | $0.00 | $0.00 | Compute only |
| Ultra-Budget | GPT-5 Nano | $0.05 | $0.40 | Extremely cheap |
| Ultra-Budget | DeepSeek V3.2 | -- | -- | Extremely cheap |
| Budget | Gemini 3.1 Flash Lite | $0.25 | $1.50 | Very affordable |
| Budget | Gemini 3 Flash | $0.50 | $3.00 | Strong multimodal value |
| Mid-Range | GPT-5.4 Mini | $0.75 | $4.50 | Excellent value; GPQA 87.5% |
| Mid-Range | Claude Haiku 4.5 | $1.00 | $5.00 | Good value |
| Mid-Range | GPT-5.6 (Luna) | $1.00 | $6.00 | Smallest GPT-5.6 tier |
| Mid-Range | Gemini 2.5 Pro | $1.25 | $10.00 | Strong value |
| Mid-Range | GLM-5.2 (hosted) | ~$1.40 | ~$4.40 | No official Z.ai per-token rate published; DeepInfra/Fireworks/Together hosting range shown |
| Premium | Gemini 3.1 Pro | $2.00 | $12.00 | Best frontier value |
| Premium | Grok 4.20 Beta | $2.00 | $6.00 | Multi-agent; 2M context |
| Premium | Grok 4.5 | $2.00 | $6.00 | Coding/agentic; cache-hit $0.50/M (-75%); 500K context |
| Premium | Claude Sonnet 5 | $2.00 (intro) | $10.00 (intro) | Introductory through Aug 31, 2026; reverts to $3/$15 (same as Sonnet 4.6) |
| Premium | GPT-5.4 | $2.50 | $15.00 | Strong reasoning |
| Premium | GPT-5.6 (Terra) | $2.50 | $15.00 | Mid GPT-5.6 tier |
| Premium | Claude Sonnet 4.6 | $3.00 | $15.00 | Quality premium |
| Frontier | Grok 4 | $3.00 | $15.00 | HLE leader; 260K context |
| Frontier | Claude Opus 4.6 | $5.00 | $25.00 | Premium quality |
| Frontier | Claude Opus 4.8 | $5.00 | $25.00 | Unchanged from Opus 4.7; SWE-bench Verified 88.6% |
| Frontier | GPT-5.6 (Sol) | $5.00 | $30.00 | Flagship GPT-5.6 tier; "Sol Ultra" reasoning-max variant priced separately |
| Reasoning | o3 | $2.00 | $8.00 | Variable with thinking |
| Max Reasoning | o3 Pro | ~$150 | -- | Maximum capability, maximum cost |
| Max Reasoning | GPT-5.4 Pro | $30.00 | $180.00 | Maximum GPT-5.4 capability |
| Max Reasoning | Claude Fable 5 | $10.00 | $50.00 | Mythos-class capability made generally available; less than half the price of Claude Mythos Preview |
To model your own workload rather than list prices, use our LLM pricing calculator or compare LLM token cost across cloud providers.
3.4 Performance & Latency
| Performance Tier | Models | TTFT* | Throughput | Best For |
|---|---|---|---|---|
| Ultra-Fast | Llama 4 Scout, Llama 3.3 70B (via Groq) | <100ms | 2,500+ tok/s | Real-time chat, high-volume |
| Fast | GPT-5.3 Codex, Step-3.5-Flash, Gemini 3.1 Flash Lite, Gemini Flash | <300ms | 350-1,500 tok/s | Interactive applications |
| Standard | Claude Sonnet 4.6, GPT-5.4, Gemini Pro | 300ms-1s | 100-500 tok/s | Production workloads |
| Deliberate | Claude Opus 4.6, GPT-5.2 | 500ms-2s | 50-200 tok/s | Quality-critical tasks |
| Thinking | o3, o3 Pro, Claude Opus (extended thinking) | 2-30s+ | Variable | Complex reasoning requiring chain-of-thought |
Best Open-Source LLMs in 2026: Open-Weight Model Comparison
For most enterprises in 2026 the strongest open-weight pick is GLM-5.2 (MIT, 753B, GPQA 91.2%) for frontier reasoning, MiniMax M2.5 (80.2% SWE-bench Verified) for coding, Qwen 3.5 for multilingual and vision work, and Phi-4 (14B) when the whole model must fit one GPU.
Open-weight models now span an enormous range of LLM parameter sizes — from 8B models that run on a single GPU to 1T Mixture-of-Experts architectures that rival frontier proprietary systems. Parameter count is one of the most over-indexed signals in model selection; what matters more is how those parameters are activated and trained, what the license permits, and how much VRAM the weights need before you can serve a single request. Licensing and jurisdiction are reviewed separately under model provenance and Chinese open-weight licensing.
| Model | Provider | Parameters | License | Context window | Minimum VRAM* | Key strengths | Best benchmarks |
|---|---|---|---|---|---|---|---|
| Frontier open weights (400B+ total parameters) | |||||||
| Kimi K2.5 | Moonshot AI | 1T MoE | Open-weight | — | ~650 GB | Coding, agentic work (Agent Swarm up to 100 agents), vision | SWE-bench 76.8%; HumanEval 99.0%; GPQA 87.6%; HLE 51.8% (tools) |
| GLM-5.2 | Zhipu AI / Z.ai | 753B (~40B active MoE) | MIT | 1M | ~490 GB | Long-horizon coding at a fraction of frontier proprietary cost | GPQA 91.2%; ARC-AGI-2 22.8% (highest open-weight); HLE 54.7% w/ tools; SWE-bench Pro 62.1% |
| GLM-5 | Zhipu AI | 744B (44B active MoE) | MIT | — | ~485 GB | Coding, multimodal, top Arena Elo among open models | SWE-bench 77.8%; GPQA 86.0%; AIME 92.7%; Arena Elo 1451 |
| GLM-5.1 | Zhipu AI | ~744B MoE | MIT | 200K | ~485 GB | Coding successor to GLM-5 (94% of Claude Opus 4.6 coding performance) | 28% coding improvement over GLM-5 |
| DeepSeek V3.2 | DeepSeek | ~685B MoE | MIT | 128K | ~445 GB | General purpose, coding, exceptional value | AIME 89.3%; MMLU-Pro 85.0%; Arena Elo 1421 |
| DeepSeek R1 | DeepSeek | ~670B MoE | MIT | 128K | ~435 GB | Deep reasoning, math, chain-of-thought | MATH-500 97.3%; GPQA 71.5%; MMLU 90.8% |
| Llama 4 Maverick | Meta | 400B MoE | Llama License | 1M | ~260 GB | General chat; MMLU leader among open models | MMLU 85.5% |
| Mid-size open weights (100B to 400B) | |||||||
| Qwen 3.5 (397B-A17B) | Alibaba | 397B (17B active MoE) | Apache 2.0 | 262K (1M extended) | ~260 GB | Reasoning, math, 201 languages, native vision, 19x faster decoding | GPQA 88.4%; AIME 91.3%; SWE-bench 76.4% |
| Qwen3-235B-A22B | Alibaba | 235B (22B active MoE) | Apache 2.0 | 262K | ~155 GB | Reasoning, math, 201 languages | GPQA 88.4%; AIME 85.7% |
| MiniMax M2.5 | MiniMax | 230B MoE | Modified-MIT | — | ~150 GB | Coding excellence, function calling, office productivity | SWE-bench 80.2%; BFCL 76.8% |
| Step-3.5-Flash | StepFun | 196B (11B active MoE) | Open-weight | 256K | ~130 GB | Ultra-fast reasoning at 100 to 350 tokens per second | AIME 99.8%; SWE-bench 74.4% |
| Mistral Large 3 | Mistral | ~123B | Apache 2.0 | 128K | ~80 GB | European compliance, 80+ languages, strong coding | MMLU ~85.5%; Arena Elo ~1418 |
| Llama 4 Scout | Meta | 109B MoE | Llama License | 10M | ~70 GB | Extreme long context; 17B active parameters per token | 10M token context window |
| GLM-4.7 | Zhipu AI | Varies | MIT | 200K | — | Strong all-rounder, 128K max output | HumanEval 94.2%; AIME 95.7%; GPQA 85.7%; HLE 42.8% |
| Single-GPU class (under 100B) — laptops, workstations and edge nodes | |||||||
| Qwen3-30B | Alibaba | 30B | Apache 2.0 | 262K | ~20 GB | Long-context retrieval on one datacenter GPU; pairs with Qwen3-Embedding-8B | Strong RAG performance at low cost |
| Phi-4 | Microsoft | 14B | MIT | — | ~10 GB | Resource-constrained retrieval and deployment | Runs on a single RTX 4090 |
Key open-source trend (2026): The gap between open-source and proprietary models has effectively closed for coding tasks — MiniMax M2.5 (80.2% SWE-bench) matches Claude Opus 4.6 (80.8%), and GLM-5 leads the Arena Elo among open models at 1451. MIT/Apache 2.0 licensed models now approach proprietary frontier models on nearly all benchmarks, with cost per token dropping 10-100x compared to proprietary APIs. Notable new entrants since early 2026: Qwen 3.5, MiniMax M2.5, Step-3.5-Flash, GLM-5/5.1/5.2, Kimi K2.5. MiniMax M2.7 (March 18, 2026, proprietary weights) introduces "self-evolving" agent capabilities, GLM-5.1 (March 27, 2026) achieves 94% of Claude Opus 4.6 coding performance, and GLM-5.2 (June 13, 2026) posts the highest ARC-AGI-2 score (22.8%) of any open-weight model while undercutting proprietary API pricing by roughly 6x.
Licenses, provenance and export control
Two facts decide whether an open-weight model is usable in a regulated environment, and neither is a benchmark. The first is the license: MIT and Apache 2.0 (DeepSeek, GLM, Qwen, Mistral, Phi-4) permit commercial use, redistribution and fine-tuning; the Llama Community License and MiniMax’s Modified-MIT add named conditions; "open-weight" without a named license means the terms have to be read before deployment. The second is provenance: which lab trained the weights, in which jurisdiction it operates, and whether your contract, sector rules or export-control obligations (ITAR, EAR, CMMC, FedRAMP boundaries) restrict which model families you may run. Weights downloaded once and served air-gapped never call home, which is why open weights are the default choice in air-gapped and export-controlled environments — but the license and jurisdiction review still happens first. The step-by-step version of that review is in fine-tuning and model provenance.
LLM Context Window Comparison: Context Length by Model
What is a context window?
An LLM context window is the maximum amount of text, measured in tokens, that a model can hold in working memory for a single request: the system prompt, the conversation so far, any retrieved documents and the answer it is generating. Once the window is full, the oldest content falls out. In 2026 advertised windows run from 128,000 tokens to 10 million, and every token in the window is billed and held in GPU memory.
Advertised context and usable context are not the same number
A published context length is a capacity limit, not a performance guarantee. NVIDIA’s RULER benchmark, which tests multi-needle retrieval, variable tracing and aggregation rather than a single needle in a haystack, shows models reliably use only 50–65% of the window they advertise. Accuracy on facts placed in the middle of a long prompt degrades well before the hard limit — the lost-in-the-middle effect — so a model sold with a 1M-token window should be planned around roughly 600–700K, and Llama 4 Scout’s 10M around 5–6.5M. The Effective Context column below is that adjusted figure.
What the window costs in memory and money
Context is not free on either side of the deployment line. On an API, every token in the window is re-billed on every turn, and several providers price beyond a threshold at a premium: GPT-5.4 doubles over 272K and Gemini 3.1 Pro doubles over 200K. On your own hardware the constraint is the KV cache, which stores the keys and values for every token already in the window and grows linearly with context length and with the number of concurrent users. The KV cache regularly exceeds the memory taken by the model weights on long-context serving, which is why the on-prem hardware sizing guide budgets it as a separate line item, and why prefix caching and FP8 or INT4 quantization are standard for long-context work. To model the token side, use the LLM token usage and cost guide.
| Model | Standard Context | Extended Context | Max Output | Effective Context* |
|---|---|---|---|---|
| Llama 4 Scout | 10M | — | — | ~5-6.5M |
| Grok 4 Fast | 2M | — | — | ~1.2-1.4M |
| Grok 4.20 Beta | 2M | — | — | ~1.2-1.4M |
| GPT-5.4 (Codex) | 272K | 1M | 128K | ~170K / ~600-650K |
| Claude Fable 5 | 1M | — | 128K | ~600-650K |
| Claude Opus 4.8 | 1M | — | 128K | ~600-650K |
| Gemini 3.1 Pro | 1M | 2M (beta) | 64K | ~600-700K |
| Gemini 3 Flash | 1M | — | 64K | ~600-700K |
| Claude Opus 4.6 | 200K | 1M (GA) | 64K | 130-200K / 600K-700K |
| Claude Sonnet 5 | 1M | — | 64K | ~600-650K |
| Claude Sonnet 4.6 | 200K | 1M (GA) | 64K | 130-200K / 600K-700K |
| GPT-5.2 / GPT-5 | 400K | — | 128K | ~240-280K |
| Grok 4.5 | 500K | — | — | ~300-325K |
| Grok 4 | 260K | — | — | ~160-170K |
| Qwen 3.5 (397B) | 262K | 1M | 65K | ~160-170K / ~600K |
| GLM-5.2 | 1M | — | — | ~600-650K |
| DeepSeek R1 / V3 | 128K | — | 64K / 8K | ~80-90K |
| Mistral Large 3 | 128K | — | — | ~80-90K |
What this means for retrieval design
When the corpus is larger than the usable window — which it almost always is — the answer is better retrieval, not a bigger window. Stuffing an entire document set into a long prompt raises cost on every turn, pushes the relevant passage into the region where accuracy falls off, and still fails once the corpus outgrows the model. A retrieval pipeline that returns the few passages that actually answer the question keeps the prompt short, cheap and accurate. That is the problem Blockify exists to solve: it distills source documents into compact, deduplicated blocks so the window carries answers instead of near-duplicate noise. For the retrieval architecture around it, see what a vector database is.
Turn This Research Into Real Strategy
Understanding LLM benchmarks is just the beginning. The AI Strategy Blueprint gives you the complete framework to evaluate, select, and deploy AI across your organization -- from model selection to ROI measurement.
- Model selection frameworks
- Build vs. buy decision guides
- ROI measurement templates
- Security & compliance checklists
- 78x accuracy improvement methodology
- Real-world case studies
Intelligence Level Taxonomy
4.1 Task Complexity Hierarchy
Understanding where your task falls on the complexity spectrum is the single most important factor in model selection. The hierarchy below moves from simplest to most complex:
Level 1: EXTRACTION -- Pull structured data from text Level 2: CLASSIFICATION -- Categorize inputs into predefined buckets Level 3: TRANSFORMATION -- Reformatting, translation, simple rewriting Level 4: SUMMARIZATION -- Condense information preserving key points Level 5: GENERATION -- Create new content following patterns Level 6: ANALYSIS -- Multi-factor reasoning about information Level 7: SYNTHESIS -- Combine information from multiple sources Level 8: MULTI-STEP REASONING -- Chain logical steps to reach conclusions Level 9: CREATIVE SYNTHESIS -- Novel solutions requiring insight + creativity Level 10: AGENTIC REASONING -- Autonomous multi-step tool use with planning
Detailed Level Descriptions
| Level | Name | Description | Example Tasks | Minimum Model Tier |
|---|---|---|---|---|
| 1 | Extraction | Pull specific fields, entities, or values from structured/semi-structured text | Name/email extraction, date parsing, regex-like tasks | Small (Haiku, GPT-5 Nano, Phi-4) |
| 2 | Classification | Assign inputs to one or more predefined categories | Sentiment analysis, topic tagging, intent detection, spam filtering | Small (Haiku, Flash-Lite, Phi-4) |
| 3 | Transformation | Convert content between formats or styles | JSON reformatting, language translation, tone adjustment, data normalization | Small-Mid (Haiku, Flash, DeepSeek V3) |
| 4 | Summarization | Condense longer content while preserving meaning and priority | Meeting notes, article summaries, report digests | Mid (Sonnet, GPT-5, Gemini Pro) |
| 5 | Generation | Create new content following specified patterns, tone, or constraints | Email drafting, product descriptions, template completion, simple code | Mid (Sonnet, GPT-5, Gemini Pro) |
| 6 | Analysis | Evaluate information considering multiple factors and perspectives | Market analysis, document review, data interpretation, code review | Mid-High (Sonnet, GPT-5.2, Gemini Pro) |
| 7 | Synthesis | Combine insights from disparate sources into coherent conclusions | Research synthesis, competitive intelligence, multi-document QA | High (Opus, GPT-5.2, Gemini 3.1 Pro) |
| 8 | Multi-Step Reasoning | Chain logical deductions across multiple steps | Math proofs, legal reasoning, complex debugging, strategic planning | High (Opus, o3, Gemini Deep Think) |
| 9 | Creative Synthesis | Generate novel solutions requiring both analytical and creative thinking | Architecture design, creative writing, novel algorithm design | Frontier (Opus, GPT-5.4, o3 Pro) |
| 10 | Agentic Reasoning | Plan, execute, and adapt multi-step workflows using tools autonomously | Autonomous coding agents, research agents, complex workflow automation | Frontier + Scaffolding (Opus, GPT-5.2 + tools) |
4.2 Capability Threshold Concept
The capability threshold is the minimum model intelligence required for a task to be completed reliably (>90% success rate). Below this threshold, the model fails unpredictably. Above it, upgrading provides diminishing returns.
Capability Threshold Visualization
Task Success Rate
100% | _______________
| ____/
90% | ____/ <-- Threshold: reliable above this line
| ___/
50% | ____/
| ___/
0% |_/
+----+----+----+----+----+----+----+----+-->
Nano Haiku Flash Sonnet GPT-5 Opus o3 o3Pro
<<<< Model Capability >>>>
Key insight: Once you pass the capability threshold, the cheapest model that clears it is the optimal choice. Spending more buys marginal quality improvements that rarely justify the cost.
Cost-Intelligence Sweet Spots
| Task Complexity | Threshold Model | Cost/M Output | Cost Multiplier vs. Next Tier |
|---|---|---|---|
| Level 1-2 (Extraction, Classification) | GPT-5 Nano / Haiku 4.5 | $0.40-$5.00 | 1x (baseline) |
| Level 3-4 (Transform, Summarize) | Gemini Flash / DeepSeek V3 | $0.28-$0.60 | 0.1-0.5x (cheaper!) |
| Level 5-6 (Generate, Analyze) | Sonnet 4.6 / GPT-5 | $10-$15 | 3-5x |
| Level 7-8 (Synthesize, Multi-step) | Opus 4.6 / GPT-5.2 | $14-$25 | 5-10x |
| Level 9-10 (Creative, Agentic) | Opus 4.6 + thinking / o3 | $25-$150+ | 10-50x |
4.3 When Small Models Are Sufficient
Small models (1B-14B parameters, or budget API tiers) are the right choice when many of the conditions below apply. If you need a refresher on what 1B vs 70B vs 1T parameters mean for cost, latency, and capability, our parameter-size guide breaks down the tradeoffs at each tier:
- Task is at Level 1-3 (extraction, classification, transformation)
- Input/output formats are well-defined and predictable
- High throughput (>100 req/sec) is required
- Latency budget is <500ms
- Cost per request must be <$0.001
- Data must remain on-premise (fine-tuned Phi-4, Qwen3-8B, Llama 3.3 8B)
- Task is domain-specific and model can be fine-tuned on domain data
- Output quality floor is more important than ceiling (consistency > brilliance)
Small Model Recommendations
| Use Case | Model | Why |
|---|---|---|
| Entity extraction at scale | Phi-4 (14B) | Runs on single GPU; fast; accurate for structured extraction |
| Classification / routing | Claude Haiku 4.5 | $1/$5; excellent instruction following |
| Simple formatting/transformation | GPT-5 Nano | $0.05/$0.40; massive context (400K) |
| On-premise sensitive data | Qwen3-8B / Llama 3.3 8B | Apache 2.0/Llama license; full data control |
| High-volume chat routing | Gemini Flash-Lite | $0.075/$0.30; extremely cheap |
4.4 When You Need Frontier Models
Frontier models (Opus, GPT-5.2+, Gemini 3.1 Pro, o3) are necessary when:
- Task requires multi-step reasoning (Level 7+)
- Ambiguous or incomplete inputs are common
- Creative or novel output is expected
- Code generation must handle complex real-world repositories
- Accuracy is safety-critical (medical, legal, financial)
- Long-document synthesis across 100K+ tokens is required
- Agentic workflows need autonomous planning and tool use
- Writing quality must be publication-grade
- Mathematical or scientific reasoning is involved
When NOT to use frontier models:
- Simple CRUD operations on data
- Template-based generation with minor variations
- Binary yes/no classification
- Data format conversion
- Any task that can be solved with regex + a lookup table
Task-to-Model Matching Framework
5.1 Task Category Matrix
| Task Category | Recommended Tier | Top Picks (Proprietary) | Top Picks (Open-Source) | Key Benchmark |
|---|---|---|---|---|
| Simple Extraction | Budget | Haiku 4.5, GPT-5 Nano | Phi-4, Qwen3-8B | IFEval |
| Text Classification | Budget | Haiku 4.5, Flash-Lite | Phi-4, Llama 3.3 8B | Custom eval |
| Translation | Mid | Sonnet 4.6, Gemini Pro | Qwen3-235B (201 langs), Mistral Large 3 | MMMLU |
| Summarization | Mid | Sonnet 4.6, GPT-5 | DeepSeek V3, Qwen3-30B | HELM |
| Content Generation | Mid-High | Claude Opus 4.6 (quality), GPT-5.4 (structure) | Llama 4 Maverick | Arena Elo (Creative) |
| Code Generation | High | Claude Opus 4.6, GPT-5.4, Gemini 3 Flash | MiniMax M2.5, MiniMax M2.7, Kimi K2.5, GLM-5/5.1 | SWE-bench |
| Code Review/Debug | High | Claude Opus 4.6, Sonnet 4.6 | MiniMax M2.5/M2.7, GLM-5/5.1, DeepSeek R1 | SWE-bench |
| Mathematical Reasoning | High-Frontier | o3, Gemini Deep Think | DeepSeek R1 | AIME 2025, MATH |
| Scientific Reasoning | Frontier | Gemini 3.1 Pro (94.3%), GPT-5.4 (92.0%), Claude Opus | Qwen 3.5 (88.4%) | GPQA Diamond |
| Document Processing/RAG | Mid-High | Gemini 3.1 Pro, Claude Opus 4.6 | Qwen3-30B (262K ctx) | RULER, MMLU-Pro |
| Creative Writing | High | Claude Opus 4.6 | Llama 4 Maverick | Arena Elo (Creative) |
| Tool Use / Function Calling | Mid-High | Claude Sonnet 4.6, GPT-5.4 | MiniMax M2.5 (BFCL 76.8%), Kimi K2.5 | BFCL v4 |
| Agentic Workflows | Frontier | Claude Opus 4.6, GPT-5.4 (OSWorld 75%) | MiniMax M2.5/M2.7, Kimi K2.5 (100 agents) | SWE-bench, BFCL v4 |
| Multimodal (Image) | Mid-High | Gemini 3 Flash, Gemini 3.1 Pro | Qwen 3.5, InternVL3-78B | MMMU Pro |
| Multimodal (Audio/Video) | Frontier | Gemini 3.1 Pro (only native option) | -- | -- |
| Customer Support Chatbot | Mid | Claude Sonnet 4.6, GPT-5 | Llama 4 Maverick | Arena Elo |
| Data Analysis | Mid-High | Claude Opus 4.6, GPT-5.2 | DeepSeek R1 | Custom eval |
| Legal Document Review | High | Claude Opus 4.6 (low hallucination) | Qwen3-235B | GPQA, IFEval |
| Medical Q&A | Specialized/Frontier | Med-Gemini, Claude Opus 4.6 | PMC-LLaMA, BioMistral | MedQA, PubMedQA |
5.2 Use Case Deep Dives
Best LLM for Coding in 2026
The coding landscape has distinct tiers (note: seven models now score within 2.8 points of each other on SWE-bench Verified). Current scores, tier by tier, are maintained on the best LLM for coding page:
- Agentic coding (fix real bugs in real repos): Claude Opus 4.6, Gemini 3.1 Pro, MiniMax M2.5, or GPT-5.4
- Everyday coding assistance: Claude Sonnet 4.6, Gemini 3 Flash, or GPT-5.4 -- excellent quality at lower cost
- Code completion/autocomplete: Smaller models fine -- Qwen3-Coder, DeepSeek-Coder, Step-3.5-Flash
- Open-source self-hosted: MiniMax M2.5 (now matching proprietary frontier), MiniMax M2.7, GLM-5, Kimi K2.5, Qwen 3.5
RAG / Document Processing
Best models for RAG need three capabilities: knowledge breadth (MMLU-Pro), reasoning ability (GPQA, BBH), and instruction following (IFEval).
Recommended setups:
- Best overall: Qwen 3.5 (256K context, native vision, Apache 2.0) or Qwen3-30B + Qwen3-Embedding-8B
- Maximum quality: Claude Opus 4.6 or Gemini 3.1 Pro (1M+ context)
- Budget: Gemini 3.1 Flash Lite ($0.25/$1.50) or Gemini 3 Flash ($0.50/$3.00) + RAG framework
- Resource-constrained: Phi-4 (14B, runs on consumer GPU)
Creative Writing
No automated benchmark reliably measures creative writing quality. Arena Elo (Creative Writing subcategory) and human evaluation are the only reliable signals.
Current ranking (subjective, based on practitioner reports):
- Claude Opus 4.6 -- best prose rhythm, subtext handling, consistent tone
- GPT-5.4 -- more structured, better at maintaining complex narrative frameworks
- Gemini 3.1 Pro -- capable but less literary; better for informational content
The LLM Selection Decision Process
How do you choose the right LLM?
Choose an LLM in five steps: (1) apply hard filters first — data privacy, deployment mode, context window, latency, and budget eliminate most of the 30+ candidates immediately; (2) classify your task's intelligence tier, from simple extraction to agentic reasoning; (3) read only the benchmarks that match that task; (4) shortlist 3–5 models and evaluate them on 100–200 examples from your real workload; (5) route production traffic across tiers rather than betting on one model. Most teams over-buy: the majority of production workloads run well on mid-tier models.
6.1 Decision Tree
The following decision tree guides you from initial task definition to a narrowed shortlist of candidate models. For multi-model deployments, pair this with our hybrid AI architecture framework to map workloads to the right combination of on-prem and cloud inference, and our analysis of edge AI vs cloud AI TCO to size the cost envelope.
Start: Define Your Task | +-- Data stays on-premise? | | | +-- YES --> GPU infrastructure available? | | | | | +-- Production GPUs --> Open-Source Self-Hosted | | | +-- High reasoning --> DeepSeek R1, Qwen 3.5, GLM-5/5.1 | | | +-- Medium tasks --> Llama 4 Maverick, MiniMax M2.5/M2.7 | | | +-- Low tasks --> Phi-4, Llama 3.3 8B, Qwen3-8B | | | | | +-- No / Limited --> Managed private cloud | | +-- Consumer GPU --> Phi-4, Qwen3-8B | | | +-- NO (Cloud API OK) --> Task Complexity Level? | | | +-- Level 1-3 (Simple) | | +-- High throughput? --> Haiku 4.5, GPT-5 Nano, Flash-Lite | | +-- Standard --> Haiku 4.5, Flash, DeepSeek V3 | | | +-- Level 4-6 (Medium) --> Primary task type? | | +-- Code --> Sonnet 4.6, Gemini 3 Flash, GPT-5.4 | | +-- Writing --> Sonnet 4.6, GPT-5 | | +-- Analysis --> Gemini Pro, Sonnet 4.6 | | +-- Multilingual--> Gemini Pro, Qwen3 | | +-- Multimodal --> Gemini Flash/Pro | | | +-- Level 7-8 (Complex) --> Budget? | | +-- Strict --> DeepSeek R1, Gemini 2.5 Pro | | +-- Moderate --> Opus 4.6, GPT-5.4 | | +-- No limit --> Test all frontier models | | | +-- Level 9-10 (Frontier) --> Task type? | +-- Coding agents --> Claude Opus 4.6 + Claude Code | +-- Math/Science --> o3 / Gemini Deep Think | +-- Creative --> Claude Opus 4.6 | +-- Computer use --> GPT-5.4 / Kimi K2.5 | +-- Multimodal --> Gemini 3.1 Pro | +-- Max reasoning --> o3 Pro | +--> Proceed to Evaluation Phase
6.2 Step 1: Define Requirements
Before looking at any model, document the following:
Task Requirements Worksheet
- TASK DESCRIPTION
- What specifically does the model need to do?
- What are example inputs and expected outputs?
- Task complexity level (1-10 from taxonomy above)
- QUALITY REQUIREMENTS
- Minimum acceptable accuracy: ___%
- Tolerance for hallucination: None / Low / Medium
- Output format: Free text / Structured JSON / Code / Mixed
- Consistency requirement: Every response identical / Mostly similar / Creative variation OK
- VOLUME & PERFORMANCE
- Expected requests per day: ___
- Peak requests per second: ___
- Maximum acceptable latency (TTFT): ___ ms
- Maximum acceptable total response time: ___ s
- CONTEXT REQUIREMENTS
- Typical input length: ___ tokens
- Maximum input length: ___ tokens
- Required output length: ___ tokens
- Need for long-context retrieval: Yes / No
- DATA & COMPLIANCE
- Data sensitivity: Public / Internal / Confidential / Regulated
- Can data leave your infrastructure? Yes / No
- Industry-specific regulations
- Geographic data residency requirements
- BUDGET
- Maximum monthly spend: $___
- Maximum cost per request: $___
- Infrastructure budget (if self-hosting): $___/month
- INTEGRATION
- Deployment mode: API / Self-hosted / Hybrid
- Tool/function calling needed: Yes / No
- Streaming required: Yes / No
- Multimodal inputs needed: Text only / +Images / +Audio / +Video
- Languages required
6.3 Step 2: Apply Hard Filters
Hard filters are binary -- models either pass or fail. Apply these to immediately eliminate unsuitable candidates.
| Filter | Eliminates If... | Example |
|---|---|---|
| Data Privacy | Provider cannot meet your data handling requirements | HIPAA data eliminates most API providers without BAA |
| Deployment Mode | Model is API-only but you need on-premise | Eliminates all proprietary if self-hosting mandatory |
| Context Window | Effective context is less than your maximum input | 128K models eliminated if processing 200K+ documents |
| Licensing | License prohibits your use case | Some models restrict commercial use or require attribution |
| Language Support | Does not support your required languages | Many models weak on low-resource languages |
| Multimodal | Lacks required input modalities | Only Gemini 3.1 Pro supports native audio+video |
| Regulatory | Provider does not meet compliance standards | SOC 2, GDPR, HIPAA, FedRAMP requirements |
| Geography | Model/API not available in your region | Some providers have geographic restrictions |
After hard filters, you should have 10-20 candidates remaining.
6.4 Step 3: Score Against Soft Criteria
Score remaining candidates 1-5 on each criterion, weighted by importance to your use case:
| Criterion | Weight (adjust per use case) | Scoring Guidance |
|---|---|---|
| Benchmark Performance | 20-30% | Use relevant benchmarks from Section 2.4 |
| Cost Efficiency | 15-25% | Cost per 1M tokens relative to quality tier |
| Latency / Throughput | 10-20% | TTFT and tokens/sec relative to your requirements |
| Context Window | 5-15% | Effective context vs. your needs (50-65% rule) |
| Provider Reliability | 10-15% | Uptime SLA, rate limits, support quality |
| Ecosystem / Tooling | 5-15% | SDK quality, documentation, community |
| Safety / Alignment | 5-15% | Hallucination rate, content filtering, refusal patterns |
| Customization | 0-10% | Fine-tuning availability, system prompt flexibility |
Example Scoring Matrix
| Model | Benchmark (25%) | Cost (20%) | Latency (15%) | Context (10%) | Reliability (15%) | Ecosystem (10%) | Safety (5%) | Weighted Score |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.6 | 5 | 2 | 3 | 4 | 5 | 4 | 5 | 3.80 |
| GPT-5.2 | 5 | 3 | 3 | 3 | 5 | 5 | 4 | 3.95 |
| Gemini 3.1 Pro | 5 | 4 | 4 | 5 | 4 | 4 | 4 | 4.25 |
| DeepSeek R1 | 4 | 5 | 3 | 2 | 3 | 3 | 3 | 3.35 |
| Claude Sonnet 4.6 | 4 | 3 | 4 | 4 | 5 | 4 | 5 | 3.95 |
6.5 Step 4: Narrow to Top 5-10
After scoring, select your shortlist:
- Top 2-3 from your scoring matrix (highest weighted scores)
- 1-2 "wild card" models that might outperform on YOUR specific task despite lower general scores
- 1 budget option to establish a cost-efficiency baseline
- 1 open-source option (if applicable) for comparison and fallback strategy
Your shortlist should be 5-8 models maximum. More than this creates evaluation burden without proportional benefit.
Example Shortlist for a Coding Assistant Use Case
| # | Model | Rationale |
|---|---|---|
| 1 | Claude Opus 4.6 | Highest coding Arena Elo (1548); SWE-bench leader (80.8%) |
| 2 | GPT-5.4 | Strong SWE-bench (~80%); native computer use; 1.05M context in Codex |
| 3 | Gemini 3.1 Pro / 3 Flash | SWE-bench 80.6% (Pro) / 78% (Flash at $0.50/$3); best context; strong value |
| 4 | Claude Sonnet 4.6 | 98% of Opus coding quality at 60% cost (79.6% SWE-bench) |
| 5 | MiniMax M2.5 | Open-source SWE-bench leader (80.2%); Modified-MIT |
| 6 | GLM-5 | Open-source; SWE-bench 77.8%; top Arena Elo among open models |
Get Chapter 1 Free + AI Academy Access
Go deeper into AI strategy with a free chapter from The AI Strategy Blueprint, plus instant access to Iternal's AI Academy with frameworks, templates, and implementation guides.
- Free Chapter 1: "The AI Strategy Imperative"
- AI Academy access with video courses
- Model selection decision templates
- ROI calculator spreadsheets
- Weekly AI strategy newsletter
Get Your Free Chapter
Model Routing & Cascade Strategies
What is model routing?
Model routing sends each request to the cheapest model that can handle it, instead of sending everything to one frontier model. A typical cascade routes simple tasks to Haiku-class models ($1/$5 per million tokens), medium tasks to Sonnet-class ($3/$15), and only the hardest 15% to Opus-class ($5/$25) — a blended cost near $10.50 per million output tokens versus $25 for frontier-only, a 58% saving with equal quality where it matters. Routing is the default enterprise architecture in 2026.
Why Route?
The most effective AI architecture in 2026 does not rely on a single model. Instead, it routes different requests to different models based on what the task actually needs. Research shows well-designed routing systems can outperform even the strongest individual models while reducing costs 50-80%. Routing rules, caching and spend limits are normally enforced once at the AI gateway layer rather than re-implemented in every application.
Routing Strategies
| Strategy | How It Works | Best For | Complexity |
|---|---|---|---|
| Static Routing | Predefined rules map task types to models | Predictable workloads with clear task categories | Low |
| Difficulty-Based Routing | Lightweight classifier estimates task difficulty, routes to appropriately-sized model | Mixed-difficulty workloads | Medium |
| Cascade | Start with cheapest model; escalate to larger model if confidence is low | Cost optimization with quality guarantee | Medium |
| Cascade Routing | Unified framework: iteratively picks best model, can skip/reorder | Maximum efficiency (up to 14% better) | High |
| RL Routing | Router learns optimal model assignment from feedback | Large-scale production with feedback loops | High |
Practical Cascade Architecture
User Request
|
v
[Classifier / Router] -- Estimates task complexity
|
|-- Simple (Level 1-3) --> Haiku 4.5 / GPT-5 Nano ($0.05-$1.00/M)
| |
| v
| [Confidence Check]
| |
| >= 0.9 --> Return Response
| < 0.9 --> Escalate to Mid-Tier
|
|-- Medium (Level 4-6) --> Sonnet 4.6 / GPT-5 ($3-$10/M)
| |
| v
| [Confidence Check]
| |
| >= 0.85 --> Return Response
| < 0.85 --> Escalate to Frontier
|
|-- Complex (Level 7+) --> Opus 4.6 / GPT-5.2 ($5-$25/M)
| |
| v
| Return Response
|
|-- Reasoning Required --> o3 / Gemini Deep Think
Cost Savings From Routing
Assuming a typical enterprise workload distribution:
| Task Complexity | % of Requests | Model Used | Cost/M Output | Blended Contribution |
|---|---|---|---|---|
| Simple (Level 1-3) | 60% | Haiku 4.5 | $5.00 | $3.00 |
| Medium (Level 4-6) | 25% | Sonnet 4.6 | $15.00 | $3.75 |
| Complex (Level 7+) | 15% | Opus 4.6 | $25.00 | $3.75 |
| Blended average | 100% | -- | -- | $10.50/M |
Compared to using Opus for everything ($25.00/M), routing saves 58% in cost while maintaining quality where it matters.
Evaluation Methodology After Shortlisting
Shortlist evaluation runs on your own workload, so pair the dataset design below with the LLM evaluation metrics that decide whether a candidate model passes.
8.1 Designing Evaluation Datasets
Dataset composition targets:
| Category | % of Dataset | Purpose |
|---|---|---|
| Happy Path | 50-60% | Common, expected inputs that represent your typical workload |
| Edge Cases | 20-30% | Atypical, ambiguous, complex inputs that test boundaries |
| Adversarial | 10-15% | Malicious or tricky inputs that test safety and error handling |
| Regression | 5-10% | Known-difficult examples from production failures |
Dataset curation strategies:
- Manual curation (highest quality): Subject matter experts create 50-200 test cases aligned with product goals. Include high-priority workflows, known failure modes, and edge cases.
- Production sampling: Pull real prompts and responses from production logs. Provides grounded, real-world data. Best for identifying drift and tracking quality over time.
- Synthetic generation: Use a strong LLM to generate test cases automatically. Fast but requires human review. Best for scaling coverage after manual curation establishes the pattern.
- Gold Standard Questions (GSQs): Labeled dataset with expert-verified ground truth answers. Most reliable for automated scoring but expensive to create.
Minimum viable evaluation set: 100-200 examples covering all categories above. For high-stakes applications, aim for 500+.
8.2 LLM-as-Judge Approaches
LLM-as-Judge uses a strong model (typically Claude Opus or GPT-5) to evaluate outputs from candidate models.
For each test case: 1. Send input to candidate model --> get response 2. Send (input, response, rubric) to judge model --> get score + reasoning 3. Aggregate scores across all test cases
Best practices:
| Practice | Why |
|---|---|
| Use a model at least as capable as candidates | Weaker judges cannot reliably evaluate stronger models |
| Define explicit rubrics with 1-5 scoring criteria | Vague instructions produce inconsistent scores |
| Include "reasoning" field in judge output | Enables auditing of judge decisions |
| Test for judge bias (position, verbosity) | Judges may prefer first response or longer responses |
| Calibrate with human agreement rate | Target >80% agreement between judge and human experts |
| Version control your judge prompts | Judge behavior changes with prompt changes |
| Use multiple judge models to reduce bias | Average scores from 2-3 different judge models |
Key frameworks:
- DeepEval -- 50+ research-backed metrics including G-Eval, hallucination detection, answer relevancy, task completion
- Langfuse -- LLM-as-judge integration with production tracing
- Arize Phoenix -- Open-source evaluation with hallucination-specific judges
- Amazon Bedrock Model Evaluation -- Managed LLM-as-judge on AWS
8.3 A/B Testing Frameworks
[Offline Evaluation] [Online A/B Testing]
| |
| Identify promising | Validate with
| candidates on static dataset | real users
| |
v v
Top 2-3 candidates -------> Deploy to % of traffic
|
Measure: completion rate,
user satisfaction, task success
|
Feed challenging examples
back into offline eval dataset
|
[Continuous Improvement Loop]
A/B testing checklist:
- Define primary success metric before starting (e.g., task completion rate)
- Calculate required sample size for statistical significance
- Randomize user assignment to prevent selection bias
- Run for minimum 2 weeks to capture variance
- Control for confounders (time of day, user type, input complexity)
- Measure secondary metrics (latency, cost, user satisfaction)
- Document all prompt versions used with each model
8.4 Automated Evaluation Pipelines
[Test Dataset]
|
v
[Evaluation Runner] -- Sends inputs to all candidate models in parallel
|
v
[Response Collector] -- Stores all (input, model, response) tuples
|
v
[Metric Calculator]
|
|-- Deterministic Metrics: exact match, regex, JSON schema validation
|-- Statistical Metrics: BLEU, ROUGE, BERTScore
|-- LLM-as-Judge Metrics: quality, relevance, hallucination
|-- Latency Metrics: TTFT, total time, tokens/second
|-- Cost Metrics: actual cost per request
|
v
[Dashboard / Report Generator]
|
v
[CI/CD Integration] -- Block deployment if metrics drop below threshold
Tools for automated evaluation:
| Tool | Type | Key Feature | Cost |
|---|---|---|---|
| DeepEval | Open-source framework | 50+ metrics, CI/CD integration, LLM-as-judge | Free |
| Langfuse | Open-source observability | Production tracing + evaluation | Free / managed |
| Braintrust | Commercial platform | Eval + prompt management + logging | Paid |
| Promptfoo | Open-source CLI | Fast model comparison, CI-friendly | Free |
| Arize Phoenix | Open-source platform | Hallucination detection, tracing | Free |
| Weights & Biases | Commercial platform | Experiment tracking + eval | Free tier |
| HELM | Academic framework | 7 metrics across 42 scenarios | Free |
8.5 Metrics Beyond Accuracy
| Metric | What It Measures | Why It Matters | How to Measure |
|---|---|---|---|
| Task Success Rate | % of outputs that fully complete the intended task | The most business-relevant metric | Human evaluation or automated checks |
| Hallucination Rate | % of outputs containing fabricated facts | Trust and liability | LLM-as-judge + spot-check |
| Latency (TTFT) | Time to first token | User experience in interactive apps | API timing |
| Latency (Total) | Time to complete full response | End-to-end user experience | API timing |
| Throughput | Requests handled per second | Scalability and capacity planning | Load testing |
| Cost Per Request | Average $ per API call | Budget planning and ROI | Provider billing |
| Cost Per Successful Request | $ per request that actually succeeds | True cost of quality | Cost / success rate |
| Instruction Adherence | % of constraints/instructions followed | Reliability for structured output | IFEval-style checks |
| Consistency | Variance in output quality across runs | Predictability | Multiple runs on same inputs |
| Safety / Refusal Rate | % of harmful requests correctly refused AND safe requests incorrectly refused | Safety vs. usability balance | Red-team testing |
| Format Compliance | % of outputs matching required format | Integration reliability | Schema validation |
Key Leaderboards & Resources
These leaderboards rank general capability. If the question is which system answers correctly on your own company documents rather than which model is strongest overall, see most powerful AI vs most accurate on your own data.
| Resource | What It Provides | Update Frequency |
|---|---|---|
| Chatbot Arena (LMSYS) | Human preference Elo ratings; most ecologically valid | Daily |
| Vellum LLM Leaderboard | Multi-benchmark comparison with scores | Weekly |
| Open LLM Leaderboard | Open-source model rankings (HuggingFace) | Continuous |
| HELM (Stanford) | Holistic 7-metric evaluation across 42 scenarios | Periodic |
| LLM-Stats | Comprehensive benchmark aggregation | Daily |
| Berkeley Function Calling | Tool/function calling evaluation (BFCL v4) | Regular |
| Artificial Analysis | Performance, latency, and pricing comparison | Continuous |
| Price Per Token | Pricing comparison across 300+ models | Daily |
| SWE-bench Leaderboard | Coding/engineering model rankings | Regular |
| Onyx AI Leaderboards | Task-specific leaderboards (coding, RAG, self-hosted) | Weekly |
Open-Source vs. Proprietary Decision Guide
Should you use an open-source or a proprietary LLM?
Use open-source models when data cannot leave your infrastructure (HIPAA, GDPR, air-gapped environments), when volume exceeds roughly one million tokens per day, or when you must fine-tune on proprietary data. Use proprietary APIs when you need maximum quality, ship fast without ML infrastructure, or run low-to-moderate volume. The performance gap has nearly closed for coding — open models now score within a few points of frontier APIs on SWE-bench — so most enterprises in 2026 run a hybrid: self-hosted open weights for sensitive data, frontier APIs for the hardest work.
Decision Matrix
| Factor | Open-Source Advantage | Proprietary Advantage |
|---|---|---|
| Data Privacy | Full control; data never leaves your infrastructure | BAAs available but data goes to provider |
| Customization | Fine-tune freely; modify architecture | Limited to prompt engineering + some fine-tuning |
| Cost at Scale | Fixed infrastructure cost; no per-token fees | No infra management; pay-per-use |
| Cost at Low Volume | High fixed cost (GPUs) regardless of usage | Pay only for what you use |
| Performance Ceiling | Narrowing gap, but still below frontier proprietary | Highest absolute performance (Opus, GPT-5.2, Gemini 3.1 Pro) |
| Deployment Speed | Days-weeks for infrastructure setup | Minutes via API |
| Reliability | You manage uptime | Provider SLAs (typically 99.9%+) |
| Provider Lock-in | None -- switch models freely | Moderate -- prompt engineering is provider-specific |
| Regulatory | Full audit trail; compliance control | Varies by provider and region |
| Support | Community + paid options | Enterprise support included |
When to Choose Open-Source
- Regulated industries requiring full data sovereignty (HIPAA, GDPR strict interpretation)
- High-volume workloads where per-token costs exceed infrastructure costs
- Need to fine-tune on proprietary domain data (though for most knowledge-injection problems, RAG vs fine-tuning tips toward RAG)
- Competitive advantage requires model customization
- Geographic/sovereignty restrictions prevent using US-based APIs
- Budget: typically economical above ~1M tokens/day
When to Choose Proprietary
- Need maximum absolute quality (top-tier coding, reasoning, or creative tasks)
- Low-to-moderate volume (<500K tokens/day)
- Fast prototyping and iteration
- No ML infrastructure team
- Need native multimodal support (especially audio/video -- Gemini only)
- Enterprise support and SLAs are required
Hybrid Strategy (Recommended for Most Enterprises)
Most teams in 2026 mix models:
- Self-hosted open-weight model for sensitive data processing (Qwen3-30B, DeepSeek V3)
- Cheap API model for high-volume routine tasks (Gemini Flash, GPT-5 Nano)
- Frontier API model for the hardest 15% of work (Opus 4.6, GPT-5.2)
Whichever mix you land on, the connection step is the same: Iternal platforms let you bring your own foundation model, whether it is self-hosted open weights or a provider API.
Quick Reference: Model Recommendations by Use Case
Tier 1: Best Overall (No Budget Constraints)
| Use Case | #1 Pick | #2 Pick | #3 Pick |
|---|---|---|---|
| Coding Agent | Claude Opus 4.6 | Gemini 3.1 Pro / Gemini 3 Flash (value) | GPT-5.4 |
| Creative Writing | Claude Opus 4.6 | GPT-5.4 | Claude Sonnet 4.6 |
| Math/Science Reasoning | o3 | Gemini Deep Think | Gemini 3.1 Pro (94.3%) |
| Abstract Reasoning | GPT-5.4 Pro (83.3%) | Gemini 3.1 Pro (77.1%) | GPT-5.4 Standard (73.3%) |
| General Knowledge Q&A | Gemini 3.1 Pro | Claude Opus 4.6 | GPT-5.4 |
| Document Processing | Gemini 3.1 Pro | Claude Opus 4.6 | Qwen 3.5 |
| Multimodal (Image) | Gemini 3.1 Pro | Gemini 3 Flash (MMMU Pro 81.2%) | GPT-5.4 |
| Multimodal (Audio/Video) | Gemini 3.1 Pro | -- | -- |
| Tool Use / Agentic | Claude Opus 4.6 | GPT-5.4 (computer use) | MiniMax M2.5 |
| Customer Support | Claude Sonnet 4.6 | GPT-5 | Gemini Pro |
Tier 2: Best Value (Cost-Optimized)
| Use Case | #1 Pick | #2 Pick | #3 Pick |
|---|---|---|---|
| Coding | Claude Sonnet 4.6 | Gemini 3 Flash (78% at $0.50/$3) | GPT-5.4 Mini |
| Writing | Claude Sonnet 4.6 | GPT-5 | Llama 4 Maverick |
| Reasoning | DeepSeek R1 | Gemini 2.5 Pro | GPT-5 |
| Classification | Haiku 4.5 | GPT-5 Nano | Gemini 3.1 Flash Lite |
| Extraction | GPT-5 Nano | Haiku 4.5 | Phi-4 |
| Translation | Gemini 3 Flash | Qwen 3.5 (201 langs) | Mistral Large 3 |
| Summarization | DeepSeek V3.2 | Gemini 3.1 Flash Lite | GPT-5 |
| RAG | Qwen 3.5 | Gemini 3.1 Flash Lite | DeepSeek V3.2 |
Tier 3: Best Self-Hosted (On-Premise)
| Use Case | #1 Pick | #2 Pick | #3 Pick |
|---|---|---|---|
| General Purpose | Qwen 3.5 (397B) | Llama 4 Maverick | DeepSeek V3.2 |
| Coding | MiniMax M2.5 (80.2%) | GLM-5/5.1 (94% of Opus) | Kimi K2.5 (76.8%) |
| Reasoning | DeepSeek R1 | Qwen 3.5 | GLM-4.7 (HLE 42.8%) |
| Small/Edge | Step-3.5-Flash (11B active) | Phi-4 (14B) | Qwen3-8B |
| Multilingual | Qwen 3.5 (201 langs) | Mistral Large 3 (80+) | Llama 4 Maverick |
| Multimodal | Qwen 3.5 (native vision) | InternVL3-78B | GLM-4.5V |
Need Help Selecting & Deploying the Right AI Models?
Iternal's AI strategy consultants help enterprises navigate model selection, build evaluation frameworks, and deploy production AI systems with measurable ROI.
AI Masterclass
AI Strategy Sprint
Transformation Program
Founder's Circle
Frequently Asked Questions
There is no single "best" LLM. The right model depends on your task, budget, latency, and data-privacy constraints. As of July 2026, Claude Fable 5 leads frontier coding and HLE, Gemini 3.1 Pro dominates scientific reasoning and multimodal tasks, GPT-5.4 excels at computer use and structured reasoning, and open-weight models like MiniMax M2.5 rival proprietary coding scores. For most organizations, a routing strategy that sends different tasks to different models beats any single-model bet.
Choose open-source when you need full data sovereignty (HIPAA/GDPR), process more than ~1M tokens per day, need to fine-tune on proprietary data, or have geographic restrictions. Choose proprietary when you need maximum quality, have low-to-moderate volume, want fast prototyping, lack ML infrastructure, or need native multimodal capabilities. Most enterprises benefit from a hybrid approach that combines both.
Benchmark scores are necessary starting points but insufficient on their own. Major concerns include: saturation (MMLU, HumanEval, GSM8K are no longer differentiating), data contamination (training data may include benchmark questions), and scaffold dependence (SWE-bench scores vary significantly with different evaluation frameworks). Always supplement benchmarks with your own domain-specific evaluation using 100-200 test cases that represent your actual workload.
Model routing sends different requests to different models based on task complexity, latency requirements, and cost constraints. A well-designed routing system can reduce costs by 50-80% while maintaining quality. For example, sending simple classification tasks to Haiku ($1/$5), medium-complexity tasks to Sonnet ($3/$15), and only the hardest tasks to Opus ($5/$25) produces a blended cost of ~$10.50/M output tokens vs. $25/M for Opus across the board -- a 58% savings.
The gap between open-source and proprietary has effectively closed for coding tasks. MiniMax M2.5 achieves 80.2% on SWE-bench Verified, matching Claude Opus 4.6 (80.8%). New entrants like GLM-5/5.1, Kimi K2.5, and Qwen 3.5 rival frontier proprietary models across most benchmarks. MIT/Apache 2.0 licensed models now offer cost per token 10-100x cheaper than proprietary APIs, making self-hosted deployments increasingly attractive for high-volume workloads.
MMLU is saturated (88-94% for top models) and no longer differentiates frontier models. Instead, use: GPQA Diamond for scientific reasoning, SWE-bench Verified or SWE-bench Pro for coding, AIME 2025 for mathematical reasoning, ARC-AGI 2 for abstract reasoning, Humanity's Last Exam (HLE) for the hardest reasoning tasks, BFCL v4 for tool/function calling, and Arena Elo from LMSYS for overall human preference. Choose benchmarks that align with your specific use case.
NVIDIA's RULER benchmark shows models reliably use only 50-65% of their advertised context window. A model with a 1M token context may only perform well up to 600-700K tokens. Llama 4 Scout advertises 10M but effectively uses ~5-6.5M. Always test with your actual document sizes and verify retrieval accuracy at the context lengths you need. Performance degrades significantly beyond the effective context threshold.
Start by eliminating models that fail your hard constraints -- privacy, deployment mode, context window, latency, and budget. Then compare the survivors on benchmarks matched to your task: SWE-bench Verified for coding, GPQA Diamond for scientific reasoning, AIME 2025 for math, BFCL v4 for tool use, and Arena Elo for overall human preference. Treat scores within 2-3% as equivalent, and confirm finalists on 100-200 examples from your own workload before committing.
The active differentiators in 2026 are GPQA Diamond (PhD-level science), SWE-bench Verified and SWE-bench Pro (real-world coding), AIME 2025 (competition math), ARC-AGI 2 (abstract reasoning), Humanity's Last Exam (hardest expert questions -- still under 51% for every model), BFCL v4 (tool calling), and LMSYS Arena Elo (human preference). MMLU, GSM8K, and HumanEval are saturated -- top models cluster above 88% -- so they confirm a capability floor but cannot rank frontier models. The full definitions and score tables are kept in the LLM benchmark repository.
Claude Fable 5 posts the highest SWE-bench Verified score in our July 2026 table (95.0%), with Claude Opus 4.8 (88.6%) and Claude Sonnet 5 (85.2%) behind it; Claude Opus 4.6, Gemini 3.1 Pro, MiniMax M2.5, and GPT-5.4 all sit near 80%. Among open-weight models, MiniMax M2.5 (80.2%) and GLM-5 (77.8%) are the strongest self-hosted picks. Scores depend heavily on the agentic scaffold, so verify on your own repositories before standardizing.
The context window is the maximum number of tokens a model can hold in working memory for one request, covering the system prompt, the conversation history, any retrieved documents and the response being generated. When the window fills, the oldest content is dropped. Advertised windows in 2026 range from 128K tokens to 10M, but NVIDIA’s RULER benchmark shows models reliably use only 50-65% of what they advertise, so plan against the effective figure rather than the headline number.
Smaller than most teams assume. A retrieval pipeline should hand the model the few passages that answer the question, not the corpus, and 8K to 32K tokens of retrieved context covers most enterprise question answering. Long windows are a fallback for documents that cannot be split, not a retrieval strategy: stuffing the window raises cost on every turn and pushes the relevant passage into the range where accuracy falls off. Iternal’s Blockify distills source documents into compact, deduplicated blocks so the window carries answers instead of near-duplicate noise.
There is no single winner, but four open-weight models cover most enterprise deployments. GLM-5.2 (MIT, 753B, GPQA Diamond 91.2%, 1M context) is the strongest all-rounder; MiniMax M2.5 matches proprietary frontier coding at 80.2% on SWE-bench Verified; Qwen 3.5 (Apache 2.0) leads multilingual and native-vision work across 201 languages; and Phi-4 (14B, MIT) runs on a single RTX 4090. Confirm the license terms and the hosting jurisdiction before you standardize, and size the VRAM before you commit to a tier — the full tiered table is in the open-weight model comparison.
As of the September 2026 review the most recent releases are GPT-5.6 (July 9, 2026, Terminal-Bench 2.1 88.8%), Claude Sonnet 5 (June 30, 2026, SWE-bench Verified 85.2%) and GLM-5.2 (June 13, 2026, MIT-licensed, 1M context, ARC-AGI-2 22.8%), alongside Claude Opus 4.8, Claude Fable 5 and Grok 4.5. Several 2026 launches published only agentic and terminal evaluations, so classic academic benchmarks are unavailable for direct comparison on those models. The dated entries are in the release log, and raw scores refresh daily in the LLM benchmark leaderboard.
Sources
Leaderboards & Benchmarks
Benchmark Explainers & Guides
Model Selection & Decision Frameworks
Model Comparisons
New Model Release Sources
Pricing & Performance
Open-Source vs. Proprietary
Evaluation & Testing
Related Resources
Explore our calculators, assessments, and guides to apply these insights to your organization. For the job-by-job version, see Choosing a Local Model: What Is Supported and Can You Bring Your Own.