I now have comprehensive current data across all major providers. Let me compile the structured markdown summary.
A few important caveats I'll flag: much of this non-Anthropic data comes from third-party aggregator/blog sources (not official docs), the model landscape is churning monthly, and some "facts" (Llama 5 specs especially) look suspiciously inflated from unreliable sources. The Claude family data is from the authoritative skill reference.
Here is my structured summary:
2026 LLM Landscape — Practical Reference (mid-2026)
Compiled: 2026-07-23. This field moves monthly; treat every version number and price as perishable. The Anthropic/Claude data is from the authoritative first-party skill reference; all other providers rely on third-party aggregators and blogs (official docs were not directly fetched), so their specifics are lower-confidence — see VOLATILE CLAIMS at the end.
Quick comparison table
| Provider | Current flagship (mid-2026) | Model ID / name | Context | Pricing tier ($/1M in-out) | Relative strength |
|---|---|---|---|---|---|
| Anthropic | Claude Opus 4.8 | claude-opus-4-8 |
1M | Premium ($5 / $25) | Agentic coding, long-horizon autonomy, knowledge work |
| Anthropic (most capable) | Claude Fable 5 | claude-fable-5 |
1M | Ultra-premium ($10 / $50) | Hardest reasoning / long-horizon agentic |
| Anthropic (balanced) | Claude Sonnet 5 | claude-sonnet-5 |
1M | Mid ($3 / $15; intro $2 / $10) | Near-Opus quality at Sonnet cost |
| Anthropic (cheap/fast) | Claude Haiku 4.5 | claude-haiku-4-5 |
200K | Cheap ($1 / $5) | Speed-critical simple tasks |
| OpenAI | GPT-5.6 Sol | gpt-5.6 (Sol tier) |
~1.05M | Premium ($5 / $30) | Frontier reasoning, computer-use, tool-heavy |
| OpenAI (balanced) | GPT-5.6 Terra | Terra tier | ~1.05M | Mid ($2.50 / $15) | GPT-5.5-class quality, half the price |
| OpenAI (cheap) | GPT-5.6 Luna | Luna tier | ~1.05M | Cheap ($1 / $6) | Summarize/classify/draft at volume |
| Gemini 3.1 Pro | gemini-3.1-pro |
1M | Mid ($2 / $12; long-ctx surcharge >200K) | Multimodal reasoning, agentic, computer use | |
| Google (cheap/fast) | Gemini 3.5 Flash | gemini-3.5-flash |
1M | Cheap | Fast; reportedly beats prior flagship on coding |
| DeepSeek (open) | DeepSeek V4 Pro | deepseek-ai/DeepSeek-V4-Pro |
1M | Very cheap ($0.44 / $0.87) | Price-performance frontier, MIT-licensed |
| DeepSeek (open, small) | DeepSeek V4 Flash | DeepSeek-V4-Flash |
— | Ultra-cheap ($0.14 / $0.28) | Cheap open-weight workhorse |
| Alibaba Qwen | Qwen 3.7 Max | qwen3.7-max |
1M | Cheap-mid ($2.50 / $7.50) | Agentic coding; closed-weights at launch |
| Mistral (open) | Mistral Large 3 | mistral-large-2512 |
256K | Open-weight (self-host) | Apache 2.0, multilingual, function calling |
| Meta (open) | Llama 5 | — | (see caveat) | Open-weight | General open flagship — specs unverified |
By provider
Anthropic — Claude (HIGH confidence: first-party reference)
- Claude Fable 5 (
claude-fable-5) — most capable widely released model. 1M context, 128K output. $10/$50. Thinking always on; refusal-fallback handling; requires 30-day data retention. Ultra-premium. - Claude Opus 4.8 (
claude-opus-4-8) — the flagship most people should default to. 1M context, 128K output. $5/$25. State-of-the-art long-horizon agentic + coding + knowledge work. Premium. - Claude Opus 4.7 / 4.6 (
claude-opus-4-7,claude-opus-4-6) — previous-gen Opus, still active. - Claude Sonnet 5 (
claude-sonnet-5) — best speed/intelligence balance; near-Opus on coding/agentic. 1M context. $3/$15 ($2/$10 intro through 2026-08-31). Mid. - Claude Haiku 4.5 (
claude-haiku-4-5) — fastest/cheapest. 200K context, 64K output. $1/$5. Cheap.
OpenAI — GPT-5.6 family (MEDIUM confidence: aggregators + OpenAI blog)
Released July 9, 2026 as three tiers with celestial names. ~1.05M context, 128K output. Requests above 272K input tokens are billed at a higher long-context rate (~$10/$45 for Sol).
- Sol (flagship) — $5/$30. Frontier reasoning, long-horizon agentic. Premium.
- Terra (balanced) — $2.50/$15. GPT-5.5-class at half the price; the production migration target. Mid.
- Luna (cheap/fast) — $1/$6. High-volume simple work. Cheap.
- Older gpt-5.5 ($5/$30), gpt-5.4 family, and gpt-5.1 ($1.25/$10) remain on the price sheet.
Google — Gemini (MEDIUM confidence)
- Gemini 3.1 Pro (
gemini-3.1-pro) — current GA flagship as of July 2026. 1M context. $2/$12 input/output; input roughly doubles above 200K-token prompts. Leads multimodal reasoning, agentic, computer use. Mid tier by price. - Gemini 3.5 Flash — live since May 2026; cheap/fast tier that reportedly beats the previous flagship on coding.
- Gemini 3.5 Pro — announced (Google I/O, May 2026) with a 2M-token context window and "Deep Think" mode, but NOT generally available as of early-mid July 2026 (slipped past June/July; August window rumored). Do not treat as shipping.
DeepSeek — V4 (open-weights) (MEDIUM confidence)
Shipped April 24, 2026. Two-tier open release, MIT-licensed weights on Hugging Face. - V4 Pro — 1.6T total / ~49B active MoE, 1M context. $0.44/$0.87 (cache-hit input ~$0.0036). Best open-weight price-performance. - V4 Flash — 284B total / ~13B active. $0.14/$0.28. Ultra-cheap. - R2 (reasoning model) — persistently rumored, not released; V4 is the real 2026 flagship.
Alibaba — Qwen (MEDIUM confidence)
- Qwen 3.7 Max — agent-first flagship, announced May 20, 2026. 1M context, native extended-thinking. $2.50/$7.50. Strong on agentic coding benchmarks. Important: closed-weights at launch (DashScope/Model Studio/OpenRouter only) — Alibaba typically open-weights smaller variants weeks/months later (e.g. Qwen3.6-35B-A3B, Apache 2.0). Qwen 3.7 Plus is the smaller sibling. Qwen 3.8 also referenced in some sources.
Mistral (open-weights) (MEDIUM confidence)
- Mistral Large 3 (
mistral-large-2512) — released Dec 2, 2025. 675B total / 41B active MoE, 256K context, text+image. Apache 2.0, self-hostable, strong multilingual + function calling. Lineup also includes Small 4, Medium 3.5.
Meta — Llama (LOW confidence — flag heavily)
- Llama 5 — reportedly released April 8, 2026, open-weight. Sources claim 600B params / 5M-token context / recursive self-improvement / Apache 2.0 — but these specs come from unreliable blog sources and read as inflated/hype. Treat all Llama 5 specifics as unverified until confirmed against Meta's official announcement.
How to choose a model for a task
Match the task to the cheapest tier that clears the quality bar — don't reflexively reach for the flagship.
1. Classify the task difficulty: - Simple/high-volume (classify, extract, summarize, route, draft): cheap tier — Haiku 4.5, GPT-5.6 Luna, Gemini 3.5 Flash, or open-weight DeepSeek V4 Flash. Big cost savings, negligible quality loss. - Balanced/production (most app workloads, RAG, chat, moderate reasoning): mid tier — Sonnet 5, GPT-5.6 Terra, Gemini 3.1 Pro. Best cost/quality. - Hardest (long-horizon agents, complex coding, deep reasoning, novel problems): premium — Opus 4.8, GPT-5.6 Sol, or Fable 5 / (Gemini 3.5 Pro once GA) for the absolute ceiling.
2. Weigh the axes that matter for your task: - Coding / agentic → Opus 4.8, Sonnet 5, GPT-5.6 Sol, Qwen 3.7 Max all strong. - Multimodal (image/audio/video) → Gemini line is the multimodal leader; Gemini 3.5 Pro (when GA) adds audio/video depth. - Very long context → most flagships now ship 1M; Gemini 3.5 Pro targets 2M (not yet GA). Watch long-context surcharges (OpenAI >272K, Gemini/Google >200K). - Cost-critical / self-hosting / data-residency → open weights: DeepSeek V4 (MIT), Mistral Large 3 (Apache 2.0), Llama 5 (open, specs unverified). Qwen 3.7 Max is not open at launch. - Cheapest per token overall → DeepSeek V4 Flash ($0.14/$0.28) is dramatically below the closed flagships.
3. Practical selection heuristics:
- Start one tier below where you think you need to be; upgrade only if quality fails your eval.
- Use effort/thinking controls (Anthropic effort, OpenAI reasoning tiers) to trade cost vs. quality within a model before jumping tiers.
- For agentic loops, a premium model at high effort often costs less total than a mid model that flails across many turns.
- Prompt caching / batch APIs cut cost 50–90% on repeated-context workloads — factor before comparing sticker prices.
- Open-weights make sense when you have volume + infra, need data control, or want no per-token cost; otherwise hosted flagships win on convenience.
VOLATILE CLAIMS (fact-check before relying on any of these)
Every item below is a specific model-name, version, date, spec, or price claim that will drift. Confidence tags: [HIGH] = first-party Anthropic reference; [MED] = third-party aggregators/blogs; [LOW] = single/unreliable source.
Anthropic (from skill reference — but skill data has a cache date and may lag):
1. [HIGH] Claude Opus 4.8 = claude-opus-4-8, 1M ctx, 128K out, $5/$25 — current flagship default.
2. [HIGH] Claude Fable 5 = claude-fable-5, $10/$50 — most capable widely-released.
3. [HIGH] Claude Sonnet 5 = claude-sonnet-5, $3/$15 with $2/$10 intro pricing through 2026-08-31 (intro expiry date will pass).
4. [HIGH] Claude Haiku 4.5 = claude-haiku-4-5, 200K ctx, $1/$5.
5. [HIGH] Opus 4.7/4.6 still active as prior-gen.
OpenAI: 6. [MED] GPT-5.6 family (Sol/Terra/Luna) released July 9, 2026. 7. [MED] Sol $5/$30, Terra $2.50/$15, Luna $1/$6; context ~1.05M, 128K output. 8. [MED] Long-context surcharge above 272K input tokens (~$10/$45 for Sol). 9. [MED] GPT-5.5 (April 23, 2026), GPT-5.4 family, GPT-5.1 ($1.25/$10) still listed. 10. [LOW] Exact GPT-5.6 model ID strings for the API (Sol/Terra/Luna → API identifiers) — not confirmed, verify against OpenAI docs.
Google Gemini: 11. [MED] Gemini 3.1 Pro is the current GA flagship (July 2026), 1M ctx, $2/$12, input ~doubles above 200K prompts. 12. [MED] Gemini 3.5 Pro announced (2M ctx, Deep Think) but NOT GA as of ~July 7, 2026 — do not treat as shipping; launch repeatedly slipped. 13. [MED] Gemini 3.5 Flash live since May 19, 2026.
DeepSeek: 14. [MED] DeepSeek V4 shipped April 24, 2026, MIT-licensed open weights. 15. [MED] V4 Pro: 1.6T/~49B MoE, 1M ctx, $0.44/$0.87. V4 Flash: 284B/~13B, $0.14/$0.28. 16. [MED] R2 rumored but not released.
Alibaba Qwen: 17. [MED] Qwen 3.7 Max announced May 20, 2026; 1M ctx; closed-weights at launch; $2.50/$7.50. 18. [MED] Qwen 3.7 Plus, 3.8 also referenced — naming/status uncertain.
Mistral:
19. [MED] Mistral Large 3 = mistral-large-2512, released Dec 2, 2025; 675B/41B MoE, 256K ctx, Apache 2.0.
Meta: 20. [LOW] Llama 5 (claimed April 8, 2026): 600B params, 5M-token context, recursive self-improvement, Apache 2.0 — sourced from unreliable blogs, specs look inflated. Verify against Meta's official announcement before citing anything about Llama 5.
Cross-cutting: 21. [MED] Long-context surcharge thresholds (OpenAI 272K, Google 200K) — verify current thresholds. 22. General note: all non-Anthropic pricing/versions come from aggregator sites, not fetched official docs. For a production wiki, re-verify each against the provider's official pricing/model pages.
Sources (non-Anthropic data): Morph OpenAI pricing, OpenAI GPT-5.6 announcement, Simon Willison GPT-5.6, eesel Gemini 3 pricing, eesel Gemini 3.5 Pro, Morph DeepSeek V4, Kingy AI open-weight models, codersera Qwen 3.7 Max, Vals.ai Mistral Large 3, RAGyfied Llama 5 (low-confidence).