원본 캡처
verified · type: reference · [meta, raw]

I now have comprehensive current data across all major providers. Let me compile the structured markdown summary.

A few important caveats I'll flag: much of this non-Anthropic data comes from third-party aggregator/blog sources (not official docs), the model landscape is churning monthly, and some "facts" (Llama 5 specs especially) look suspiciously inflated from unreliable sources. The Claude family data is from the authoritative skill reference.

Here is my structured summary:


2026 LLM Landscape — Practical Reference (mid-2026)

Compiled: 2026-07-23. This field moves monthly; treat every version number and price as perishable. The Anthropic/Claude data is from the authoritative first-party skill reference; all other providers rely on third-party aggregators and blogs (official docs were not directly fetched), so their specifics are lower-confidence — see VOLATILE CLAIMS at the end.


Quick comparison table

Provider Current flagship (mid-2026) Model ID / name Context Pricing tier ($/1M in-out) Relative strength
Anthropic Claude Opus 4.8 claude-opus-4-8 1M Premium ($5 / $25) Agentic coding, long-horizon autonomy, knowledge work
Anthropic (most capable) Claude Fable 5 claude-fable-5 1M Ultra-premium ($10 / $50) Hardest reasoning / long-horizon agentic
Anthropic (balanced) Claude Sonnet 5 claude-sonnet-5 1M Mid ($3 / $15; intro $2 / $10) Near-Opus quality at Sonnet cost
Anthropic (cheap/fast) Claude Haiku 4.5 claude-haiku-4-5 200K Cheap ($1 / $5) Speed-critical simple tasks
OpenAI GPT-5.6 Sol gpt-5.6 (Sol tier) ~1.05M Premium ($5 / $30) Frontier reasoning, computer-use, tool-heavy
OpenAI (balanced) GPT-5.6 Terra Terra tier ~1.05M Mid ($2.50 / $15) GPT-5.5-class quality, half the price
OpenAI (cheap) GPT-5.6 Luna Luna tier ~1.05M Cheap ($1 / $6) Summarize/classify/draft at volume
Google Gemini 3.1 Pro gemini-3.1-pro 1M Mid ($2 / $12; long-ctx surcharge >200K) Multimodal reasoning, agentic, computer use
Google (cheap/fast) Gemini 3.5 Flash gemini-3.5-flash 1M Cheap Fast; reportedly beats prior flagship on coding
DeepSeek (open) DeepSeek V4 Pro deepseek-ai/DeepSeek-V4-Pro 1M Very cheap ($0.44 / $0.87) Price-performance frontier, MIT-licensed
DeepSeek (open, small) DeepSeek V4 Flash DeepSeek-V4-Flash Ultra-cheap ($0.14 / $0.28) Cheap open-weight workhorse
Alibaba Qwen Qwen 3.7 Max qwen3.7-max 1M Cheap-mid ($2.50 / $7.50) Agentic coding; closed-weights at launch
Mistral (open) Mistral Large 3 mistral-large-2512 256K Open-weight (self-host) Apache 2.0, multilingual, function calling
Meta (open) Llama 5 (see caveat) Open-weight General open flagship — specs unverified

By provider

Anthropic — Claude (HIGH confidence: first-party reference)

  • Claude Fable 5 (claude-fable-5) — most capable widely released model. 1M context, 128K output. $10/$50. Thinking always on; refusal-fallback handling; requires 30-day data retention. Ultra-premium.
  • Claude Opus 4.8 (claude-opus-4-8) — the flagship most people should default to. 1M context, 128K output. $5/$25. State-of-the-art long-horizon agentic + coding + knowledge work. Premium.
  • Claude Opus 4.7 / 4.6 (claude-opus-4-7, claude-opus-4-6) — previous-gen Opus, still active.
  • Claude Sonnet 5 (claude-sonnet-5) — best speed/intelligence balance; near-Opus on coding/agentic. 1M context. $3/$15 ($2/$10 intro through 2026-08-31). Mid.
  • Claude Haiku 4.5 (claude-haiku-4-5) — fastest/cheapest. 200K context, 64K output. $1/$5. Cheap.

OpenAI — GPT-5.6 family (MEDIUM confidence: aggregators + OpenAI blog)

Released July 9, 2026 as three tiers with celestial names. ~1.05M context, 128K output. Requests above 272K input tokens are billed at a higher long-context rate (~$10/$45 for Sol). - Sol (flagship) — $5/$30. Frontier reasoning, long-horizon agentic. Premium. - Terra (balanced) — $2.50/$15. GPT-5.5-class at half the price; the production migration target. Mid. - Luna (cheap/fast) — $1/$6. High-volume simple work. Cheap. - Older gpt-5.5 ($5/$30), gpt-5.4 family, and gpt-5.1 ($1.25/$10) remain on the price sheet.

Google — Gemini (MEDIUM confidence)

  • Gemini 3.1 Pro (gemini-3.1-pro) — current GA flagship as of July 2026. 1M context. $2/$12 input/output; input roughly doubles above 200K-token prompts. Leads multimodal reasoning, agentic, computer use. Mid tier by price.
  • Gemini 3.5 Flash — live since May 2026; cheap/fast tier that reportedly beats the previous flagship on coding.
  • Gemini 3.5 Pro — announced (Google I/O, May 2026) with a 2M-token context window and "Deep Think" mode, but NOT generally available as of early-mid July 2026 (slipped past June/July; August window rumored). Do not treat as shipping.

DeepSeek — V4 (open-weights) (MEDIUM confidence)

Shipped April 24, 2026. Two-tier open release, MIT-licensed weights on Hugging Face. - V4 Pro — 1.6T total / ~49B active MoE, 1M context. $0.44/$0.87 (cache-hit input ~$0.0036). Best open-weight price-performance. - V4 Flash — 284B total / ~13B active. $0.14/$0.28. Ultra-cheap. - R2 (reasoning model) — persistently rumored, not released; V4 is the real 2026 flagship.

Alibaba — Qwen (MEDIUM confidence)

  • Qwen 3.7 Max — agent-first flagship, announced May 20, 2026. 1M context, native extended-thinking. $2.50/$7.50. Strong on agentic coding benchmarks. Important: closed-weights at launch (DashScope/Model Studio/OpenRouter only) — Alibaba typically open-weights smaller variants weeks/months later (e.g. Qwen3.6-35B-A3B, Apache 2.0). Qwen 3.7 Plus is the smaller sibling. Qwen 3.8 also referenced in some sources.

Mistral (open-weights) (MEDIUM confidence)

  • Mistral Large 3 (mistral-large-2512) — released Dec 2, 2025. 675B total / 41B active MoE, 256K context, text+image. Apache 2.0, self-hostable, strong multilingual + function calling. Lineup also includes Small 4, Medium 3.5.

Meta — Llama (LOW confidence — flag heavily)

  • Llama 5 — reportedly released April 8, 2026, open-weight. Sources claim 600B params / 5M-token context / recursive self-improvement / Apache 2.0 — but these specs come from unreliable blog sources and read as inflated/hype. Treat all Llama 5 specifics as unverified until confirmed against Meta's official announcement.

How to choose a model for a task

Match the task to the cheapest tier that clears the quality bar — don't reflexively reach for the flagship.

1. Classify the task difficulty: - Simple/high-volume (classify, extract, summarize, route, draft): cheap tier — Haiku 4.5, GPT-5.6 Luna, Gemini 3.5 Flash, or open-weight DeepSeek V4 Flash. Big cost savings, negligible quality loss. - Balanced/production (most app workloads, RAG, chat, moderate reasoning): mid tier — Sonnet 5, GPT-5.6 Terra, Gemini 3.1 Pro. Best cost/quality. - Hardest (long-horizon agents, complex coding, deep reasoning, novel problems): premium — Opus 4.8, GPT-5.6 Sol, or Fable 5 / (Gemini 3.5 Pro once GA) for the absolute ceiling.

2. Weigh the axes that matter for your task: - Coding / agentic → Opus 4.8, Sonnet 5, GPT-5.6 Sol, Qwen 3.7 Max all strong. - Multimodal (image/audio/video) → Gemini line is the multimodal leader; Gemini 3.5 Pro (when GA) adds audio/video depth. - Very long context → most flagships now ship 1M; Gemini 3.5 Pro targets 2M (not yet GA). Watch long-context surcharges (OpenAI >272K, Gemini/Google >200K). - Cost-critical / self-hosting / data-residency → open weights: DeepSeek V4 (MIT), Mistral Large 3 (Apache 2.0), Llama 5 (open, specs unverified). Qwen 3.7 Max is not open at launch. - Cheapest per token overall → DeepSeek V4 Flash ($0.14/$0.28) is dramatically below the closed flagships.

3. Practical selection heuristics: - Start one tier below where you think you need to be; upgrade only if quality fails your eval. - Use effort/thinking controls (Anthropic effort, OpenAI reasoning tiers) to trade cost vs. quality within a model before jumping tiers. - For agentic loops, a premium model at high effort often costs less total than a mid model that flails across many turns. - Prompt caching / batch APIs cut cost 50–90% on repeated-context workloads — factor before comparing sticker prices. - Open-weights make sense when you have volume + infra, need data control, or want no per-token cost; otherwise hosted flagships win on convenience.


VOLATILE CLAIMS (fact-check before relying on any of these)

Every item below is a specific model-name, version, date, spec, or price claim that will drift. Confidence tags: [HIGH] = first-party Anthropic reference; [MED] = third-party aggregators/blogs; [LOW] = single/unreliable source.

Anthropic (from skill reference — but skill data has a cache date and may lag): 1. [HIGH] Claude Opus 4.8 = claude-opus-4-8, 1M ctx, 128K out, $5/$25 — current flagship default. 2. [HIGH] Claude Fable 5 = claude-fable-5, $10/$50 — most capable widely-released. 3. [HIGH] Claude Sonnet 5 = claude-sonnet-5, $3/$15 with $2/$10 intro pricing through 2026-08-31 (intro expiry date will pass). 4. [HIGH] Claude Haiku 4.5 = claude-haiku-4-5, 200K ctx, $1/$5. 5. [HIGH] Opus 4.7/4.6 still active as prior-gen.

OpenAI: 6. [MED] GPT-5.6 family (Sol/Terra/Luna) released July 9, 2026. 7. [MED] Sol $5/$30, Terra $2.50/$15, Luna $1/$6; context ~1.05M, 128K output. 8. [MED] Long-context surcharge above 272K input tokens (~$10/$45 for Sol). 9. [MED] GPT-5.5 (April 23, 2026), GPT-5.4 family, GPT-5.1 ($1.25/$10) still listed. 10. [LOW] Exact GPT-5.6 model ID strings for the API (Sol/Terra/Luna → API identifiers) — not confirmed, verify against OpenAI docs.

Google Gemini: 11. [MED] Gemini 3.1 Pro is the current GA flagship (July 2026), 1M ctx, $2/$12, input ~doubles above 200K prompts. 12. [MED] Gemini 3.5 Pro announced (2M ctx, Deep Think) but NOT GA as of ~July 7, 2026 — do not treat as shipping; launch repeatedly slipped. 13. [MED] Gemini 3.5 Flash live since May 19, 2026.

DeepSeek: 14. [MED] DeepSeek V4 shipped April 24, 2026, MIT-licensed open weights. 15. [MED] V4 Pro: 1.6T/~49B MoE, 1M ctx, $0.44/$0.87. V4 Flash: 284B/~13B, $0.14/$0.28. 16. [MED] R2 rumored but not released.

Alibaba Qwen: 17. [MED] Qwen 3.7 Max announced May 20, 2026; 1M ctx; closed-weights at launch; $2.50/$7.50. 18. [MED] Qwen 3.7 Plus, 3.8 also referenced — naming/status uncertain.

Mistral: 19. [MED] Mistral Large 3 = mistral-large-2512, released Dec 2, 2025; 675B/41B MoE, 256K ctx, Apache 2.0.

Meta: 20. [LOW] Llama 5 (claimed April 8, 2026): 600B params, 5M-token context, recursive self-improvement, Apache 2.0 — sourced from unreliable blogs, specs look inflated. Verify against Meta's official announcement before citing anything about Llama 5.

Cross-cutting: 21. [MED] Long-context surcharge thresholds (OpenAI 272K, Google 200K) — verify current thresholds. 22. General note: all non-Anthropic pricing/versions come from aggregator sites, not fetched official docs. For a production wiki, re-verify each against the provider's official pricing/model pages.

Sources (non-Anthropic data): Morph OpenAI pricing, OpenAI GPT-5.6 announcement, Simon Willison GPT-5.6, eesel Gemini 3 pricing, eesel Gemini 3.5 Pro, Morph DeepSeek V4, Kingy AI open-weight models, codersera Qwen 3.7 Max, Vals.ai Mistral Large 3, RAGyfied Llama 5 (low-confidence).