Editorial investigation
The Reasoning Gap: How Claude Fable 5.1 Redefined the Quality Ceiling
An investigation into the emergent structural behaviors of the latest frontier update and its impact on the latent quality scores.
Explore full analysisLoading the current published AI model ranking…
Quality score uses HLE and SWE-Bench Pro from the published roster. Provider speed and pricing never change this score. For exact weights and coverage rules, see the Methodology page.
View
Quality view
Quality scores verified reasoning and coding evidence without applying price, routing speed, or provider availability as a bonus or penalty.
Value view
Value keeps cheap-but-weak models out with the HLE floor, then rewards qualified models that deliver more intelligence at lower standard conversation cost.
Intelligence view
Intelligence ranks the same published roster by Humanity's Last Exam in thinking mode with no tools. Supporting benchmarks stay visible, but they do not change the public axis label.
Coding view
Coding is a raw benchmark lens over the same published roster. SWE-Bench Pro determines rank, while accepted companion coding evidence stays visible when exact-variant benchmark rows are verified and does not change order.
Buyer-facing table
Score explanation
Why the quality score changes
Quality is recalculated from verified HLE and SWE-Bench Pro. Provider responsiveness and AA-Omniscience never affect this score or rank.
ARC-AGI-2 is withheld from this primary table until accepted coverage reaches 60% of the published roster.
Editorial investigation
An investigation into the emergent structural behaviors of the latest frontier update and its impact on the latent quality scores.
Explore full analysis"Quality is what still makes sense after the excitement fades."
PickAIModel Editorial
Quality note
AA-Omniscience
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
Claude Fable 5.1 combines the strongest current locked-variant HLE and SWE-Bench Pro evidence with a 1M-token context window. Its $10 per million input-token and $50 per million output-token prices make it a specialist choice for difficult coding and knowledge work, while safeguard fallback means integrations should surface the model identity returned by the API.
Quality score
AA-Omniscience
AA-Omniscience
AA-Omniscience
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
Claude Opus 5 combines a 1M-token context window with strong exact-variant HLE and SWE-Bench Pro results. Its $5 per million input-token and $25 per million output-token prices keep it in the premium tier, so it is best reserved for work where capability matters more than unit cost.
Quality score
AA-Omniscience
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
GPT-5.6 Sol is OpenAI’s flagship tier for complex coding and professional work. Its strong independent HLE and official SWE-Bench Pro evidence come with the family’s highest API price, so it is best reserved for work that justifies maximum capability.
Quality score
AA-Omniscience
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
Gemini 3.8 Flash combines a 1M-token context window with current HLE and SWE-Bench Pro evidence, multimodal input, and unusually low introductory token prices. Google positions it for software engineering and agentic workflows; the standard $1.50/$7.50 prices begin January 1, 2027.
Quality score
AA-Omniscience
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
Claude Sonnet 5 is the right default Sonnet model for most production teams that want stronger coding, agentic, and multi-step performance without moving every workload to Opus. It replaces Sonnet 4.6 in the current Claude lineup, carries stronger accepted-source HLE and SWE-Bench Pro evidence, and adds a 1M token context window that materially improves RAG, repository-scale coding, and document-heavy workflows. The migration still deserves an engineering review: the public row uses Anthropic vendor-reported benchmark evidence, the introductory API price ends after August 31, 2026, and plan-level availability or limits can vary by Claude surface. Treat it as the new default, test your real prompts and retrieval payloads, and route to Opus only when the task genuinely needs the higher-cost tier.
Quality score
AA-Omniscience
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
GPT-5.6 Terra is the balanced tier for everyday professional and coding workloads. It gives up some peak performance versus Sol in exchange for lower token pricing, making it the practical default when both capability and cost matter.
Quality score
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
This model is still under editorial review. We will publish a verdict as soon as we have completed our review of the AI model.
Quality score
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
Qwen3.7 Max is the optimal choice when your pipeline demands rigorous, multi-step logical deduction, complex code generation, or scientific analysis, and when cost-efficiency at scale is a primary constraint.
Quality score
AA-Omniscience
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
Gemini 3.1 Pro is the ultimate all-in-one creative partner. It does more than chat; it builds. From generating cinematic video and studio-quality music to managing your life through seamless Google Workspace integration, it turns complex tasks into instant results. It is the fastest, most versatile tool for turning ideas into reality without needing a technical degree. True multimodality means it can create stunning video, professional images, and high-fidelity music in seconds. Its massive context window lets it remember entire books or long documents, so you do not have to repeat yourself. It works inside Gmail, Docs, and Drive to automate daily chores. It also delivers high-level reasoning and instant answers without the lag of older models. If you want an AI that acts as a creative studio, personal assistant, and expert researcher all in one subscription, Gemini 3.1 Pro is the gold standard.
Quality score
AA-Omniscience
HLE
Conversation
3K tokens/chat
Context
Verdict
Leaderboard verdict
Full verdict
GPT-5.6 Luna is the fastest and lowest-cost tier in the family. Its accepted HLE and SWE-Bench Pro evidence keep it competitive for high-volume workflows, while difficult tasks should still be tested against Terra or Sol.
Quality score