Scores are deterministic. AI writes editorial prose only. The public site reads a published snapshot through the DAL, never from a live benchmark scrape. When the methodology changes, the update is versioned and applied only through a republished snapshot, so buyers always read the same rules that produced the current roster.
The two scores that determine the public leaderboard.
Quality and Value are the only official PickAI scores. They rank the same published roster with different rules, so buyers can compare raw capability separately from buying power.
Quality Score
Capability-first benchmark pool
0.556 HLE + 0.278 SWE-Bench Pro + 0.167 operational capability
55.6% HLE - The dominant broad-capability anchor.
27.8% SWE-Bench Pro - The canonical coding input when an exact current model row is verified.
Quality and Value are scores. Intelligence and Coding are benchmark lenses.
The product has two official scores and two raw benchmark views. Intelligence exposes Humanity's Last Exam in thinking mode with no tools. Coding exposes SWE-Bench Pro directly. Both keep companion benchmarks visible without turning them into extra PickAI scores.
Official scores
Quality ranks capability across the shared benchmark pool. Value ranks the smartest qualified models for the lowest normal-use token cost.
Raw benchmark views
Intelligence and Coding show direct benchmark evidence beside the same roster. They surface supporting signals, but they do not rewrite the official score logic.
What stays separate
Display-only signals stay visible without becoming hidden score inputs.
Subscription pricing can appear on Quality pages without changing the Quality score.
Terminal-Bench 2.1 and 2.0 appear in separate columns beside SWE-Bench Pro only when the current roster has accepted exact-variant coverage.
Score coverage
What the official scores include, and what the raw benchmark views show.
Y means the signal is part of the ranking rule. ? means the signal is tracked and shown when available, but the row note defines whether it affects a disclosed component. X means the signal stays out of the scoring formula.
Official PickAI scores
These cards define the published order.
Quality and Value are formula-driven. The score order is public, deterministic, and separate from the benchmark-only tabs.
Quality Score
Quality is capability-first: HLE, SWE-Bench Pro, and OpenRouter operational capability are the only scored inputs. OpenRouter responsiveness is display-only. Price is not a standalone Quality input.
HLE
Primary capability anchor.
SWE-Bench Pro coding
Canonical coding input when an exact current model row is verified.
Supporting buyer signal
Conversation Value stays disclosed beside the rank.
Monthly plan price / API-equivalent standard conversation cost
Shared basket Standard conversation size - Each model uses the same fixed input/output token basket for like-for-like buying-power comparison.
API rates Published input and output pricing - The benchmark uses accepted API pricing facts from the current snapshot and does not guess missing prices.
Plan price Monthly subscription basis - The selected paid tier is converted into an API-equivalent conversation count.
Disclosure Shown without changing Quality - The Quality leaderboard can show the benchmark for context while Quality rank remains capability-first.
PickAI Conversation Value is a proprietary buying-power benchmark. It helps explain subscription value in buyer terms, but it does not change the Quality leaderboard order.
AI tools badges
Visual editorial signals for the tools directory.
Badges on the AI tools page are selected through AI editorial review. They are not scores, do not affect model rankings or tool order, and are never assigned because of affiliate payouts. The question-mark badge means the tool has not been reviewed deeply enough yet to assign a more specific signal.
Most PopularMost PopularBroad public awareness, strong adoption, or frequent discussion makes this tool especially recognizable.
Broad public awareness, strong adoption, or frequent discussion makes this tool especially recognizable.
Highly Rated
Trust & governance
Why you can trust what you are reading.,
This page is public and buyer-facing, but the safeguards behind it are still explicit. Publication rules, update cadence, and operational controls all exist to stop unverified benchmark noise from leaking into the public ranking.
How publication works
Editorial prose is allowed. Score edits are not.
Scores are read-only after snapshot generation.
The public site reads a published Supabase snapshot instead of fetching live benchmark pages.
Methodology changes ship through a versioned update and a republished snapshot, not a hidden runtime override.
Evidence & benchmark sources
Accepted reputable citations only, grouped by the evidence they support.
Public ranking rows are published only when the underlying benchmark and pricing inputs can be tied to accepted reputable sources that any buyer can fact-check. Unsupported sources are omitted rather than estimated.
Reasoning & novel problem solving
Primary benchmark evidence for raw reasoning views and supporting problem-solving context.
Primary public leaderboard source for MathArena Expected Performance.
FAQ
Questions buyers ask before they trust a ranking.
What is the difference between Quality and Value?
Quality ranks the published roster using HLE, SWE-Bench Pro, and OpenRouter operational capability. Responsiveness is display-only. Value asks which qualified model is smartest for the lowest normal-use token cost, using HLE and published API prices.
Why can the Intelligence view rank models differently from Quality?
The Intelligence view is the Humanity's Last Exam lens in thinking mode with no tools over the same published roster. It surfaces HLE evidence directly, while Quality remains the broader deterministic ranking.
Why does the Coding view show more than one benchmark?
SWE-Bench Pro sets the rank, while Terminal-Bench 2.1 and 2.0 stay in separate companion columns so buyers can compare agentic terminal evidence without mixing benchmark versions.
Display only OpenRouter responsiveness - Time to first token and generated-token throughput; never a Quality or Value input.
This staged formula version uses fixed normalization ranges established from its reviewed 11-model reference cohort, so the public top-10 cap cannot change scores. Responsiveness is display-only until OpenRouter provides a stable documented performance feed. Price is not applied as a direct Quality bonus or penalty; published token costs affect only the separate Value conversation-affordability calculation. Unsupported capability signals stay display-only until comparable accepted evidence exists.
Value Score
Smartness and conversation affordability
0.40 smartness + 0.60 conversation affordability
40% Smartness - A model must clear HLE 20 first, then stronger HLE scores earn more credit.
60% Conversation affordability - The buying-power signal comes from a shared normal-use conversation cost using published API token prices.
A model must score at least 20 on HLE to qualify. Standard conversation affordability carries more weight than raw HLE, and token costs are scaled so large price gaps count fairly.
OpenRouter responsiveness remains display-only until a stable documented performance feed is available.
Quality leaderboard rows can show Conversation Value beside capability evidence without changing the Quality rank.
AI-written prose can explain the roster, but it never edits numeric values.
OpenRouter operational capability
Context, tools/structured output, and multimodal support.
OpenRouter responsiveness
TTFT and throughput are display-only and never change scores or ranks.
Companion coding benchmarks
Displayed beside Coding when exact-variant evidence is verified; not separate Quality inputs.
Subscription price
Displayed for buyer context; not a standalone Quality input.
Token costs
Can feed disclosed conversation-affordability buyer components when active; never guessed or used as a hidden override.
Value Score
Value ranks qualified models by a plain buyer question: what is the smartest model for the lowest normal-use token cost?
HLE
Required floor and smartness anchor.
Standard conversation cost
Measured with the same normal-use conversation basket for every model.
Subscription price
Shown as context, not scored separately.
Missing pricing
Not guessed. A model needs published input and output token prices for Value.
Raw benchmark views
These cards explain the benchmark-only tabs.
Intelligence and Coding remain direct benchmark views over the same roster. Companion metrics stay visible, but they do not become hidden score inputs.
Intelligence View
This remains a Humanity's Last Exam view in thinking mode with no tools over the same roster. Supporting benchmarks stay visible, but they do not create a third PickAI score.
HLE (thinking, no tools)
Ranking signal.
MathArena
Displayed as supporting evidence when available.
ARC-AGI-2
Displayed as supporting evidence when available.
AA-Omniscience
Independent factual-reliability evidence displayed when available; never a Quality or Value input.
Context window
Not part of the raw intelligence view.
Speed
Not part of the raw intelligence view.
Price
Not part of the raw intelligence view.
Coding View
This remains a raw SWE-Bench Pro view over the same roster when public SWE-Bench coverage exists. Terminal-Bench 2.1 and 2.0 stay in separate companion columns only when accepted exact-variant coding evidence is available, and they do not create a fourth PickAI score.
SWE-Bench Pro
Ranking signal.
LiveCodeBench Pro
Displayed as supporting evidence only when an accepted exact-variant row is available.
Terminal-Bench 2.1
Displayed separately as supporting agentic terminal evidence only when an accepted exact-variant row is available.
Why the separation matters
Buyers can see more context without losing the ranking logic.
Capability stays capability
Quality and Value do the ranking work. Supporting context does not leak into the official order.
Buyer context stays visible
Pricing, app access, benchmark companions, and editorial notes stay available where they help a purchase decision.
Raw views stay raw
Intelligence and Coding expose direct benchmark evidence rather than a blended third or fourth PickAI score. For Intelligence, the ranked HLE row follows the site policy: thinking mode, no tools.
Trust stays explicit
The page tells you which signals rank, which signals inform, and which signals are intentionally excluded.
Highly RatedPublic review or reputation signals suggest this tool is especially well regarded.
Public review or reputation signals suggest this tool is especially well regarded.
Most InnovativeMost InnovativeThe tool stands out for a distinctive workflow, product concept, or AI implementation.
The tool stands out for a distinctive workflow, product concept, or AI implementation.
Creator FavoriteCreator FavoriteThe tool is especially useful for creative work such as design, video, images, audio, or writing.
The tool is especially useful for creative work such as design, video, images, audio, or writing.
Automation PickAutomation PickThe tool is especially useful for workflows, agents, integrations, or operational automation.
The tool is especially useful for workflows, agents, integrations, or operational automation.
Research PickResearch PickThe tool is especially useful for research, source review, search, or evidence-heavy knowledge work.
The tool is especially useful for research, source review, search, or evidence-heavy knowledge work.
Best for TeamsBest for TeamsThe tool is especially useful for team, meeting, workplace, or collaboration workflows.
The tool is especially useful for team, meeting, workplace, or collaboration workflows.
Needs ReviewNeeds ReviewWe have not reviewed this tool deeply enough yet to assign a specific badge.
We have not reviewed this tool deeply enough yet to assign a specific badge.
Refresh remains manual and admin-gated; there is no public or scheduled refresh trigger.
AI-written editorial copy refreshes after publication without touching numeric score values.
The operational diagnostics that explain candidate-pool blockers and publication withholding live in the protected status console, not on the public methodology page.
Benchmarks under review
These inputs are tracked publicly before they are allowed into the score formula.
SWE-Bench Pro / Terminal-Bench 2.1 / Terminal-Bench 2.0
These coding benchmarks test different facets of software capability. SWE-Bench Pro feeds Quality and ranks the Coding view when coverage exists, while Terminal-Bench 2.1 and 2.0 appear in separate columns only when accepted exact-variant companion rows are available.
ARC-AGI-2
ARC-AGI-2 tests novel visual pattern reasoning without prior exposure. It will be considered for inclusion only when comparable verified scores exist across all up to 10 models and a methodology version update is published.
ARC-AGI-3
ARC-AGI-3 is a watchlist benchmark for interactive agent reasoning. It will not enter the formula until stable comparative data exists across the published roster.
The rules we do not bend
01
Scores are immutable
Quality Score, Value Score, and benchmark sub-scores are read-only numbers. AI can write editorial prose, but it never changes the score fields.
02
Snapshot is the source of truth
The public site reads from the snapshot through the DAL. It does not fetch benchmark sources or call AI at request time, and unsupported benchmark packs are withheld rather than rendered.
03
Quality and value stay independent
There is no combined score. Quality and Value remain the only official PickAI scores, while available raw benchmark tabs stay direct views over the same roster.
04
Up to 10 models
Public pages and endpoints are capped at 10 models. The published roster tracks the latest stable up to 10 models in the current snapshot, with no overflow pagination or hidden extra rows.
05
AI prose is labeled
Any AI-assisted verdict, pros/cons list, or best-for block carries a visible “AI-assisted, editorially reviewed” label and data-ai-generated="true".
When updates happen & what we refuse
Update cadence
Admin-reviewed refresh: the public snapshot updates only when the protected refresh action is run after source review.
Narrow refresh scope: only newly discovered supported models and existing models with missing benchmark facts are targeted.
Operator diagnostics stay in the protected status console rather than the public methodology page.
Hard exclusions
We do not let AI change numeric scores.
We do not rank more than 10 models on public pages.
We do not expose public refresh triggers.
We do not add affiliate links to leaderboard rows or this page.
AI-written editorial copy is labeled wherever it appears
First-party leaderboard maintained by the benchmark creators. Report the scores as Aider publishes them, using the aider harness and the open-source benchmark dataset on GitHub.
NVIDIA-hosted model card for DeepSeek V4 Pro with model metadata and benchmark table. Treat as accepted tier-3 evidence and cite exact model rows only.
Aggregator comparison page that includes Aider Polyglot. Use only as a tertiary cross-check.
PickAI Conversation Value remains our disclosed buying-power benchmark at published API rates. It is shown as buyer context without becoming a hidden Quality-score input.
Where do I compare price, app access, and buyer notes in more detail?
Use the models index and the model detail pages when you want pricing, app access, buyer-facing guidance, and the supporting benchmark context next to the ranking.