AI Intelligence per Dollar — Frontier Model Benchmark Efficiency (2020–2026)
Tracks how much AI capability a dollar buys, from GPT-3 (2020) to today's frontier. Composite capability is the plain average of three benchmark scores — swapped twice as labs retired saturated evaluations — divided by published API output-token price. Covers OpenAI, Anthropic, Google, DeepSeek, Meta Llama, Moonshot Kimi and xAI through GPT-5.6, Claude Opus 4.6, Gemini 3.1 Pro, DeepSeek V4-Flash, Kimi K3 and Grok 4.5.
Data
| Period | Model | Era | Knowledge (MMLU / GPQA-D) | Code (HumanEval / SWE-bench-V / TB2.1) | Reasoning (MATH / HLE) | Composite | Cost / 1M Output (at release) | Intelligence-per-Dollar |
|---|---|---|---|---|---|---|---|---|
| 2026-Q3 | Kimi K3 (Moonshot AI) | Era 3 | 93.5% | 88.3% | 44.3% | 75.4 | $15.00 | 5 |
| 2026-Q2 | GPT-5.6 Sol (OpenAI) | Era 3 | 94.1% | 88.8% | 47.2% | 76.7 | $30.00 | 2.6 |
| 2026-Q2 | Grok 4.5 (xAI) | Era 3 | 93.1% | 81.6% | 40.3% | 71.7 | $6.00 | 12 |
| 2026-Q1 | DeepSeek V4-Flash (DeepSeek) | Era 3 | 89.4% | 61.8% | 32.1% | 61.1 | $0.28 | 218.2 |
| 2026-Q1 | Gemini 3.1 Flash-Lite (Google) | Era 3 | 86.9% | n/a | 16.0% | n/a | $1.50 | n/a |
| 2026-Q1 | Gemini 3.1 Pro (Google) | Era 3 | 94.1% | 73.8% | 44.7% | 70.9 | $12.00 | 5.9 |
| 2026-Q1 | GPT-5.4 nano (OpenAI) | Era 3 | 81.7% | n/a | 26.5% | n/a | $1.25 | n/a |
| 2026-Q1 | Claude Opus 4.6 (Anthropic) | Era 2 | 91.3% | 80.8% | 40.0% | 70.7 | $25.00 | 2.8 |
| 2025-Q1 | Claude 3.7 Sonnet (Anthropic) | Era 2 | 78.2% | 63.7% | 8.0% | 50 | $15.00 | 3.3 |
| 2025-Q1 | DeepSeek R1 (DeepSeek) | Era 2 | 71.5% | 49.2% | n/a | n/a | $2.19 | n/a |
| 2024-Q4 | DeepSeek V3 (DeepSeek) | Era 1 | 88.5% | 65.2% | 61.6% | 71.8 | $1.10 | 65.3 |
| 2024-Q3 | Llama 3.1 405B (Meta) | Era 1 | 87.3% | 89.0% | 73.8% | 83.4 | $3.50 | 23.8 |
| 2024-Q2 | GPT-4o (OpenAI) | Era 1 | 88.7% | 90.2% | 76.6% | 85.2 | $15.00 | 5.7 |
| 2023-Q1 | GPT-4 (OpenAI) | Era 1 | 86.4% | 67.0% | 42.2% | 65.2 | $60.00 | 1.1 |
| 2023-Q1 | GPT-3.5-turbo (OpenAI) | Era 1 | 70.0% | 53.9% | 34.1% | 52.7 | $2.00 | 26.4 |
| 2020-Q2 | GPT-3 (OpenAI) | Era 1 | 43.9% | 14.0% | 4.0% | 20.6 | $60.00 | 0.3 |
About this Dataset
This page tracks Intelligence-per-Dollar: the AI capability a dollar of API spend buys, from GPT-3 in 2020 through the mid-2026 frontier. Composite capability is the plain, unweighted average of three published benchmark scores covering Knowledge, Code and Reasoning, and Intelligence-per-Dollar divides that composite by the model's published price per million output tokens. Both formulas hold fixed across the full 2020–2026 span; what changes is which specific benchmark measures each axis, because the industry has retired two full benchmark generations as frontier models saturated them in turn. Composite scores are not directly comparable across those generations at face value — a lower score under the current, harder benchmark set does not mean a less capable model, a point covered in detail below. One pattern holds across all three generations regardless: the model with the best raw benchmark score is rarely the model with the best Intelligence-per-Dollar ratio.
How the calculation works
Composite is the plain average of three benchmark scores; Intelligence-per-Dollar divides that by the model's published price per million output tokens. Both formulas are fixed: nothing is normalised, rebased, or weighted. The three benchmarks behind "Composite" have changed twice, though, because the industry retired each generation once frontier models saturated it.
A model is only scored in an era if all three of that era's benchmarks are published for it. A model missing one is shown with the missing cell marked n/a and no computed ratio, rather than an average over whatever happens to exist.
Three benchmark generations
Era 1 (2020–2024) used MMLU, HumanEval and MATH, the benchmark trio that defined the GPT-3-to-GPT-4 years. The last frontier release in this dataset to report all three together is DeepSeek V3, in December 2024. No model since has published a complete MMLU/HumanEval/MATH trio: Anthropic's Claude 3.x and 4.x model cards stopped reporting HumanEval outright, and by 2025 most labs had moved on to harder evaluations as MMLU and HumanEval approached their ceilings.
Era 2 (2024–2026) replaced the trio with GPQA Diamond, SWE-bench Verified and Humanity's Last Exam, and that generation was retired too. OpenAI stopped publishing SWE-bench Verified for its frontier models in February 2026, stating the benchmark no longer measured frontier coding progress; Artificial Analysis removed it from its own evaluation suite the same year. None of GPT-5.6, Grok 4.5 or Kimi K3, the models actually shipping in mid-2026, have a published SWE-bench Verified score anywhere. Era 3 replaces just that one axis with Terminal-Bench v2.1, keeping GPQA Diamond and HLE unchanged. Two models in this dataset, Gemini 3.1 Pro and DeepSeek V4-Flash, published scores under both Era 2's SWE-bench Verified and Era 3's Terminal-Bench v2.1, which is what makes the handover checkable directly rather than assumed; the table and chart list each of them once, under its more current Era 3 figure.
Anthropic's current flagship, Claude Opus 5, illustrates how far this fragmentation has gone: its own launch page reports Frontier-Bench, CursorBench, ARC-AGI 3, OSWorld and GDPval, and none of the six benchmarks tracked on this page. It does not appear in this dataset because none of its scores on those six benchmarks are published anywhere. Claude Opus 4.6, one generation earlier, still reported GPQA Diamond, SWE-bench Verified and HLE, and anchors this dataset's Era 2 instead.
The cheapest tier tells a consistent, separate story: agentic coding benchmarks haven't reached it yet. GPT-5.4 nano and Gemini 3.1 Flash-Lite both have published GPQA Diamond and HLE scores, but neither has a published SWE-bench Verified or Terminal-Bench v2.1 result — those evaluations require running an agent against real repositories, which labs and third-party harnesses evidently don't do routinely for their smallest models. Both appear in the table with the code column marked n/a and no computed Intelligence-per-Dollar, by the same completeness rule applied everywhere else on this page. DeepSeek V4-Flash is the one exception: it is DeepSeek's own economy-tier model, and it is the only sub-$0.30 model in this dataset with a complete benchmark trio in either 2026 era.
Reading the composite across eras
The table's Composite column is not one continuous scale, and reading it top to bottom as if it were will produce the wrong conclusion. Era 1's composite (MMLU, HumanEval, MATH; GPT-4o's 85.2 is the highest figure this dataset records) and Era 2/3's composite (GPQA Diamond, SWE-bench Verified or Terminal-Bench v2.1, HLE; Claude Opus 4.6 scores 70.7) are built from benchmarks aimed at very different difficulty targets. GPQA Diamond is designed so that PhD-level experts working outside their own specialty score only 65 to 74 percent on it. Humanity's Last Exam was built to sit at the frontier of expert human knowledge; models scored close to zero percent on it at introduction, and even GPT-5.6 Sol, this dataset's highest Era 3 composite at 76.7, manages only 47.2 percent on that one axis. SWE-bench Verified and Terminal-Bench v2.1 both require multi-step agentic work against real code repositories, not the isolated single-function generation HumanEval tested.
None of that difficulty gap is built into the composite formula, which stays a plain three-way average in every era, deliberately, so any reader can recompute a row straight from the table without trusting a hidden weighting. The consequence is that a lower Era 2 or Era 3 composite than an Era 1 one is not a sign of weaker capability — it is more likely the opposite, a sign the yardstick got harder. Claude 3.7 Sonnet's Era 2 composite of 50 sits well below GPT-4o's Era 1 figure of 85.2, but that is two different rulers, not one model outperforming another.
Cost-efficiency: the value leaders
DeepSeek holds the best Intelligence-per-Dollar ratio in every era it appears in, though the margin is smaller than an earlier, uncorrected pricing figure once suggested. DeepSeek V3, priced at $1.10 per million output tokens, illustrates why: a $0.28 launch-window promotional rate applied for about six weeks before reverting to $1.10 on 2025-02-08, which is the figure used here. At that price, DeepSeek V3 delivers about 65.3 composite points per dollar, the highest Era 1 ratio in this dataset. Its 2026 successor, DeepSeek V4-Flash, leads Era 3 outright at 218.2 points per dollar, well ahead of Grok 4.5's 11.95, the next-highest Era 3 ratio. Neither DeepSeek model tops its era on raw capability (GPT-4o and GPT-5.6 Sol do that), but open-weight pricing an order of magnitude below the frontier makes DeepSeek's models the cheapest way to buy a given level of measured capability at every point this dataset covers.
What this doesn't measure
Intelligence-per-Dollar figures for reasoning models carry a caveat the ratio itself doesn't show. Every model in Era 2 and Era 3 that does explicit chain-of-thought reasoning generates internal reasoning tokens ahead of its visible output, typically billed at the same output rate; a demanding query can produce many more reasoning tokens than response tokens. This dataset prices every model on its published per-token rate only, with no adjustment for reasoning-token volume, so ratios for reasoning-heavy models on hard queries are likely more favourable here than they turn out to be in production. Buyers evaluating a reasoning model for a specific workload should measure their own token consumption rather than relying on the ratio shown here.
Other methodological considerations for professional use of this dataset:
- Completeness, not coverage: A model's composite is computed only when all three of its era's benchmarks are published for it. Partial trios are shown with the missing figure marked n/a and no computed ratio — never averaged over whatever exists.
- Price is release-era, not current: Costs reflect the published output-token price at or near each model's release. Several older models, GPT-3, GPT-3.5-turbo, GPT-4, have since been retired from their providers' pricing pages entirely; their prices here are the last published figures while they were commercially available.
- Open-weight pricing uses one host: DeepSeek and Moonshot sell first-party API access, so their prices are DeepSeek Platform and Moonshot's own rates respectively. Meta does not sell Llama directly; the Llama 3.1 405B price is Together AI's hosted rate, named because Meta publishes no price of its own.
- Two benchmark retirements, not one: MMLU and HumanEval saturated and were dropped between Era 1 and Era 2; SWE-bench Verified followed the same path into Era 3. A future refresh of this dataset should expect a fourth benchmark generation.
- Benchmarks measure benchmarks: none of the six evaluations on this page capture reliability, calibration, agentic tool use in production, or fit for a specific workflow.
Capability and value, not the same ranking
GPT-4o leads Era 1 on raw capability; GPT-5.6 Sol leads Era 3. Neither comes close to DeepSeek's Intelligence-per-Dollar. At the frontier, GPT-5.6 Sol, Kimi K3 and Grok 4.5 cluster within five composite points of each other on Era 3's axes — provider gaps at that level have largely closed. Below the frontier, DeepSeek's open-weight pricing is the deciding factor, not benchmark performance.