{"version":1,"asset_type":"statistical_series","data_type":"time_series","slug":"ai-model-benchmark-performance","url":"https://apiardata.com/statistics/ai-model-benchmark-performance","html_url":"https://apiardata.com/statistics/ai-model-benchmark-performance/","title":"AI Intelligence per Dollar — Frontier Model Benchmark Efficiency (2020–2026)","description":"Tracks how much AI capability a dollar buys, from GPT-3 (2020) to today's frontier. Composite capability is the plain average of three benchmark scores — swapped twice as labs retired saturated evaluations — divided by published API output-token price. Covers OpenAI, Anthropic, Google, DeepSeek, Meta Llama, Moonshot Kimi and xAI through GPT-5.6, Claude Opus 4.6, Gemini 3.1 Pro, DeepSeek V4-Flash, Kimi K3 and Grok 4.5.","domain":"papers","category":"AI & Research","keywords":["AI & Research","Benchmark Performance","Large Language Models","AI Capabilities","AI Economics"],"publisher":"OpenAI; Anthropic; Google DeepMind; DeepSeek Platform API pricing; Moonshot AI (Kimi) API pricing; xAI API pricing; Meta Llama model card (via Together AI hosting); GPT-Fathom (arXiv:2309.16583); Epoch AI Benchmark Database; Artificial Analysis Intelligence Index","frequency":"Per model release","geography":"Global — frontier commercial and open-weight models","temporal_coverage":"2020/..","last_updated":"2026-07-28","last_updated_text":"July 28, 2026","data_as_of":"Q3 2026","variable_measured":"Composite capability score (0–100, three-benchmark mean); MMLU 5-shot accuracy; HumanEval pass@1; MATH accuracy; GPQA Diamond accuracy; SWE-bench Verified resolve rate; Terminal-Bench v2.1 accuracy; Humanity's Last Exam accuracy; API output-token price","measurement_technique":"Plain unweighted mean of three published benchmark accuracy scores per era (Knowledge, Code, Reasoning axes), divided by the published USD price per 1 million output tokens at or near each model's release; the benchmark instruments changed twice as the industry retired saturated evaluations — MMLU/HumanEval to GPQA Diamond/SWE-bench Verified, then SWE-bench Verified to Terminal-Bench v2.1","license":"https://apiardata.com/data-license","is_accessible_for_free":true,"sources":[{"name":"OpenAI — API Pricing","url":"https://developers.openai.com/api/docs/pricing"},{"name":"OpenAI — Introducing GPT-4o","url":"https://openai.com/index/hello-gpt-4o/"},{"name":"Anthropic — Claude Opus 5","url":"https://www.anthropic.com/news/claude-opus-5"},{"name":"Google DeepMind — Gemini 3.1 Flash-Lite Model Card","url":"https://deepmind.google/models/model-cards/gemini-3-1-flash-lite/"},{"name":"Epoch AI — SWE-bench Verified Tracking","url":"https://epoch.ai/benchmarks/swe-bench-verified"},{"name":"Epoch AI — Benchmark Hub","url":"https://epoch.ai/benchmarks"}],"meta":[{"label":"Frequency","value":"Per model release"},{"label":"Coverage","value":"2020–2026 (16 observations across three benchmark generations)"},{"label":"Observations","value":"16"},{"label":"Geography","value":"Global — frontier commercial and open-weight models"},{"label":"Last updated","value":"July 28, 2026"}],"kpis":[{"label":"Current Frontier Composite (GPT-5.6 Sol, 2026)","value":"76.7","unit":"0–100 composite; GPQA Diamond + Terminal-Bench v2.1 + HLE, averaged","trend":{"direction":"up","value":"Highest raw composite of any model in this dataset"}},{"label":"2026 Frontier Output Price Range","value":"$6–$30","unit":"per 1M tokens; Grok 4.5 to GPT-5.6 Sol","trend":{"direction":"down","value":"Down from GPT-3's $60/1M in 2020, before adjusting for capability"}},{"label":"Peak Intelligence-per-Dollar (DeepSeek V4-Flash, 2026)","value":"218.2","unit":"composite points per $1M output tokens, Era 3","trend":{"direction":"up","value":"Highest ratio in this dataset; DeepSeek also leads Era 1 at 65.3 points per dollar"}},{"label":"Benchmark Generations Tracked","value":"3","unit":"MMLU/HumanEval/MATH → GPQA-D/SWE-bench-V/HLE → GPQA-D/Terminal-Bench-v2.1/HLE","trend":{"direction":"neutral","value":"Two retirements since 2020 as each generation saturated"}}],"series":[{"key":"era1","label":"Era 1 — MMLU + HumanEval + MATH"},{"key":"era2","label":"Era 2 — GPQA-Diamond + SWE-bench-Verified + HLE"},{"key":"era3","label":"Era 3 — GPQA-Diamond + Terminal-Bench v2.1 + HLE"}],"observation_count":13,"observations":[{"period":"2020-06","era1":0.34,"era2":null,"era3":null},{"period":"2023-03","era1":26.35,"era2":null,"era3":null},{"period":"2023-04","era1":1.09,"era2":null,"era3":null},{"period":"2024-05","era1":5.68,"era2":null,"era3":null},{"period":"2024-07","era1":23.83,"era2":null,"era3":null},{"period":"2024-12","era1":65.27,"era2":null,"era3":null},{"period":"2025-02","era1":null,"era2":3.33,"era3":null},{"period":"2026-01","era1":null,"era2":2.83,"era3":null},{"period":"2026-02","era1":null,"era2":null,"era3":5.91},{"period":"2026-03","era1":null,"era2":null,"era3":218.21},{"period":"2026-04","era1":null,"era2":null,"era3":11.95},{"period":"2026-05","era1":null,"era2":null,"era3":2.56},{"period":"2026-07","era1":null,"era2":null,"era3":5.03}],"table":{"columns":[{"key":"period","label":"Period"},{"key":"model","label":"Model"},{"key":"era","label":"Era"},{"key":"knowledge","label":"Knowledge (MMLU / GPQA-D)"},{"key":"code","label":"Code (HumanEval / SWE-bench-V / TB2.1)"},{"key":"reasoning","label":"Reasoning (MATH / HLE)"},{"key":"composite","label":"Composite"},{"key":"costPerMToken","label":"Cost / 1M Output (at release)"},{"key":"intelligencePerDollar","label":"Intelligence-per-Dollar"}],"rows":[{"period":"2026-Q3","model":"Kimi K3 (Moonshot AI)","era":"Era 3","knowledge":"93.5%","code":"88.3%","reasoning":"44.3%","composite":"75.4","costPerMToken":"$15.00","intelligencePerDollar":"5.03"},{"period":"2026-Q2","model":"GPT-5.6 Sol (OpenAI)","era":"Era 3","knowledge":"94.1%","code":"88.8%","reasoning":"47.2%","composite":"76.7","costPerMToken":"$30.00","intelligencePerDollar":"2.56"},{"period":"2026-Q2","model":"Grok 4.5 (xAI)","era":"Era 3","knowledge":"93.1%","code":"81.6%","reasoning":"40.3%","composite":"71.7","costPerMToken":"$6.00","intelligencePerDollar":"11.95"},{"period":"2026-Q1","model":"DeepSeek V4-Flash (DeepSeek)","era":"Era 3","knowledge":"89.4%","code":"61.8%","reasoning":"32.1%","composite":"61.1","costPerMToken":"$0.28","intelligencePerDollar":"218.21"},{"period":"2026-Q1","model":"Gemini 3.1 Flash-Lite (Google)","era":"Era 3","knowledge":"86.9%","code":"n/a","reasoning":"16.0%","composite":"n/a","costPerMToken":"$1.50","intelligencePerDollar":"n/a"},{"period":"2026-Q1","model":"Gemini 3.1 Pro (Google)","era":"Era 3","knowledge":"94.1%","code":"73.8%","reasoning":"44.7%","composite":"70.9","costPerMToken":"$12.00","intelligencePerDollar":"5.91"},{"period":"2026-Q1","model":"GPT-5.4 nano (OpenAI)","era":"Era 3","knowledge":"81.7%","code":"n/a","reasoning":"26.5%","composite":"n/a","costPerMToken":"$1.25","intelligencePerDollar":"n/a"},{"period":"2026-Q1","model":"Claude Opus 4.6 (Anthropic)","era":"Era 2","knowledge":"91.3%","code":"80.8%","reasoning":"40.0%","composite":"70.7","costPerMToken":"$25.00","intelligencePerDollar":"2.83"},{"period":"2025-Q1","model":"Claude 3.7 Sonnet (Anthropic)","era":"Era 2","knowledge":"78.2%","code":"63.7%","reasoning":"8.0%","composite":"50.0","costPerMToken":"$15.00","intelligencePerDollar":"3.33"},{"period":"2025-Q1","model":"DeepSeek R1 (DeepSeek)","era":"Era 2","knowledge":"71.5%","code":"49.2%","reasoning":"n/a","composite":"n/a","costPerMToken":"$2.19","intelligencePerDollar":"n/a"},{"period":"2024-Q4","model":"DeepSeek V3 (DeepSeek)","era":"Era 1","knowledge":"88.5%","code":"65.2%","reasoning":"61.6%","composite":"71.8","costPerMToken":"$1.10","intelligencePerDollar":"65.27"},{"period":"2024-Q3","model":"Llama 3.1 405B (Meta)","era":"Era 1","knowledge":"87.3%","code":"89.0%","reasoning":"73.8%","composite":"83.4","costPerMToken":"$3.50","intelligencePerDollar":"23.83"},{"period":"2024-Q2","model":"GPT-4o (OpenAI)","era":"Era 1","knowledge":"88.7%","code":"90.2%","reasoning":"76.6%","composite":"85.2","costPerMToken":"$15.00","intelligencePerDollar":"5.68"},{"period":"2023-Q1","model":"GPT-4 (OpenAI)","era":"Era 1","knowledge":"86.4%","code":"67.0%","reasoning":"42.2%","composite":"65.2","costPerMToken":"$60.00","intelligencePerDollar":"1.09"},{"period":"2023-Q1","model":"GPT-3.5-turbo (OpenAI)","era":"Era 1","knowledge":"70.0%","code":"53.9%","reasoning":"34.1%","composite":"52.7","costPerMToken":"$2.00","intelligencePerDollar":"26.35"},{"period":"2020-Q2","model":"GPT-3 (OpenAI)","era":"Era 1","knowledge":"43.9%","code":"14.0%","reasoning":"4.0%","composite":"20.6","costPerMToken":"$60.00","intelligencePerDollar":"0.34"}]},"qa":[{"question":"How is the composite capability score calculated?","answer":"Composite is the plain, unweighted average of three published benchmark accuracy scores, each covering a different capability axis: Knowledge, Code, and Reasoning. The calculation is Composite equals Knowledge plus Code plus Reasoning, divided by three. No score is normalised against another model, no ceiling or floor is applied, and no benchmark is weighted more heavily than another. Which specific benchmark measures each axis changes across the three eras on this page — Knowledge is MMLU through 2024 and GPQA Diamond from 2024 onward; Code is HumanEval through 2024, SWE-bench Verified from 2024, and Terminal-Bench v2.1 from 2026; Reasoning is MATH through 2024 and Humanity's Last Exam from 2024 onward — but the formula itself never changes. A model is only scored if all three of its era's benchmarks are published for it; a model missing one is shown with that cell marked n/a and no computed composite."},{"question":"How is Intelligence-per-Dollar calculated, and can I verify it myself?","answer":"Intelligence-per-Dollar equals Composite divided by the model's published price in US dollars per one million output tokens. Every figure needed to check a row is visible in the table: for example, DeepSeek V4-Flash's Era 3 row shows a Composite of 61.1 and a price of $0.28 per million output tokens; 61.1 divided by 0.28 is approximately 218, matching the Intelligence-per-Dollar figure shown. The ratio measures commercial value, capability per dollar of API spend, not underlying training or inference efficiency, and it uses the published output-token price only, with no adjustment for input tokens or for the reasoning tokens some models generate before their visible response."},{"question":"Why does the page use three benchmark generations instead of one?","answer":"Because no single benchmark trio has been reported consistently across this dataset's full 2020–2026 span. MMLU, HumanEval and MATH defined AI capability measurement through 2024, but labs stopped publishing them once frontier models approached each benchmark's ceiling. GPQA Diamond, SWE-bench Verified and Humanity's Last Exam took over for 2024–2026, until SWE-bench Verified itself was retired in early 2026 in favour of Terminal-Bench v2.1. Rather than force incompatible eras into one line or quietly drop the earlier ones, this page tracks all three eras separately, using the identical formula and the same three capability axes, Knowledge, Code, Reasoning, throughout. Two models, Gemini 3.1 Pro and DeepSeek V4-Flash, published scores under both the second and third benchmark generations' code axis; the table and chart show each of them once, under its more current Era 3 figure, rather than as two separate rows."},{"question":"Why were MMLU and HumanEval dropped, and is SWE-bench Verified really gone too?","answer":"Both benchmarks saturated. MMLU's accuracy ceiling is estimated at roughly 93 percent because of labelling errors in the original dataset, and frontier models were effectively at that limit by 2025; HumanEval's original 164-problem test set stopped distinguishing frontier models once pass rates above 95 percent became common, and several providers have been found to include similar problems in training data. Anthropic's Claude 3.x and 4.x model cards stopped reporting HumanEval well before either of those milestones. SWE-bench Verified followed a similar path: OpenAI publicly stated in February 2026 that the benchmark no longer measured frontier coding progress, and none of the frontier models released in the months since, including GPT-5.6, Grok 4.5, and Kimi K3, have a published SWE-bench Verified score anywhere, including from independent evaluators. Terminal-Bench v2.1 has taken its place for that generation of models."},{"question":"Why does a lower composite score in Era 2 or Era 3 not mean a less capable model than a higher Era 1 score?","answer":"Because the three eras use benchmarks built to very different difficulty targets, and the composite formula does not adjust for that difference. Era 1's benchmarks, MMLU, HumanEval and MATH, are now considered largely saturated, with frontier models scoring close to each one's ceiling by 2024. GPQA Diamond, used in Era 2 and Era 3, is designed so that PhD-level experts working outside their own specialty score only 65 to 74 percent on it. Humanity's Last Exam was built to sit at the edge of expert human knowledge; even GPT-5.6 Sol, this dataset's highest Era 3 composite at 76.7, scores just 47.2 percent on that one axis. SWE-bench Verified and Terminal-Bench v2.1 both require multi-step work against real code repositories rather than the isolated single-function problems HumanEval used. None of that difficulty gap is built into the composite formula, which stays a plain three-way average in every era on purpose, so a reader can recompute any row from the table alone without trusting a hidden weighting. The practical result is that GPT-4o's Era 1 composite of 85.2 and Claude 3.7 Sonnet's Era 2 composite of 50.0 are not directly comparable numbers. They were produced by different, harder-to-satisfy yardsticks, and the drop reflects a tougher benchmark, not a weaker model."},{"question":"Which price is used for each model, and why do retired models still have a price?","answer":"Every model is priced at its published API rate in US dollars per one million output tokens, at or as close as possible to its release date, not at today's price, which for several older models no longer exists. GPT-3, GPT-3.5-turbo, and GPT-4 have all been withdrawn from their provider's current pricing pages; the figures used here are the prices those models carried while commercially available. Open-weight models are priced at the rate their developer sells directly where one exists, DeepSeek's own platform for DeepSeek models, Moonshot's own platform for Kimi models. Meta does not sell API access to Llama directly, so the Llama 3.1 405B figure is Together AI's hosted rate for that model, named explicitly because it is a third-party price rather than a first-party one."},{"question":"Why do reasoning models likely cost more in practice than their Intelligence-per-Dollar ratio suggests?","answer":"Most models in the Era 2 and Era 3 benchmark generations generate reasoning tokens, an internal chain-of-thought stream that precedes the visible response, and these are typically billed at the same rate as ordinary output tokens. A demanding query can produce several times more reasoning tokens than visible output tokens, and the volume is not directly controllable by the caller; it scales with how difficult the model judges the task to be. This page's Intelligence-per-Dollar ratio uses only the published per-token price, with no adjustment for reasoning-token volume, on the same basis for every model. That makes the ratio a fair comparison of headline pricing, but it means reasoning models will typically look more cost-efficient here than they turn out to be on genuinely hard, demanding production workloads. Teams evaluating a reasoning model for a specific use case should measure actual token consumption on their own query distribution rather than relying on this ratio alone."},{"question":"Are there any pricing details that could make cross-model price comparisons misleading?","answer":"Three specific issues are worth flagging. Anthropic's tokenizer changed starting with Claude Opus 4.7; the company states the same text now produces roughly 30 percent more tokens than it did on Claude Opus 4.6 and earlier models, so per-token prices for Claude models released before and after that change are not directly comparable on a per-word basis, even when the dollar figures look similar. Some prices are introductory and scheduled to change on a fixed future date; DeepSeek V3 is an example from this dataset itself, launched in December 2024 at a promotional $0.28 per million output tokens that lasted about six weeks before reverting to its standing $1.10 rate on 2025-02-08, which is the figure used here. More generally, a price shown on this page reflects the rate in effect at the time of the most recent data refresh, and buyers should check the provider's current pricing page before making a purchasing decision, since API prices change more often than benchmark scores do."}]}