AI adoption, in numbers.
Each dimension is a short data essay from a public source — with the chart, the story behind it, and the open math for how the number was reached. Pick a dimension to explore.
Pick a dimension…
Capabilities
From text to work: what AI can already do
To match GPT-4 on GPQA Diamond, the output price per token fell from US$60 to US$2.19 per million — same test, nearly 30× cheaper, in under two years.
In little over three years, models jumped from 'completing sentences' to solving graduate-level problems, closing code issues on their own and sustaining tens-of-billions companies. This dimension measures that jump through three lenses — what they know, what they can execute, and how much money it already moves — and ends on the exam built to be unbeatable.
Cognitive capability — by year and by price
Each dot is a frontier model: when it shipped (horizontal), how much it scored on GPQA Diamond — 'Google-proof' graduate questions (vertical) — and which price bracket it plays in (color). The ceiling rises while the lighter dots show the same intelligence getting cheap: DeepSeek-R1 delivered 71% at US$2 per million tokens.
- 2023-03GPT-435.7%US$60/M
- 2024-03Claude 3 Opus50.4%US$75/M
- 2024-05GPT-4o53.6%US$10/M
- 2024-06Claude 3.5 Sonnet59.4%US$15/M
- 2024-12o178%US$60/M
- 2025-02Claude 3.7 Sonnet78.2%US$15/M
- 2025-03Gemini 2.5 Pro84%US$10/M
- 2025-08GPT-588.4%US$10/M
- 2025-11Gemini 3 Pro91.9%US$12/M
- 2026-02Gemini 3.1 Pro94.1%US$12/M
What this chart concludes
The vertical axis is GPQA Diamond accuracy (100% = perfect). In 2023 GPT-4 scored 36% at US$60/M tokens; DeepSeek-R1 soon hit 71% — double — at US$2.19. Hence the headline: beating GPT-4 got 96% cheaper. Capability up, price down.
- Method
- GPQA Diamond (198 questões de pós-graduação, 'à prova de Google'), score oficial pass@1 na data de lançamento. Preço = tabela de saída do provedor (US$/M tokens); cor = faixa de preço, linha = recorde acumulado.
- Source
- Epoch AI — GPQA Diamond
- Snapshot from
- July 05, 2026
Agentic capability — solving, not just answering
Answering well is one thing; executing a task end-to-end is another. SWE-bench Verified measures exactly that: 500 real GitHub issues the model must resolve by editing the repo until tests pass. In under two years, the frontier went from a third to almost everything.
- 2024-08GPT-4o33%
- 2024-103.5 Sonnet49%
- 2025-023.7 Sonnet62.3%
- 2025-05Opus 472.5%
- 2025-08GPT-574.9%
- 2025-11Gemini 376.2%
- 2025-11Opus 4.580.9%
- 2026-04Opus 4.787.6%
- 2026-06Fable 5 👑95%
Claude Fable 5 across three agentic suites today
- SWE-bench Verifiedcódigo de ponta a ponta95
- Terminal-Bench 2.1tarefas no terminal88
- τ-benchatendimento com ferramentas89.2
- Method
- SWE-bench Verified: 500 issues reais do GitHub resolvidas de ponta a ponta (editar o repo até os testes passarem) — o padrão de capacidade agêntica de código. Escada = recorde por data; barras = líder atual em três suítes. Nota: em 2026 um estudo de Berkeley mostrou que essas suítes podem ser 'gameadas' — leia como tendência.
- Source
- SWE-bench (leaderboard oficial)
- Snapshot from
- July 05, 2026
The economics behind it — who earns, and how much
Capability turned into revenue at an unprecedented pace. OpenAI opened the decade ahead, but in 2026 Anthropic crossed the curve — pulled by enterprise coding usage. Figures are annualized run-rate (the month times twelve), reported by the companies and the press.
Valuation, latest round
The end of the ruler
Humanity's Last Exam — the test built to be unbeatable
Once models started acing tests like MMLU, benchmarks lost their point: you could no longer tell who was ahead. The answer, launched in January 2025 by the Center for AI Safety with Scale AI, was to assemble the hardest exam humanity could write.
How it's made
- 1About a thousand experts from 500+ institutions submitted over 70,000 questions in their fields — from topology to virology.
- 2A question only makes it in if the best models of the day get it wrong: easy ones are dropped on the spot, an adversarial filter against the frontier itself.
- 3Survivors go through two rounds of peer review until 3,000 remain — 2,500 public and 500 secret, to catch models that 'memorized' the test.
- 4They're closed-ended (76% exact short answers, 24% multiple choice; ~14% need an image), so they can be graded automatically and compared fairly.
Best score over time (%)
- 2025-01o18%
- 2025-04o317%
- 2025-08GPT-525.3%
- 2025-11Gemini 3 Pro37.5%
- 2026-02Gemini 3.1 Pro45%
- 2026-06Fable 553.3%
The first models got 3% to 8% right. Eighteen months later, the frontier passed the halfway mark — on an exam designed not to be beaten any time soon. That's the pace this dimension tries to capture.
- Method
- Center for AI Safety (CAIS) + Scale AI. 3,000 questões (2,500 públicas + 500 secretas) em 100+ disciplinas, filtradas contra a fronteira. Lançado em 24/01/2025.
- Source
- Humanity's Last Exam (CAIS + Scale AI)
- Snapshot from
- July 05, 2026