MST0052 · Lecture 2 · Extra
ChatGPT
a conversation with a model
o1 and reasoning
more computation before answering
Coding agents
Claude Code and Codex
Fable 5.1 and GPT-6 Astra
work across many steps
The figure shows results relative to a human baseline. Source: Stanford HAI, AI Index 2026, ch. 2 (CC BY-ND 4.0).
Source: METR and TH1.1. Own figure of METR's published estimates; bars show 95% intervals.
A fixed workflow follows predetermined steps. An agent can choose the next step based on intermediate results.
Source: Anthropic, Building effective agents; MCP: Tools. The economist example is an illustration.
Cheapest model reaching a fixed test level. Selected series from Epoch AI (through February 2025).
What is frontier level today soon becomes cheap – and possible to run on your own machine.
Source: Epoch AI, LLM inference price trends (March 2025) and Frontier AI capabilities can be run at home (Aug. 2025).
Artificial Analysis Intelligence Index v4.3
Max setting.
Open weights: can be downloaded and self-hosted, but require substantial hardware.
Source: Artificial Analysis; Edwards and Emberson, Epoch AI, Open models lag state-of-the-art closed models by 4 months.
The figure shows 2021–2025 for two occupational groups. The number on the right comes from a broader analysis through June 2026.
Source: Stanford HAI, AI Index 2026 (CC BY-ND 4.0); Brynjolfsson, Chandar & Chen, Canaries update. US observational data.
“OpenAI Says It Has Cracked One of Math’s ‘Millennium Problems’”
Read the article ↗
“On the Navier–Stokes Millennium Prize Problem”: the company’s own account of the result
Read OpenAI’s write-up ↗