01 / 04
A benchmark dies the day a frontier model scores past 90% on it, and the field's flagship exam took 48 months to get there. The 57-subject knowledge exam released in September 2020 opened with the best model at 43.9%, far below the 89.8% that human experts average. In September 2024 a frontier model posted 92.3%, above the experts and inside the test's own error, so the score stopped separating models.
02 / 04
The two big tests that came after it died in 17 and 24 months, respectively. The 2021 test held grade-school math word problems, and the best setup in its launch paper solved about 55% of them, yet GPT-4 posted 92% seventeen months later. The 2023 test was built so answers cannot be found by searching, i.e. skilled non-experts with full web access managed just 34%, and it still fell in November 2025 when a frontier model posted 91.9%.
03 / 04
Once you line the three lifespans up from the same starting line, you will notice everything after the 2020 exam fell at or within two years. Harder questions did not buy more time, and that is because difficulty only lowers the score a test starts at, while the speed of the climb is set by model progress. Basically, a harder exam starts from a lower score, but it does not stay difficult for longer.
04 / 04
So, what is the key takeaway? A benchmark score is a dated measurement, so the number means little without knowing when it was recorded. It is better to treat any public score more than a year old as expired. The 2025 exam was assembled by subject experts to be the final one, and frontier models opened it in single digits, reached 37.5% by month 10, and sit near half at month 19, right on the two-year pace. So for anything you ship, the yardstick that keeps working is a private eval on your own tasks, refreshed as models move.