Syed Tanveer Jishan.

Why AI benchmarks that lasted four years now die in two years

A public benchmark stops working the day frontier models score past 90% on it, because beyond that line the test no longer tells models apart. The flagship 2020 exam took four years to get there, and the two big tests that followed died in 17 and 24 months.

Published ·Updated

01 / 04

A benchmark dies the day a frontier model scores past 90% on it, and the field's flagship exam took 48 months to get there. The 57-subject knowledge exam released in September 2020 opened with the best model at 43.9%, far below the 89.8% that human experts average. In September 2024 a frontier model posted 92.3%, above the experts and inside the test's own error, so the score stopped separating models.

02 / 04

The two big tests that came after it died in 17 and 24 months, respectively. The 2021 test held grade-school math word problems, and the best setup in its launch paper solved about 55% of them, yet GPT-4 posted 92% seventeen months later. The 2023 test was built so answers cannot be found by searching, i.e. skilled non-experts with full web access managed just 34%, and it still fell in November 2025 when a frontier model posted 91.9%.

03 / 04

Once you line the three lifespans up from the same starting line, you will notice everything after the 2020 exam fell at or within two years. Harder questions did not buy more time, and that is because difficulty only lowers the score a test starts at, while the speed of the climb is set by model progress. Basically, a harder exam starts from a lower score, but it does not stay difficult for longer.

04 / 04

So, what is the key takeaway? A benchmark score is a dated measurement, so the number means little without knowing when it was recorded. It is better to treat any public score more than a year old as expired. The 2025 exam was assembled by subject experts to be the final one, and frontier models opened it in single digits, reached 37.5% by month 10, and sit near half at month 19, right on the two-year pace. So for anything you ship, the yardstick that keeps working is a private eval on your own tasks, refreshed as models move.

Sources and method

A test is counted as dead the first time a frontier model reports 90% or better on it under standard prompting, and its lifespan is the months from the test's release to that report. One earlier 90.04% MMLU claim from December 2023 used a nonstandard 32-sample voting method, so the count starts from the first standard-methodology pass in September 2024. The curves connect each test's published start and end scores, and the month-by-month shape between those points is drawn to read clearly. The open bar for the 2025 exam shows its age so far, and the dashed extension to the two-year line is a projection, marked apart from the reported data.