01 / 04
The top 10% of users on a flat AI plan burn 70% of the tokens and pay 10% of the revenue. The straight line is what subscribers pay, and the curve far below it is what they actually burn. What you will notice is that by the 90% mark subscribers have put in 90 cents of every revenue dollar and burned only 30% of the tokens. So the shaded gap between the lines is the light users covering the heavy ones.
02 / 04
The plans were priced for a chat era, where the curve sat close enough to the line that one flat fee fit everybody. That chat era is the solid curve, and the dashed line is the same curve once agents arrive. Basically that means agents drag the curve far below the line, because one agentic coding task burns about 1,000 times the tokens of a chat exchange. The agent re-reads its whole history at every step and loops whenever a step fails.
03 / 04
Agents pulled that curve back down, and within three weeks of April 2026 four of the biggest AI coding vendors broke the flat plan the same way. OpenAI moved its coding agent to pay-as-you-go, Anthropic cut third-party agents off flat plans and pulled Claude Code from its $20 tier, GitHub froze Copilot Pro signups, and Cursor put frontier models behind a paid mode. So the paying line bends toward the burning curve, i.e. a flat base with a meter on top.
04 / 04
So, what is the solution? Give the price the same shape as the usage. Your goal is to keep a flat base that covers the median user, and then start a meter where the heavy tail begins, around the 90th percentile. This way nine users in ten never meet a meter, and the consumption-heavy users go from paying 10% of the revenue to 42%. To put it simply, the number to engineer is the percentile where the meter should start.