Loading...

AI Capability Is Doubling Every Few Months. Your Planning Cycle Isn't.

"Evaluations intended to be challenging for years are saturated in months."

— Stanford HAI, 2026 AI Index Report

Ask three different research groups how fast AI is improving and you will get three different numbers, but they all point the same direction: fast, and getting faster. METR's latest time-horizon data puts the doubling time for AI task capability at roughly four to seven months depending on which window you measure. The UK AI Security Institute (AISI) independently clocked cyber-task capability doubling every 4.7 months as of February 2026, down from an 8-month estimate just three months earlier. Meanwhile, benchmarks built to stay hard for years are being maxed out in months. If you are planning an AI rollout around what a model can or cannot do today, the ground is moving faster than most roadmaps assume.

The Doubling-Time Data, Not the Hype

The most rigorous public measurement of AI progress speed comes from METR, a research group that tracks the length of real-world tasks AI agents can complete at 50% reliability. In its original 2025 analysis, that task-length frontier was doubling roughly every 7 months, a trend METR said had held for about six years. In January 2026, METR released an updated version of its benchmark, Time Horizon 1.1, with a larger and more carefully calibrated task suite, 228 tasks instead of 170, including a jump from 14 to 31 tasks that take 8 or more hours to complete. Under the revised methodology, the post-2023 doubling time came in at 130.8 days (about 4.3 months), roughly 20% faster than the original 165-day estimate. Measured from 2024 onward only, it drops further, to 88.6 days, a little under 3 months.

AISI ran a parallel, independently constructed measurement focused specifically on offensive cyber tasks, and arrived in the same neighborhood. In February 2026, AISI estimated the task-length frontier for autonomous cyber capability had doubled every 4.7 months since late 2024, up from its own November 2025 estimate of 8 months. AISI was careful to flag that the newest models in its sample "substantially exceeded" that trend, and that it remains unclear whether this is a genuine acceleration or a one-time jump from a few unusually capable releases. Both research groups land on the same honest caveat: the exponential fits are real, useful signals, not laws of nature, and a handful of outlier models can swing the estimate meaningfully.

What "Doubling Time" Actually Means

A roughly 4-to-5-month doubling time means that, if the trend holds, the length of a task an AI agent can reliably complete today becomes twice as long again in less than half a year. That compounds fast: four doublings in under two years is a 16x jump in autonomous task length, not a steady, linear improvement you can schedule a single training program or governance review around once and move on.

Benchmarks Are Becoming the Bottleneck, Not the Models

The 2026 AI Index Report from Stanford's Institute for Human-Centered AI (HAI) makes a related and arguably more uncomfortable point: the tools used to measure AI progress are failing to keep pace with the progress itself. Frontier models gained 30 percentage points in a single year on Humanity's Last Exam, a benchmark explicitly built to be hard for AI systems for years to come. On OSWorld, a benchmark for AI agents operating a real computer interface, accuracy rose from roughly 12% to 66.3% in 2025, closing in on the human baseline of 72.35%. The report's own framing is blunt: "evaluations intended to be challenging for years are saturated in months." We have written before about what this means practically for autonomous AI agents inside the enterprise, where the benchmark a vendor quoted six months ago may already understate what the same model can now do.

There is a second, less comfortable layer to this: the benchmarks themselves are not fully trustworthy. The AI Index found invalid-question rates ranging from 2% on MMLU Math to 42% on GSM8K, meaning a meaningful share of the questions used to score "AI progress" are flawed or ambiguous. Separate research cited in the same report suggests that leaderboard rankings on platforms like Chatbot Arena may partly reflect models adapting to that specific platform's quirks rather than general capability gains. Put together, this means two things can be true at once: AI capability really is advancing at a doubling-time pace of a few months, and the specific number you read off any single leaderboard should be treated with real skepticism.

A researcher in an AI lab looking at a steep exponential growth curve on a display, illustrating the accelerating pace of AI capability improvement

The Jagged Frontier: Fast Does Not Mean Even

The pace of improvement is not uniform across tasks, and that unevenness matters as much as the speed itself. Google's Gemini Deep Think won gold at the 2025 International Mathematical Olympiad with 35 points, up from a silver-medal 28 points in 2024, a genuine jump in formal mathematical reasoning. Yet on ClockBench, a benchmark that simply asks a model to read the time on an analog clock face, the top model managed only 50.6% accuracy against a 90.1% human baseline. The AI Index describes this as a "jagged frontier," and it is visible elsewhere too: agents went from roughly 12% to 66% success on OSWorld's computer-use tasks, yet still fail about one in three attempts on structured benchmarks overall. This unevenness is exactly why organizations evaluating where an agentic AI loop is safe to deploy cannot rely on a single capability score; a model that is superhuman at competition math can still misread a clock face or fail a basic structured task a third of the time.

Competitive dynamics between frontier labs add another layer of volatility to the pace story. DeepSeek-R1 briefly matched the top US model's performance in February 2025, and as of March 2026 the AI Index measured the leading US model ahead of the leading Chinese model by just 2.7% on its composite scoring, with the gap fluctuating in the single digits over the preceding year. That gap closing and reopening within a single-digit band, rather than one country pulling decisively ahead, is itself a symptom of how fast the frontier moves: a model that leads by a few points in March can be matched or passed by autumn.

The Planning Problem This Creates

If task-length capability doubles every 4 to 7 months, and the specific benchmark numbers used to justify that estimate come with real error bars and gaming concerns, then any enterprise AI plan written around "what the model can do today" has a shelf life measured in months, not years. We have made a related case for why the corporate world needs to prepare for capability jumps, not just incremental gains, and the doubling-time data gives that argument a concrete number to plan against.

What This Means for How You Plan Around AI

None of this is a reason to chase every headline capability claim, and the researchers behind these numbers are the first to say so. METR has repeatedly flagged that its own trend line could be off by a meaningful margin given how sensitive doubling-time estimates are to task selection and human-baseline choices, and AISI explicitly says it is too early to tell whether its newest, faster reading represents a durable acceleration or a temporary jump from one or two standout models. The responsible read of this data is not "AI will be able to do anything in two years," it is "whatever capability ceiling you are planning against today will likely be wrong within a single budget cycle." That has concrete implications: governance and risk reviews built around current model limitations need a refresh cadence closer to quarterly than annual, and any AI deployment decision gated on "the model cannot reliably do X yet" needs an explicit re-test date rather than a one-time sign-off. We cover the governance side of that discipline in more depth in our piece on generative AI ethics and risk management, which is really a planning-cadence problem as much as an ethics one once you accept that the underlying capability is not static.

The honest summary is that the pace of AI evolution is real, measurable, and faster than most planning cycles, but it is also uneven, hard to pin to a single number, and actively outrunning the benchmarks built to track it. Build your AI strategy around that volatility rather than around a snapshot of what today's model can or cannot do, and you will be adjusting on your own terms instead of being caught by the next doubling.

Share This Article