Intelligence Was Rarely the Bottleneck

AI makes reasoning cheap. Scarcity moves.

In August, OpenAI published ten results in mathematics and theoretical computer science produced by an internal version of Astra, then unreleased. Each resolved or made substantial progress on a long-standing open problem, across fields from high-dimensional geometry to coding theory to group theory. The model subsequently formalized the arguments in Lean. OpenAI said the tokens needed to find all ten solutions would have cost roughly $2,000 at GPT-5.6 Sol API rates.

That number is striking, and it is easy to misread. It covers the marginal token cost of finding these ten successful results. It is not the cost of training the model, and it is no guide to the price of solving the next arbitrary open problem. Even so, it represents something new: mathematical reasoning that once depended on scarce expert time can now, in some cases, be purchased as compute.

Then, while I was writing this essay, the ceiling moved again.

On September 8, OpenAI released a proposed resolution of the Navier-Stokes existence and smoothness problem, one of the Millennium Prize Problems. The result was produced by an internal model that OpenAI says is substantially more capable than Astra. The group that found the solution involved on the order of 10,000 concurrent agents. The agents reached the result roughly 88 hours after the effort began, using about 130 billion output tokens on Navier-Stokes alone; Lean formalization and verification took another 17 hours using Astra. Three days later, the Clay Mathematics Institute said the problem had “apparently been settled,” while emphasizing that its evaluation process is deliberately unhurried.

I am not going to argue with the significance of this. A class of intellectual work that used to be measured in expert-years can now have a compute budget.

But notice what happened next. One effort cost roughly $2,000 in tokens; the other used industrial-scale parallelism and an unreleased model. The observations are not directly comparable, and turning them into a single “cost of discovery” would be false precision. What they do show is that once reasoning capacity expands, other scarce inputs become visible very quickly.

Mathematics is not the easy case intellectually. It is the unusually clean case economically.

Nothing has to be grown, dosed, manufactured, recruited or approved. A proof can ultimately be adjudicated inside mathematics itself. And once a proof has been formalized, a system such as Lean provides a remarkably cheap and deterministic layer of checking. Producing the formalization is not free, and Lean cannot tell us whether the theorem is important or whether its formal statement perfectly captures what mathematicians intended. But it can verify that the proof follows from its formal premises without asking another mathematician to reconstruct the whole argument from scratch.

Where I part ways with some of the reaction to these releases is the next inference: if intelligence is becoming this cheap, science itself is about to become cheap.

Science is not one production function. In formal disciplines, reasoning can sit close to the binding constraint. In empirical disciplines, reasoning is one input in a longer chain that also contains evidence, experiments, capital, physical execution, permission and time.

Make one of those inputs radically cheaper and the others do not disappear.

The bottleneck moves.

#The bottleneck test

A frontier model changes an economic system to the extent that it relaxes a constraint that was actually binding.

That sounds obvious when written down. It is not how model releases are usually interpreted. We see a capability improve by an order of magnitude and instinctively project something like that improvement onto the output of the entire system. That projection only works when the improved capability was what held the system back.

I find three questions useful.

First, what is actually binding? It may be cognition. It may instead be capital, experimental throughput, physical capacity, licensure, trust, procurement or institutional authority.

Second, what does it cost to check the answer? Delegation depends not only on the cost of producing work, but also on the cost of deciding whether that work is safe and good enough to use. If a skilled human has to redo the entire task to verify it, cheap generation creates much less usable capacity than the headline capability suggests.

Third, what becomes scarce next? Once one input gets cheaper, activity expands until another input starts binding. AI does not abolish scarcity. It changes its location.

Mathematics does well on all three tests. Reasoning sits near the center of the production function, and formal verification can make checking much cheaper than reproduction. The ten August results and the Navier-Stokes effort are radically different in scale, but both show the same thing: mathematical reasoning that was once extremely scarce can now be deployed in quantities large enough to change what gets attempted.

That is a real repricing.

#The empirical world

Now apply the same test to drug development, where some of the strongest claims about AI and science are being made.

One easy argument is that drug development is enormously expensive while molecular design represents only part of the bill, so making molecular design cheaper cannot matter very much. I do not think that argument works. Cost share is not the same thing as constraint. A relatively inexpensive upstream decision about which biological mechanism to pursue can determine whether a vastly more expensive downstream program succeeds or fails.

The deeper distinction is about the kind of uncertainty involved.

A large share of the remaining uncertainty in drug development is empirical rather than purely inferential. Does perturbing this target actually alter human disease? Does an effect in a cell survive in an animal? Does it survive in a heterogeneous patient population? What dose produces sufficient exposure without unacceptable toxicity? Which apparently causal relationship disappears once biology encounters an actual human body?

AI may reason about those questions extraordinarily well. It can extract more signal from existing data, propose mechanisms, design molecules, identify patterns humans missed and choose better experiments.

What it cannot do is observe an outcome that the world has not yet produced.

At some point nature has to answer.

That distinction matters because AI can make reasoning over existing information much cheaper without making new causal information about the physical world equally cheap.

The empirical record so far is broadly consistent with that distinction. A 2026 Nature Reviews Drug Discovery assessment concluded that evidence of AI’s clinically relevant impact remains “disappointingly limited.” The authors’ claim was specific: impressive performance on models and proxy endpoints has not yet translated into comparable evidence that patients receive safer or more effective medicines faster.

There is encouraging evidence earlier in the pipeline. A 2024 analysis of drugs discovered by AI-native biotechnology companies found Phase I success rates of roughly 80 to 90 percent, substantially above historical averages. That is important evidence that AI can help produce molecules with useful drug-like properties.

The Nature Reviews authors read that figure more cautiously. They point out that most of these programs pursue established disease biology and chemistry, which lowers the chance of hitting the safety problems that sink a first-in-class candidate.

But the same study found Phase II success of roughly 40 percent, on a much smaller sample, approximately in line with historical rates. Phase II is where proof of efficacy in actual patients becomes much more central.

Better molecules matter. Better therapies matter more.

Generating them is not the same problem.

#When abundance creates a selection problem

This is where I think some forecasts of scientific acceleration skip a step.

Cheap cognition increases the supply of hypotheses. Experimental throughput does not automatically increase with it.

Imagine a research group that could previously generate one hundred plausible hypotheses and afford to test twenty. Now AI can generate ten thousand. That larger candidate pool can be valuable. If the ranking system contains useful signal, having more candidates increases the chance that something unusually good is available to select.

But the economic value of additional generation now depends increasingly on selection.

The scarce capability becomes knowing which twenty of the ten thousand deserve an experiment.

That can itself become an AI problem. If a model becomes so good at target selection or experiment prioritization that one experiment delivers the information that previously required ten, it has effectively expanded experimental capacity without building another laboratory.

That would be transformative.

The difficulty is that reliable ranking eventually has to close the loop through outcomes. In biology those outcomes are expensive, conditional and noisy. Human biology does not provide clean labels. Negative results are often missing from public datasets or locked inside companies; when chemists recovered failed reactions from their notebooks, a model trained on them beat one trained on the published literature. Early-stage proxies may correlate only weakly with the clinical endpoint we ultimately care about.

So the constraint can move from hypothesis generation to hypothesis selection, and from hypothesis selection to the generation of sufficiently good ground truth to improve and validate the selection system.

A model may become extraordinarily good at telling us what to test next. The test still has to be run.

And the world runs on its own clock.

There are at least three ways that clock could speed up enough to weaken this argument.

The first is better ranking. If AI prospectively and reproducibly improves target and mechanism selection enough to create a large Phase II or Phase III success-rate advantage, fewer experiments will be required to produce each successful medicine. In that case AI would be relaxing the empirical constraint indirectly, by wasting less experimental capacity.

The second is automation. Closed-loop laboratories can increase the rate at which reality returns answers. Lila Sciences has raised $550 million in total capital to build what it calls AI Science Factories. AstraZeneca’s iLab is explicitly attempting to automate the design-make-test-analyze cycle of medicinal chemistry.

I see those investments as evidence for the bottleneck argument, not against it. Once reasoning and design become abundant, enormous value attaches to making the physical loop faster.

The experience of Berkeley’s A-Lab is instructive. The autonomous laboratory was a genuine technical achievement, but a subsequent Nature correction clarified its novelty claims and required manual reanalysis of experimental results; four previously reported successes became inconclusive. The lesson is that closing more of the experimental loop does not automatically close the interpretive one.

The third is regulation. If sufficiently validated computational evidence eventually substitutes for categories of experimental or clinical evidence, development could accelerate dramatically. In that case the trial itself would not have become physically faster. The evidentiary requirement would have changed, because a different source of information became trusted.

Any of these developments could move the bottleneck again. That is what the framework predicts.

#Why sectors are the wrong unit

This is why I increasingly think that asking “which sectors will AI disrupt?” is too coarse a question.

What matters is constraint structure, and that can vary more between workflows inside an industry than between industries.

I have seen this in medicines regulation in Africa. A regulator sounds like the last institution one would put on a list of organizations primed for AI transformation. The work is conservative, procedurally constrained and high stakes.

But look at the production function.

A substantial part of regulatory work consists of reviewing structured documentation against explicit requirements, identifying gaps, comparing evidence with standards and determining which questions need to be answered. Qualified reviewers are scarce. AI can perform a first pass over the documentation and surface the issues. A human reviewer still exercises regulatory judgment and remains accountable for the decision, but checking a well-structured analysis may take substantially less time than producing it from scratch.

In that workflow, cognition and reviewer throughput can genuinely be binding.

Now take a maternal-health product. In global health, part of my work has involved asking how much R&D investment is required to produce an additional unit of health impact. Suppose tomorrow the intellectual act of identifying a promising product became essentially free.

The expected R&D cost might fall less than the capability headline suggests if much of the remaining cost sits in preclinical validation, clinical trials, regulatory work, manufacturing development, scale-up and failed programs along the way.

If instead we care about the total cost of producing health impact, another set of constraints appears: procurement, distribution, training, stocking, adoption and continued use.

The health product looks scientific, while the regulator looks bureaucratic. Yet some regulatory workflows may accelerate faster because their scarce input is precisely the kind of cognition AI makes cheaper.

The same pattern appears almost everywhere. Organizations whose primary purpose is physical still contain large amounts of screen-based cognitive work: claims review, freight documentation, procurement, invoicing, compliance filings and license renewal. Much of that work operates against written rules and produces digital outputs that can be checked.

The industry label tells us less than the workflow.

The better questions are whether the output can be produced digitally, how expensive it is to verify, how reversible mistakes are, and what physical or institutional authority remains necessary after the cognitive work is done.

#The wrong denominator

There is another mistake in how frontier models get priced.

GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, against $4 and $20 for GPT-5.6 Sol. Look only at the price sheet and intelligence appears to have become more expensive.

But tokens are an engineering unit, not an economic outcome.

On Terminal-Bench 4.0, for example, OpenAI reports that Astra completes 57.9 percent of tasks compared with 37.3 percent for Sol, while costing about 9 percent less per task at the tested settings, despite charging 2.5 times as much per token. Independent results from Artificial Analysis make the same broader point: several Astra reasoning settings sit on its intelligence-versus-cost frontier.

Astra will not be cheaper for every workload. But the price sheet cannot tell you which ones.

The relevant quantity is closer to cost per successfully completed task. In a real production system it is closer still to cost per accepted outcome, including verification, correction and the failures discarded along the way.

A model that costs twice as much per token but needs one-third as many tokens, retries or human corrections may be cheaper. A model that generates a brilliant answer that a senior expert needs two hours to validate may be far more expensive than its API bill suggests.

Once you use that denominator, the economic ledger becomes clearer.

The cost of some forms of cognitive production falls sharply. The marginal value of complementary inputs can rise: expert verification, experimental capacity, proprietary data, institutional authority, access to patients, manufacturing capacity and physical execution.

Other constraints may change much more slowly. A factory still has finite throughput. A clinical study still needs participants and follow-up. A medicine still has to be manufactured and distributed. A regulator still has to assume legal responsibility for a decision.

AI can reduce friction around all of those things, sometimes dramatically. But eliminating friction is different from eliminating the constraint itself.

#A dated claim

Predictions that cannot lose are entertainment, so here is one with a date attached.

By September 2027, I expect the strongest demonstrated productivity gains from frontier AI to remain concentrated in workflows where much of the output is digital, mistakes can be corrected, and verification is relatively cheap. Software tasks with automated tests are an obvious example. Formal mathematics is another. So are bounded document and analytical workflows in which outputs can be checked against explicit rules.

Even there, the gains are less settled than the discourse suggests. A randomized trial by METR found sixteen experienced open-source developers completing real tasks 19 percent slower with early-2025 AI tools, while believing afterwards that they had been 20 percent faster. Sixteen developers do not settle the question, and the tools have moved since. But perceived speedup and measured speedup are different quantities, and the denominator problem turns up even in the workflow everyone treats as the easy case.

I do not expect similarly large gains merely because a field requires a great deal of intelligence.

In drug development, I expect continued evidence that AI improves molecular design, preclinical workflows and some early clinical outcomes. I do not expect that, by September 2027, AI will have produced a large and repeatable Phase II or Phase III efficacy advantage of the kind that would establish that the central problem of clinical translation has been solved.

If that advantage appears, I will have underestimated how much clinically useful causal information was already latent in existing data.

If autonomous experimentation dramatically increases the rate at which reliable new ground truth is produced, I will have underestimated how quickly the physical constraint could move.

And if workflows dependent on slow, expensive physical verification begin matching the productivity gains of cheaply verifiable digital work without corresponding improvements in ranking, experimentation or evidentiary standards, then the framework itself will need revisiting.

That is where I would put the claim today.

The most important economic consequence of abundant intelligence may not be that everything suddenly accelerates. It may be that every system begins to reveal what it was actually waiting for.

Which is a reason to take frontier AI more seriously, and to ask what happens after intelligence becomes cheap.

The next scarce input is where the value will move.

Intelligence was a bottleneck.

It was rarely the only one.