Part 3 of Intelligence Explosion: Evidence, Bottlenecks, and Control. Research checked October 4, 2026.
Twenty million researchers would be a startling resource. Twenty million researchers waiting for the same experiment would be a startling queue.
The CASP paper’s supplementary workforce calculation captures the first possibility. It divides an estimated daily token capacity by a benchmark-derived estimate of tokens per researcher-day. The resulting order of magnitude is around twenty million researcher equivalents, conditional on expert-level systems operating at comparable runtime costs. This is a capacity illustration, not a measurement of twenty million independent scientists producing frontier discoveries. CASP paper, supplementary materials, printed page 12.
The distinction exposes the central bottleneck question: what else must grow for additional cognitive work to produce additional knowledge?
The model and its assumptions
The paper formalizes part of the problem with a stylized model. Let A represent software quality and E effective research labor. It assumes:
dA/dt = A^(1 − β) × E^λ
Here λ describes returns to research labor, while β captures how discovering further proportional improvements becomes harder. If research is fully automated and effective labor rises in direct proportion to software quality, E = kA. Substitution gives:
g = (1/A) × dA/dt = c × A^(λ − β)
The proportional growth rate g rises with A when λ exceeds β. Under these assumptions, sufficiently strong feedback can overcome increasing research difficulty. The paper’s central illustrative parameters, λ = 1.40 and β = 1.01, imply a difference of 0.39. Each doubling of software quality then multiplies its growth rate by approximately 1.31. CASP supplement, printed pages 12–13.
The algebra is straightforward. The mapping to reality is the difficult part. Software quality compresses several quantities into one variable: capability, training efficiency, inference efficiency, and perhaps the quality of research organization. These do not necessarily translate into research output in the same way.
How sensitive is the acceleration estimate?
To illustrate sensitivity, I recalculated the time needed for a tenfold rise in the proportional growth rate, keeping the first software-quality doubling at 4.5 months in each case. These are mathematical scenarios, not probability estimates or calendar forecasts.
| Assumed λ − β | Doublings needed for a tenfold growth-rate increase | Elapsed time in the idealized model |
|---|---|---|
| 0.10 | 33.2 | 60.5 months |
| 0.39 | 8.5 | 17.1 months |
| 0.80 | 4.2 | 9.5 months |
| 0 | No increase in the proportional growth rate | Never |
These scenarios hold the initial doubling interval equal, not the instantaneous initial growth rate. They exclude changing compute constraints, data limits, and validation delays. Their value is diagnostic: a conclusion that sounds like an eighteen-month forecast rests heavily on uncertain parameter values and the persistence of the model’s assumptions.
Compute and research returns
Historical innovation research gives a reason to take diminishing returns seriously. Bloom and colleagues found increasing research effort alongside declining research productivity across several domains. Their semiconductor example estimated that sustaining the familiar chip-density doubling required more than eighteen times as many researchers as in the early 1970s. That does not establish the same relationship for automated AI research, but it undermines the assumption that more researchers mechanically yield proportionally more progress. Are Ideas Getting Harder to Find?.
Compute adds a different constraint. Whitfill and Wu analyzed labor and research compute at four leading laboratories. Two alternative specifications produced different conclusions about whether the inputs substitute for or complement each other. Accounting for the scale of frontier experiments changed the estimated relationship. This is evidence that the question remains empirically unsettled, rather than proof that compute either guarantees or prevents an explosion. Will Compute Bottlenecks Prevent an Intelligence Explosion?.
The distinction is practical. Better reasoning can save compute by rejecting weak ideas before experimentation. It may be unable to substitute for a large experiment that reveals an effect visible only at scale. A million proposed hypotheses do not resolve that uncertainty until something discriminates among them.
There is also an accounting problem. If an algorithm’s advantage grows as training compute grows, historical gains may reflect both better software and larger training runs. Treating all the improvement as a return to research labor can exaggerate the strength of a software-only loop. Anson Ho’s analysis of algorithmic progress identifies this scale dependence as a major uncertainty. Epoch AI’s analysis.
Serial work creates another limit. In a deliberately simplified example, suppose 80% of a research cycle becomes ten times faster and the other 20% stays unchanged. The entire cycle becomes approximately 3.57 times faster: 1 / (0.2 + 0.8/10). Even infinite speed in the accelerated portion caps the total gain at fivefold. This illustration assumes a fixed, serial workflow; real teams can redesign tasks and shift bottlenecks. It explains why automation percentages cannot be read directly as speedups.
Verification and informative feedback
Verification is especially consequential because it can become slower relative to generation. If agents propose ten times as many experiments, an unchanged review process can accumulate a larger backlog. Automating review can help, but only if its judgments continue to discriminate between real gains, statistical noise, and invalid shortcuts.
Synthetic data has the same structure. Shumailov and colleagues demonstrate degradation from indiscriminate recursive use of generated data in the settings they study. That does not imply that all synthetic-data methods fail. It shows why generating more examples and obtaining more reliable information are different operations. Model-collapse study.
There is also positive evidence for learning from generated attempts. DeepSeek-R1-Zero used reinforcement learning with correctness rewards for verifiable tasks, including math answers and code tested against predefined cases. This supplies a feedback signal beyond the model merely imitating its own output. Its success does not establish equally reliable feedback for open-ended research, but it explains why model collapse is not a universal objection to automated learning. DeepSeek-R1 technical paper.
My expectation is that bottlenecks will move as systems improve: from coding to experiment selection, from experiment selection to execution, and from execution to reliable validation. This is a hypothesis about development dynamics, not an observed universal sequence. It suggests measuring the marginal effect of an extra unit of each resource rather than assuming one permanent constraint.
The important question is whether automated research can repeatedly remove the next constraint before it limits the feedback loop. Demonstrating self-modification answers only the beginning of that question. Demonstrating sustained progress across changing bottlenecks would answer much more.
The ContextOS connection
For teams applying this reasoning, the ContextOS cost-per-trusted-outcome essay provides a related accounting perspective: evaluate the resources needed to obtain an accepted result, including retries and verification. A cheaper proposal is useful only when the complete workflow benefits.
