Pitfalls in AI Self-Improvement

The evolutionary misconceptions behind recursive self-improvement

Die Leiter des Auf- und Abstiegs (The Ladder of Ascent and Descent), from a 1512 edition of Ramon Llull. Wikimedia Commons (public domain).
Die Leiter des Auf- und Abstiegs (The Ladder of Ascent and Descent), from a 1512 edition of Ramon Llull. Wikimedia Commons (public domain).

The AI-Doomer discussion has gone mainstream, producing bizarre bromances, a weird fight between the right and effective altruists, and the NYT treating once-fringe AI Doomsday as though it is actual news.

The Takeoff Story Takes Off

A key assumption behind doomsday scenarios is the notion of Recursive Self-Improvement (RSI). It’s an essential step in the narrative chain from the models we have today to something that can presumably end humanity. The idea, in its simplest form, is that at some point in time AI becomes smart enough to “take off” and begin improving itself. At this point, the story goes, positive feedback sets in and AI quickly becomes some super-intelligent, super-capable system that we did not create, do not understand, and do not control. What happens next depends on who you ask. The doomers suggest it will destroy us, either intentionally or by accident. The Boomers believe it will act as a god-like “machine of loving grace”, curing cancer and ending our need to work.

I’ve found myself astounded by how mainstream these discussions have become, given their origins in Harry Potter fan fiction. Scholarship hasn’t gotten much better in the intervening years, and most of what I can find is still fan fiction. For example, Project 2027 suggested we’re on track for AI with “100x Human thinking speed” (?) by September of next year. It was boosted by credulous coverage in places like the New York Times.

A lot of AI 2027 is just story telling. There’s a little bit of vaguely-quantitative fitting models to data but it’s messy and would fail a stats class. They start with a benchmark put out by METR, ostensibly measuring coding capability at AI R&D tasks. They then fit a logistic curve to a handful of points well-before the inflection point, hoping to predict when a given threshold score is reached; a prediction that missed.

Next, they set out to estimate a bunch of different “gaps” for AI to cross to go from attaining that score to being a “super-human coder.” Then they forecast takeoff, which they define as “the time between a superhuman coder and wildly superhuman capabilities.” I encourage you to give their methodology a read, but most of it is just stacking guesses and noisy sigmoids. For this, they’ve been awarded (at least) $2M in funding.. Earnest science seems to be the wrong line of work these days.

Their model has one notable feature: It doesn’t ask if we’ll have RSI, it just asks when. It likewise doesn’t seem to ask if we’ll have super-intelligence, just when. These conclusions are baked into the model guesswork, and it can only have one outcome. Sayash Kapoor and Arvind Narayanan have pushed back, highlighting that there are plenty of plausible ways in which AI might fail to “take-off”. These include:

  • Amdahl’s Law: The Speedups AI affords do not make too much of a difference because of bottlenecks elsewhere
  • Expanding Pie: Tasks requiring human researchers to progress grow as AI capabilities expand such that automation doesn’t really speed up R&D
  • Collective slowdown: This one is my favorite. The gist is that AI-assisted research might seem to increase the productivity of individual researchers while slowing collective progress. For example, even though it might broaden the methods one researcher uses it could narrow the set of methods used across researchers.

None of these are mutually exclusive and you should really check out their work on bottlenecks to progress. Nevertheless, the core of these debates---the notion that AGI is upon us via recursive AI self improvement---feels unmistakably rooted in common misunderstandings of evolutionary biology. This blog post is going to be a bit of a doozy. We’ll start by correcting some common misconceptions and then apply that corrected perspective to an AI apocalypse.

Misconceptions in natural (or artificial) selection.

Evolutionary biology often invokes the notion of fitness, which can be explained roughly as the (relative) (prob)ability of an organism to survive and pass on its gene(s) to its offspring (and their offspring (and their offspring(and their offspring))…). While in theory it gets quite complicated, empirical biologists often do simple things to measure fitness such as counting eggs or taking the ratio of bacteria on petri dishes.

The brilliant insight of Darwin and Wallace was that organisms vary in their heritable traits, such that those which are associated1 with higher fitness will be more likely to persist in the population. Perhaps more colorful birds have more eggs, or bacteria producing more of some enzyme will reproduce quicker. Whatever the case may be, “survival of the fittest” leaves many of us with a mental model of selection that is directional, increasing a trait on some axis. Something like this:

Trait selection

Figure 1. Classic directional selection: a trait's value is pushed in a direction by higher fitness associated with larger values of that trait.

When teaching evolutionary biology, some students invariably adopt this directional selection model of evolution and mentally map it to “The Road to Homo Sapiens” by Rudolph Zallinger. This monkey-to-man depiction of evolution is contentious for plenty of reasons, one of which is that it depicts something like a chimpanzee as less-evolved than us humans.

The March of Progress

The March of Progress (Rudolph Zallinger, 1965). Image via Wikimedia; rights remain with the copyright holder.

The reality is, of course, that the chimpanzees have been subjected to natural selection for the very same millennia as us and are every bit as evolved. Darwin, perhaps anticipating this misconception, famously wrote “I think” above a bushy depiction of speciation. Each species at the tips of branches (the present moment) is every bit as evolved.

Darwin's 1837 tree of life sketch from his notebook

Figure 2. Darwin’s 1837 notebook sketch of branching speciation. Wikimedia Commons (public domain).

Natural selection produces not only directional selection but also stabilizing selection and disruptive selection. Stabilizing selection sees highest fitness in a Goldilocks zone, imposing costs for (perhaps literally) being too hot or too cold. Disruptive selection does the exact opposite, imposing a cost for fence-sitting and making organisms choose a side (e.g., dark or light colored). Selection isn’t the only game in town and is accompanied by genetic drift, where genes in a population change somewhat arbitrarily.2

Directional, stabilizing, and disruptive selection

Figure 3. Directional, stabilizing, and disruptive selection

If all of this is not complicated enough, traits close to one another on the genome wind up having somewhat linked fates as they are likely to get inherited together. Maybe one is under directional selection and the other disruptive, what then? Selection navigates all sorts of Trade-offs and produces traits that have no adaptive function even though they appear to, called spandrels. This is before we get into the nitty gritty debates about inclusive fitness and multi-level selection. These forces pulling and pushing traits around happen all at once, simultaneously for all species from bacteria to blue whales in a complex cacophony of concert and conflict. Suffice to say, the March of Progress is an oversimplification.

The rot at the core of Recursively Self-Improving AI Super Intelligence comes from a glance at Zallinger’s monkey-to-man or skimming a book on evolution. It becomes all too easy to slip into thinking that human intelligence is more of something than chimpanzees or even roundworm intelligence because we’re further along some direction3. That intelligence can compound, creating feedbacks and leading to super-intelligence. Yet if it feels intuitively obvious that we are more intelligent than roundworms, is your model of intelligence simply human-like-ness or something universal? If we gave you control of a roundworm robot, with its sensory input, would you outperform them on their 302 neurons 4 on their home turf? If you think so, click here.

Evolutionary Biology for Technologists

I earned my PhD in ecology and evolutionary biology, but these days mostly study Homo sapiens. In EEB, we’re taught to avoid just telling verbally compelling “Just so” stories about evolution and ecology. “Giraffes have long necks to eat high leaves” is the classic example. It seems almost obviously true, but even very compelling verbal logic can be misleading. From our first year of graduate school, we are encouraged to make (or consider) explicit mathematical theory which can show us the perhaps counterintuitive predictions of our assumptions. These testable predictions allow us to evaluate theory against data. We can then revise theory, gather more data, etc… You know, do science stuff.

Rationalism, which birthed a lot of the current discussion around AI, lives and breathes the idea that we can arrive at truth by verbal logic alone. Despite frequently invoking mathy-sounding terms like P(doom)P(doom), it’s just a bunch of “just so” stories stacked on top of one another. Adding numbers doesn’t make it scientific, nor does calling it superforecasting. These stories remain infused with the “March of Progress” misconception of how traits evolve, and what intelligence is.

It’s important to note that their argument is an argument about evolution. Boiled down:

  1. Intelligence is directional.
  2. We can select for intelligence in machines5.
  3. At some point they can manage their own selection, creating a runaway selection.

Therefore:

  1. Superintelligence is an inevitability.

Climbing mount intelligence

Cheetahs clearly experienced a period of intense directional selection for speed and now are capable of running 70mph. Wait, you might ask, if they run at 70mph why don’t their prey just run at 80mph? Then they should run at 90mph? Then their prey would go 100? Perhaps, you’d posit, that we’re witnessing the early stages of a runaway increase in cheetah speed and they’ll be traveling at Mach 1 or even Mach 3 in no time. Humans will no longer be safe, hunted down by cheetahs that hit us like a ballistic missile. A cheetah apocalypse is upon us, unless we act now6.

As absurd as it seems, it’s the same dynamical argument as Recursive Self-Improvement (RSI). Humans build increasingly intelligent AI until we hit a point where it improves itself, after which runaway self-improvement leads to superintelligence. That’s the argument that has people so scared they’re donating billions. Gatherings of very influential people discuss where we are “on the curve”, by which they mean the growth rate of AI capabilities relative to “AI Safety”.

Given all this concern, it is remarkable that influential reports, resort to just drawing curves that align with their intuition. What if we instead tried to apply, you know, scientific thinking7 to the question of RSI?

A toy model of RSI

For the moment, we can formalize RSI into a model, keeping things as simple as possible. We need terms for how smart artificial intelligence is (AA) and human effort (HH). Specifically, the ASI argument is about how much faster AI will get, given how smart it is. We’ll call this A˙\dot{A}:

A˙=cH+r(A)c,H,A>0\begin{aligned} \dot{A} &= cH + r(A)\\ c,H,A &> 0 \end{aligned}

The constant, cc, just helps us scale human and AI effort at improving AI. What matters more is this function, r(a)r(a). It’s the math that tells us how much AI can contribute to its own growth in intelligence. In the verbal model, it does nothing until a critical point at which AI begins aiding in its own development. Here’s a simple choice that captures the spirit of the verbal model:

r(A)={0,A<Ac,ρA,A≥Ac,ρ>0\begin{aligned} r(A) &= \begin{cases} 0, & A < A_c,\\[0.25em] \rho A, & A \ge A_c, \end{cases} \\ \rho &> 0 \end{aligned}

Once some threshold is crossed (AcA_c), AI starts helping to improve itself alongside human effort. Below that, it does nothing. There’s a graphical way we can analyze these kinds of models, by plotting A˙\dot{A} vs. AA as seen in the figure below. This tells us how much smarter it’ll get in the next bit of time, given how smart it is currently. If the line is above zero, AI will get smarter; if it’s below zero it will become less smart (we’ll come back to this).

Diagram showing positive values of A dot for all values of A

Figure 4. $\dot{A}$ vs $A$ with an RSI threshold, the Yud model.

The dynamics of this model match the story we keep hearing; humans push AI’s intelligence forward (cH>0cH>0) until it starts pulling its own weight (A>AcA > A_c) and winds up becoming super intelligent (A→∞A\to \infty as t→∞t\to\infty). All of the parameters are positive logically, so no other story can possibly emerge from this math. All we can tweak is how fast or slow superintelligence arrives. The dynamics of this model seem to match the intuition of Doomers and Boomers alike. These lengthy discussions about AI safety, pacing the frontier, etc… are all just fiddling with cHcH or regulating AcA_c and ρ\rho. But is the model realistic?

Adding costs

What if we add a tiny little tweak? The current model implicitly assumes that the cost to store, manage, evaluate, adjust, and maintain a one thousand parameter toy model is the same as today’s frontier systems with trillions of parameters. Worse yet, it assumes each of those costs is precisely zero, or near enough to ignore.

This assumption seems inarguably wrong, given the amount of capital being invested into AI and datacenters at this very moment. Let’s add a term that accounts for costs which scale with capability. We’ll call this −δA-\delta A, and make it the simplest choice. Simple is good because it allows us to have a model that says “let’s not forget about costs” without baking in a desired result. Here’s our new model:

A˙=cH+r(A)−δAc,H,A>0\begin{aligned} \dot{A} &= cH + r(A) - \delta A\\ c,H,A &> 0 \end{aligned}

Add even this tiny tweak and things get interesting. Now, the takeoff story runs into a problem if cHδ<Ac\frac{cH}{\delta} < A_c. This would be the case if human effort doesn’t manage to get AI to where it is starting to improve itself. Whether or not we’re there already is one topic at a big meeting; folks certainly are using AI when doing AI R&D but for all the reasons Arvind and Sayash outlined it’s possible that we haven’t actually hit AcA_c as defined in this overly simple model.

Here too, we can analyze the model visually, but we’re gonna take a slightly different approach because the model does more than “line goes up”. Below, stable-state intelligence (i.e., how smart will AI ever get) is plotted against the cost rate, δ\delta, while holding other values in the model constant. Unlike the plot above, height here tells us where intelligence AA settles at the end of time (e.g. t→∞t\to\infty). If costs of model maintenance are large enough (green line), our human effort simply won’t be enough to get AI to where it can begin improving itself. This is the “no RSI, no way, never” regime. Finite artificial intelligence, below human intelligence.

Figure 5. Long-run intelligence (A) vs costs.

Perhaps you think we’ve achieved RSI, and takeoff has begun. What’s interesting is that with costs, RSI no longer guarantees takeoff. There’s a regime (clay-colored line) where AI contributes to its own development but still runs up against cost drag and settles at finite intelligence. This may be below human intelligence, perhaps a bit above. Alternatively, it might be as smart as a dozen people but not two dozen. Finite either way. True runaway takeoff needs self-improvement to outpace that drag of maintenance cost with increased intelligence (i.e., ρ>δ\rho > \delta, left-most regime of the graph).

Of course all models are wrong, but how is this one informative? It shows us that incorporating costs complicates the simple story of RSI yielding take-off. We suddenly have scientific questions to ask about how those costs and capabilities scale. Maybe our model needs refinement as we pull in data. For example, cost here is simply and generously linear with intelligence. It’s certainly conceivable costs could be superlinear, sublinear, or dynamic with AA and require its own model C˙\dot{C}.

We also might want to model dynamics of human investment in AI (H˙\dot{H}), perhaps we get disillusioned and disinvest if progress is too slow and get excited when it’s fast, investing more. Maybe we want to tease apart costs of development and costs of maintenance. All of this is the kind of thing we teach to undergrads in EEB, and theory can be merged with data to go from learning qualitatively about the world (as above) to making actual predictions. You know, science stuff.

If anyone were seriously worried about AGI, this is the kind of thing we’d be doing. It’s what we do for other risks, from nuclear war and the spread of disease to conservation and climate change. In each case, we make and refine models with data---we don’t just draw curves and vibe it out or sum up guesses. Unfortunately, the “future of AI” discussions are dominated by captivating cranks and “useful idiots” who command so much attention.

Back to the cheetah. We could almost replace the terms in our model above and wind up with a C-worthy model of why cheetahs aren’t infinitely fast. Even though there may be selection there are always costs such that directional selection hits a wall and becomes something closer to stabilizing selection. Verbal vibe models forget that costs exist, and of course lead to ASI-runaway superintelligence yet it’s no more plausible, nor worrisome, than our infinitely fast Cheetah doomsday.

For doomers' eyes only

For the doomers hate-reading this, one very fun feature of the model above is that it provides a means by which you can avert AI Doomsday without solving the alignment problem, regulating models, or nuking data centers. It’s an existing societal lever that we pull, all the time, for way less important reasons. I suspect it’s even one you could pull directly for less than the projected EA investment in solving the AI apocalypse, and could pull indirectly for orders of magnitude less. It might even pull itself. Do you see it?

The topological problem

RuBisCO, on which all of life depends, is pretty shitty at its job. If you reach back into the recesses of high-school biology, RuBisCO is part of the dark side of photosynthesis, and its job is to pluck CO2\mathrm{CO}_2 out of the atmosphere and jam it onto things with long names that eventually become glucose (sugar).

The trouble is that RuBisCO gets confused and frequently winds up grabbing atmospheric Oxygen, O2\mathrm{O}_2 instead. Rather than storing energy as sugar, it winds up burning it in (photorespiration). It’s roughly the equivalent of working a fast food register and giving the customer some money from the till plus their burger. Yet it makes this mistake somewhere around 20% of the time. Given that it’s the enzyme behind converting light into food and sequestering extra atmospheric CO2\mathrm{CO}_2, there’s a lot of interest in understanding if, why, and how it is “slow across the tree of life”8.

So why is RuBisCO shit? When it first evolved, the atmosphere had very little O2\mathrm{O}_2 and there was no meaningful risk of accidentally grabbing it and engaging in photorespiration. Photosynthetic organisms kept pumping oxygen into the atmosphere, producing the great oxidation event where most of life on earth went extinct. Ooops. Ever since then, RuBisCO has been making mistakes. But why hasn’t it gotten better?

There’s a saying in Maine, that “you can’t get there from here.” A state with many peninsulas, some destinations may seem close but require a very long drive up north to Highway 1, and back down south on the adjacent peninsula.

National Atlas map of Maine showing peninsulas and coastline

Figure 6. Maine. National Atlas of the United States via Wikimedia Commons (public domain).

So too with RuBisCO. It isn’t great, but the pathway and organisms it’s embedded in have undergone substantial selection to convert light into sugars. It’s a part of that clockwork. Mutations to replace RuBisCO would likely reduce photosynthetic output overall until other parts of the system caught up with it. Organisms en route to better RuBisCo would get outcompeted. We’re stuck with it not because it’s the best, but because it was first.

In my opinion this is the biggest problem with RSI arguments, my model above, and the discussion of AI futures more broadly. Each assumes that R&D for intelligence is a single, smooth-enough hill that we’re climbing. Disagreements arise from whether the hill is infinitely tall or not, and whether we’ll accelerate our way up or start to slow down. Nevertheless, they sort of assume that we can, at least in theory, get there from here.

For example, even under Amdahl’s law, progress towards more intelligent systems continues albeit hampered by bottlenecks. We’ll get formal again soon, but let’s start by just imagining a ball being pushed up the intelligence hill by humans, machines, or some combination per our model in the last section. The height of that hill is going to be determined by costs, settling at (say) cHδ−ρ\frac{cH}{\delta-\rho}, even if ASI occurs but doesn’t overcome take-offs.

Figure 7. Models generally assume we're climbing the lone intelligence hill

Competing architectures

The architecture behind today’s AI is largely Generative Pre-trained Transformers (GPTs). Perhaps there exists an entirely different way of constructing artificial intelligence. Let’s imagine two other architectures, TBDs and BFDs, that have lower costs, yielding higher finite (TBDs) or infinite (BFDs) intelligence.

Figure 8. Each architecture could have its own unique limits

Here’s the rub, even if we “discover” these architectures, they’ll be at a competitive disadvantage. The trouble is, we can’t know how far a given architecture will go before we develop it. Let’s say that, unbeknownst to us, we only need to invest $100B into BFDs to hit “RSI” and achieve superintelligence. Unfortunately, even after $30B you still have only a 2022-quality chatbot. Can you convince venture capitalists in 2035 to give you more for your shitty chatbot? BFDs don’t get funding, the big models are no longer maintained and all that exists are a few toy open-weight models (as promised, an example where models can get less intelligent). Here’s that story mathed out in a toy model.

Formalizing competition for human effort

We can formalize our competing architectures. The math is going to have a few more terms in it, but it’s the same basic idea as before:

A˙1=c1H1+r1(A1)−δ1A1,A˙2=c2H2+r2(A2)−δ2A2\begin{align*} \dot A_{\mathrm{1}} &= c_1 H_1 + r_1(A_1) - \delta_1 A_1,\\ \dot A_{\mathrm{2}} &= c_2 H_2 + r_2(A_2) - \delta_2 A_2 \\ \end{align*}

One can think of better models, but the simplest thing we can do is just double the model above for two competing technologies. Each has its own take-off point and its own unique costs. Let’s assume that humans help the technologies along proportional to their current benchmark scores on some intelligence test. In other words:

H1=HA1A1+A2H2=HA2A1+A2\begin{align*} H_1 &= H\frac{A_1}{A_1+A_2} \\ H_2 &= H\frac{A_2}{A_1+A_2} \\ \end{align*}

Because this is a blog post, let’s consider one simple parameter space for the model. We’re going to imagine that A1A_1 is our GPTs, capable of perhaps RSI but only finite intelligence. An alternative approach, BFDs, A2A_2, creates artificial intelligence but does so in a qualitatively distinct way. These technologies are mostly the same, except that BFDs/A2A_2 has better cost scaling (i.e. δ\delta is smaller) and is a bit more efficient for RSI. We set parameters such that, on their own, BFDs would become infinitely super-intelligent. What happens in competition?

Phase plane of two competing AI architectures

Figure 9. Phase plane for two competing architectures

We’re going to analyze this one visually again, here with a phase-plane diagram. The grey arrows show where we wind up from where we start. The orange is a scenario in which GPTs/A1A_1 have a head-start by the time BFDs are discovered. Human effort goes into GPTs and we never find out what BFDs can do. Had the opposite happened, and BFDs been discovered first (green line), we’d wind up inventing super-intelligence and abandoning GPTs.

Hopefully it’s intuitively plausible that existing technology can choke off R&D from eventually-better alternatives. Is this where we find ourselves with ChatGPT?

Drawing an analogy to RuBisCO, GPTs were developed on an internet free from AI slop, but then quickly dumped a bunch of it onto the internet. Like RuBisCO, GPTs every so often wind up training on AI-generated slop, mistaking it for human-generated content. There’s a little inefficiency from popping into existence and then changing the world around you, analogous to the one RuBisCO incurs. We built an entire ecosystem around it---just as plants did---embedding its inefficiencies into our digital world.

The broader point is that even if super-intelligence is technically possible, it might require an alternative architecture to the one we’ve currently put so much stock into. An implicit assumption of RSI-Doom is that the current architecture is the one that will lead to super-intelligence. It seems incredibly unlikely that the first architecture we stumble on that goes as far as GPTs have has infinite capacity for intelligence. Assuming such an architecture is technically possible, it might no longer be competitively possible given the investments we’ve made into GPTs. We might be stuck with the AI equivalent of RuBisCO.

Scala Artificialis

For millennia, it had been taken as obvious that life and nature had a hierarchy. God on top, followed by angels and current local religious leaders, then other humans, mammals, birds, lizards, fish, insects and so on until we reach rocks. Any questions can be resolved by consulting the chart, learning that Lichens are more godly than truffles. Versions of this hierarchy go back to Aristotle’s ‘History of Animals’, but the idea is most firmly associated with the scala naturae of medieval times. It was literally hard to argue with. What commoner would be bold enough to argue a peasant can be smarter than a king?

1512 woodcut of Ramon Llull's Ladder of Ascent and Descent, a staircase of being from stone to God labeled Scala intellectus

Die Leiter des Auf- und Abstiegs (The Ladder of Ascent and Descent), from a 1512 edition of Ramon Llull. Wikimedia Commons (public domain).

Aristotle’s ranking system that birthed scala seemed to have two components; a suite of perfection-like traits and how close various species were to humans in those traits. In Book IX, he writes:

The characters of animals, as has been observed, differ in respect to timidity, to gentleness, to courage, to tameness, to intelligence, and to stupidity…In a general way in the lives of animals many resemblances to human life may be observed.

These lines set the groundwork for beliefs i) That these traits like intelligence are things that one can have more or less of and ii) humans have the most of these traits relative to life on earth. We see this same thinking about Artificial Intelligence, when groups like Project 2027 say things like “between a superhuman coder and wildly superhuman capabilities.” Given how deeply rooted scala natura is in our cultural history, it’s very easy to imagine AI progressing from less intelligent than us to more intelligent than us, and to think we might be able to benchmark it as with Humanity’s Last Exam. RSI, we’re told, can rocket AI to the upper-echelons of the scala artificialis, becoming “machine gods”9.

One of the dangerous features of Darwin’s “I think…” bush is that it breaks the scala natura10. All creatures that exist today are equally evolved, even related if we go back far enough. Human-like traits are just what our lineage has at this snapshot in its history, not some pinnacle of achievement as God worked their way up to make more complex forms of life. It breaks any sort of natural argument for why our intelligence is more of something than the intelligence of other species.

Climbing where, exactly.

To ground our thinking, we can very easily rank extant mammals by body mass from the Etruscan shrew to the Blue Whale, creating a scala gravitatis. We can easily do the same with body temperature (though most are quite close), lifespan, length, and other things that are natural units. It may seem subtle, but even something like number of offspring is no longer in natural units and starts requiring judgement calls. Naked-mole rats, for example, cooperatively rear the young of a single breeding queen. Who gets assigned the offspring? Do we divide them somehow? These are not hard questions in practice, but they do require that we operationalize our definition of offspring. The choices we make infuse our values into that definition, even for something like offspring counts. Consider twins, chimeras, and conjoined twins. How we count each of those categories infuses our value-laden judgements of what constitutes one living being into the definition.

Once we get to something like intelligence, our definitions do not come from the natural world but are fundamentally tied into the kinds of decision-making and learning processes that we value. Across nature there are plenty of examples of animals doing cognitive tasks that would make Einstein throw up his arms in frustration. We don’t put those on IQ tests or Humanity’s Last Exam, which is a choice we make; an echo of the scala in our definitions of IQ.

Supposing Super-intelligence

Let’s formalize this. Imagine that there exists some large, but finite, number (kk) of “intelligence” components that individually and less objectionably exist. Atomic capabilities that we could perhaps pin down the mechanism by which they occur and select for in a species. Each individual (ii) has some underlying ability in every single conceivable trait.

ai=(ai,1,…,ai,k),k>1\mathbf{a}_i=(a_{i,1},\ldots,a_{i,k}),\quad k > 1

A given task yields a score that is some function of these abilities. Maybe we count how many acorns you can bury in the fall and dig up in the spring. We can include IQ-associated tasks too, such as Raven’s Progressive Matrices or even a full IQ test. For task tt and individual ii, the score they’ll receive is:

sit=wt⊤ai+εit s_{it} = \mathbf{w}_t^\top \mathbf{a}_i + \varepsilon_{it}

Here, wt\mathbf{w}_t is how much each ability matters for that task and εit\varepsilon_{it} is noise. Clearly forms of spatial memory will be very important for caching, less so for Raven’s Progressive Matrices. If you’re a species who depends on cached food, you are selected for higher levels of all of the ai,ja_{i,j} that yield more cached food retrieved and not so much for the Progressive Matrices associated traits. OpenAI, Anthropic, DeepMind and other AI Labs companies, are working to increase their models’ abilities aa by selecting across various reward (training) and benchmarks, all of which can be thought of as sits_{it}.

If “intelligence” is a monolithic concept that spans species, machines, and machine-gods alike, it would require that all of the abilities for a given species, LLM, or person are projections of their deep-seated intelligence IiI_i:

ai=Ii v\mathbf{a}_i = I_i\,\mathbf{v}

with v\mathbf{v} being a common positive direction in ability-space. Every task score is just IiI_i rescaled by how much that task loads on v\mathbf{v}:

sit=Ii (wt⊤v)+εit.s_{it} = I_i\,(\mathbf{w}_t^\top \mathbf{v}) + \varepsilon_{it}.

Nick Bostrom defines super-intelligence as “any intellect that greatly exceeds the cognitive performance of humans in virtually all domains of interest”. In our model, this would be:

IM≫IH,I_M \gg I_H,

which implies that for (virtually) every task tt, machines can do it better per Bostrom’s definition:

sMt≫sHts_{Mt} \gg s_{Ht}

If we further take the scala natura assumption that humans are smarter than animals:

IH≫IA,I_H \gg I_A,

We wind up at the conclusion that not only will intelligence be able to do everything that we can do, but it should be able to do everything that every animal on earth can do. Indeed, long before AI becomes super-intelligent relative to humans, it should possess the ability to do virtually all animal cognitive tasks. Monarch Butterflies can Zero-shot migrate across North America based on very noisy sensory cues, can our state of the art LLMs do that? Perhaps each new model we can do a Groundhog Day test of their migratory capacity, resting assured that superintelligence has not yet arrived as the little butterfly drones crash in the ocean.

If, for some reason, our newest Grok model cannot pilot a little butterfly drone across the continent nor construct a termite mound, then we have to conclude that either animals are smarter than superintelligence, which is smarter than us

Ibutterfly≫IM≫IHI_{\textrm{butterfly} } \gg I_M \gg I_H

The alternative is to accept something is wrong with a simple model of unidirectional intelligence. The easiest alternative is to simply abandon the notion of monolithic intelligence (II). IQ and such does not rescue broad intelligence, because it’s simply a test that weights an axis along which humans covary in their traits. Measuring this axis of covariance doesn’t indicate that the heavier concept of some universal, cross-species (and machines) intelligence exists. IQ is a huge topic, but if we changed the tasks on the tests, we’d get a different score for folks. This is why we have all sorts of notions of intelligence (e.g. emotional, spatial, etc… ). Each is just a set of weights on traits.

All this is to say, the notion of super-intelligence demands that we have a coherent way of thinking about intelligence as a monolithic axis along which humans, machines, gods, and animals sit. It requires a scala artificialis. That, in turn, implies that before we surpass human intelligence Claude should climb past amoeba, ant, butterfly, lizard, fish, bird, chimpanzee, elephant, dolphin, and maybe blue whale. If this doesn’t seem coherent we need to rethink intelligence.

Perhaps Bostrom or whoever walks back intelligence to simply mean scoring higher on an IQ test, Humanity’s Last Exam, or some other test or benchmark. Maybe scoring better on all of them (which just redefines the benchmark). We narrow the definition to simply mean better at some meta-test that humans take, but not better at all cognitive tasks that all species do. That narrow definition doesn’t exactly imply that it’s a direction one can travel in forever, to the point of curing cancer or destroying humanity.

We also wind up with (at least) one more problem for superintelligence: collective human intelligence. I’ll save the details of this for another post, but the intelligence of all humans collectively, over time, is certainly greater than the intelligence of one human. Even if a machine exceeds every individual on every test we can conceive of, there remains a chasm between that and exceeding the totality of collective human ingenuity, perseverance, capability and intelligence (however defined). As a result, even once/if RSI pushes an architecture up mount intelligence and we can find no test by which any human exceeds it, it still doesn’t have us beat.

Closing thoughts.

While working on this piece, the Bulletin of Atomic Scientists put a fantastic piece by Sarah Kreps, “AI doom has a precision problem”.. It is honestly one of the best things I’ve read on the topic. For example:

Even if more capable AI accelerates research, it does not necessarily translate into proportionate capability gains because they require corresponding increases in compute, energy, and data, which may not scale at the same rate and become binding constraints.

This is the point of our second toy model, where we add costs. Of course a sentence as written above is much simpler than math and likely to appeal to a broader range of people. So why formalize?

This blog post was motivated by reading through reports from a host of “institutes” on “AI Safety”. Project-2027 was somewhat unique in that they fit a sigmoid to a few data points here and there, and sprinkled on guesstimation11. Yet even they didn’t seem to release their code and data (unless I missed it), such that their math and assumptions aren’t laid bare. Many places just pull p(Doom) out of their ass, or anchor onto the 10% zombie statistic that is floating around. These institutes, even well-funded ones at universities, seem comfortable just drawing curves without considering the actual underlying dynamics. Hundreds of millions spent, and this is the level of rigor that characterizes “AI Safety”---or at least the concern which motivates it.

A point of comparison is warranted. In the mid 1800s, Eunice Foote discovered the Greenhouse effect via an ingenious experiment involving various gases in glass containers exposed to light. Decades later, in 1896, Svante Arrhenius published the world’s first climate model. His model made predictions of temperatures throughout the year, by latitude across varying degrees of CO2\textrm{CO}_2 concentration.

Table VII from Arrhenius: variation of temperature caused by a given variation of carbonic acid, by latitude and season

Table VII from Svante Arrhenius, “On the Influence of Carbonic Acid in the Air upon the Temperature of the Ground,” Philosophical Magazine (1896). Temperature change (°C) for different levels of atmospheric $\textrm{CO}_2$ , by latitude and season.

How did he arrive at these numbers? If you read the paper linked above, it was through a very thorough set of equations he derived for the physics governing temperature change with the greenhouse effect. It even includes things like albedo feedback. Most astonishingly, the predictions for temperatures he comes up with aren’t bad. They’re off by perhaps a degree or so for a doubling of CO2\textrm{CO}_2 relative to modern IPCC calculations. Of course a degree means a lot with climate change, but not bad for calculations by hand a decade before the Ford Model T.

Pages 256–257 of Arrhenius 1896 paper on carbonic acid and ground temperature, showing radiation-balance equations

Svante Arrhenius, “On the Influence of Carbonic Acid in the Air upon the Temperature of the Ground,” Philosophical Magazine (1896), pp. 256–257.

I encourage you to read Svante Arrhenius’ paper and compare his methodology and rigor to AI 2027. Pay attention to how Arrhenius handles assumptions versus how Kokotajlo and co. glaze right past them, just asserting what will happen. In the 130 years since Arrhenius, climate models have evolved considerably and exist in various forms. The catastrophe they predict isn’t baked in, it’s driven by measurable quantities of CO2\textrm{CO}_2 and known/plausible models of geophysical processes.

Project 2027 and the various AI safety outfits, by contrast, seem almost allergic to spelling out their assumptions into a proper model of the process by which they predict catastrophe. Above, with the simplest models I can imagine, I wanted to show why. These models don’t tell us what the future will bring but force us to think more clearly about how we’d answer that question. We find that RSI will not universally lead to superintelligence if costs exist, which they most assuredly do. We find that competition across technologies could be a barrier and one without a trivial solution. We also reveal that Bostrom-style super-intelligence implies our GPTs should be able to do all of the things that butterflies, squirrels and fish do---or that intelligence is not a scala. Notably my models do not preclude RSI from leading to superintelligence, they just constrain how that could occur once we are clear about very simple assumptions. They also hint at a ton of levers one could pull to avert such an AI doomsday, but I’ll leave that as an exercise for the Doomers.

As this is a blog post, I haven’t put nearly the time into the theory above that I do academic publications, or that I would if this were my day job. Each of the models could be extended in various ways, probed and analyzed, blended with data. There’s lots of work that one could do just to map out the theoretical landscape, and the data are rich as well. There’s a huge taxonomy of AI models dating back decades; one could fit that phylogeny and perhaps ask questions about the structure of artificial intelligence. Setting humans aside, is there even a scala artificialis?

Yet for reasons that are beyond me, the level of rigor behind concerns over AI safety is less than what we had for climate science in 1896. Nevertheless, hundreds of millions of dollars are invested in fixing the ill-defined problem, it captivates politicians, and has journalists at reputable outlets devoting substantial coverage. It’s apparently headline-grabbing news when CEOs and randos just say shit about p(doom). Academics are even getting involved, taking seriously a problem we would disregard were it not for the substantial amount of funding available if one plays along with doomers’ concerns. At the end of the day, if the effective altruists funding this whole thing dedicated even a tiny fraction of their budget to a red-team of people with proper training in theory and statistics, they’d sleep easy at night and could get back to buying mosquito nets.

tl;dr? If this is an existential crisis, it’s embarrassing we’ve not been more rigorous. If it isn’t, it’s embarrassing we’ve taken it so seriously in the absence of rigor.

Coda: an actual model, albeit a preprint.

After writing this, I was sent this paper out of METR which does make a mathematical model of RSI. Below are my super quick thoughts from a rapid read-through so I could point folks who reply somewhere when this invariably gets linked to me.

The paper is a bit odd, in that it builds out a very large complicated world with various bottlenecks and then winds up riddled with assumptions that boil it down into ~3 terms (section 4.1) which if multiplied together and greater than 1.0 will ostensibly takeoff. To me this seems an exercise in trying to rhetorically acknowledge complexity but then brush it aside with many, many untested assumptions. It’d be more honest, imo, to simply state the 3 term model. That model, simply put, suggests RSI takes off if:

A×B×C>1.0A \times B \times C > 1.0

where:

  • AA: return to R&D after diminishing returns on ideas.
  • BB: how much newer models raise effective R&D.
  • CC: how much algorithmic efficiency gains raise capabilities.

(A/B/C are my simplification of their terms). They plug in 1 for A, 6.5 for C and solve for B: 1÷6=0.151 \div 6 = 0.15. It all seems very precise! They write:

Opus 4.8 scores one ECI unit more than Claude Opus 4.7. Thus under our model, Anthropic would meet the self-sustaining acceleration condition if adopting Claude Opus 4.8 led to 15% higher research productivity than Opus 4.7.

Later:

Claude Opus 4.8 scores 16 ECI points more than Claude 3.7 Sonnet, which was released in February 2025 alongside Claude Code. At the time, estimates for AI uplift on engineering tasks varied but were not conclusively larger than 0.12 More recently, Anthropic engineers surveyed in the Claude Mythos Preview system card reported a 4X uplift from Claude use (Anthropic, 2026). The 4X uplift is very likely to be an overestimate, but even if it were true it would imply a 9% increase in productivity per unit of ECI, below the threshold for self-sustaining acceleration over this period.

Here’s the rub. That 1 for A? Let’s consider just the numerator: they take it from this 2024 paper as ln⁡(3)\ln(3). That comes from the estimate of 8 month doubling time, as best as I can tell. However, the uncertainty there is huge with a CI from 5 to 14 months which would push the numerator around from around 0.59 to 1.7. That makes the critical value (B) range from 9-26%. The denominator’s terms add more uncertainty still. There’s a similar problem with CC, which takes a slope from a regression but not its uncertainty. On top of this, I don’t see any reason to believe AA and CC are constant, independent, and time-invariant.

For all the math, the authors boil off most of the model using either simplification or “calibration” to just focus on BB, how much newer models raise effective R&D. We wind up with a zombie-statistic, 15% that seems like a magical take-off number and are told we’re not there yet (9%) but it sure seems close! The authors then seem to try and even-handedly consider the evidence that RSI is/will occur. In a section called “Best evidence in favor of acceleration:”, one of their bullet points says:

“Jack Clark predicted a 60% chance that by the end of 2028, there will be “an AI system powerful enough that it could autonomously build its own successor”

It’s truly weird to consider a CEO’s guesstimate as “best evidence”. That said, I am glad to see an example of something like RSI being treated with math. At the same time, the math here reads like rhetoric designed to extract a “number to watch”, consistent with METR’s goals… rather than to really sort the science. A lot of the dynamics, including things like saturation and diminishing returns, are boiled off into those fixed-point nowcast numbers for A and B, both of which yeet uncertainty into the void and generate an over-confident estimate. An earnestly curious approach, I would think, wouldn’t be so hasty.

Finally, I would note that, in reading the references in this preprint, it is striking how much they cite Epoch.AI, blogs, Preprints, Working papers, their own githubs/blogs, popular science books, CEOs, news articles and other things that just aren’t peer-reviewed research. As a peer reviewer, I would raise that as a massive red flag. This necessitates two things: On one hand, perhaps I’ve missed formal theory that exists in this parallel grey-literature of EA-funded institute blogs and such. On the other hand, this parallel ecosystem of homogeneously funded “research institutes” largely abstaining from the independent evaluation of peer-review sure looks like an intellectual bubble inside of an economic bubble.

Refs and Notes

  1. Associated is necessary here because the trait might not cause fitness. Sometimes it’s just along for the ride with something that is. If you want to really have your mind explode, start thinking about ways in which you could get apparent directional selection for one trait when the underlying selection is disruptive or stabilizing for the actual causal trait. ↩

  2. If you’re ever surrounded by theoretical biologists, just mention neutral theory and duck out. ↩

  3. This is a hop, skip and a jump from the logic used by fucking racists and eugenicists who assume that because some groups of people transitioned from hunting and gathering to agriculture they’re further along some intelligence-mediated directional technology tree than modern hunters and gatherers. This idea (unilineal social evolutionism) has just been eaten alive by scholarship over the last century and is well understood to be straight up bullshit.. ↩

  4. IF I’m being honest, even the notion that nervous systems are necessary for “intelligence” is a bit of a bold claim. Plants and forests exhibit ingenious responses to their environment, albeit over time-scales we don’t tend to recognize as cognition and behavior. Immune systems, bacterial colonies and plenty of things that are not nervous systems seem to adaptively make use of information in their environments. ↩

  5. Here too we find ourselves a hop skip and a jump away from eugenics, where the idea is that selective breeding will push traits like intelligence (and everything else) in some better direction to develop the Übermensch. ↩

  6. If you think I’m being unfair here, you should read the section of Soares and Yudkowsky’s book where they talk about the various animal gods. ↩

  7. Well, scientific thinking with the degree of rigor and thoroughness one can reasonably expect out of a blog post on a postdoc’s personal website during job application season. ↩

  8. Benoit de Pins et al., Rubisco is slow across the tree of life, PNAS (2025). ↩

  9. Writing stuff like this, I feel as though I’m being a bully and exaggerating perspectives and then I remember that’s literally the name of a new podcast that Effective Altruists totally real journalists funded to the tune of $480K ↩

  10. Darwin’s half-cousin, Francis Galton, took whatever progress towards egalitarianism “I think” could have done and trashed it by applying Darwin’s ideas to humans and creating eugenics ↩

  11. Calling it superforecasting doesn’t make it not guesstimation when you’re guesstimating. ↩

← Back to writing