Who knew self-improvement could be so terrifying?

Illustration by Oliver Burston/Ikon Images
Contemplating a future where machines break free of human meddling
“Recursive self-improvement” might not sound that bad: more like a mind-numbing efficiency exercise than (possibly) a death knell for humanity. But words can be deceiving.
Set off by visions of catastrophe, some coming from inside the house, leaders and researchers at Anthropic, Open AI, and other AI firms have spent the past month delivering public warnings about the speed with which their creations seem to be slipping free of human control, including through calls for a coordinated AI slowdown, government-stamped or not.
At the center of the storm is “recursive self-improvement.” According to Anthropic, this describes an “AI system capable of fully autonomously designing and developing its own successor.” In more ominous terms, it raises the possibility of a future in which humans are demoted to removable obstacle — and then removed. Possible methods range from bioweapons to a full-scale infrastructure attack.
Jonathan Zittrain is the George Bemis Professor of International Law at Harvard Law, a professor of computer scientist at SEAS, and co-founder of the Berkman Klein Center for Internet and Society. We asked him about progress toward an era of self-improving AIs — and the threat such a scenario represents for humankind.
This interview has been edited for clarity and length.
Gazette: What’s a good layperson’s definition of “recursive self-improvement” — is it basically AI working without human prompt?
Jonathan Zittrain: A recursive function is one that takes its output as an input — it’s sort of plugged into itself. To answer, “Who are your ancestors?,” you might say, “Who are my parents?” and take the answer to that and say, “Who are their parents?” — and then keep going. Recursive self-improvement is the idea of asking a current AI to create a new and better AI — a child, by analogy — and then to ask the child to create its own child, and so on. Before you know it there are a lot of descendants, and by hypothesis each one is newer and better than the next. How do we know what counts as “better”? Maybe refining the current definition is one of the things we ask each new AI to do. That illustrates just how much descendants a few generations down might be quite different than those that came before — and also how hard it is to really get something like this going.
Is there any evidence that these models are closing in? If so, an informed guess on timetable?
Much depends on someone’s view of just how clever and broad-thinking current AI systems are, and whether there’s a ceiling on how capable they can be — at least how capable without needing human help and imagination. Some people figure that the most prominent systems, driven by large language models, are at core pattern-matchers to what has come before, not well suited to making the kinds of conceptual leaps that improvement requires — that is, unless improvement only needs existing architectures with faster computations and more memory in the spirit of “more, bigger, faster, louder.” Others are convinced that today’s systems, even if limited in imagination, have enough human ideas to cook with that they can rapidly prototype and evaluate systems that don’t look like themselves. For example, solving how to connect models more with the real world, to act within the world, and to evaluate how they’ve done and adapt their behavior to do better, might produce a new model that really is better — and that can then produce its improved successor, and so on.
Anyone who thinks they have a firm timetable strikes me as deluded or selling something, but it’s hard not to observe that (1) some of the most aggressive suppositions about AI capabilities from 10 or 15 years ago have borne out and (2) many people in the leadership of today’s frontier AI labs believe they are tantalizingly (and worryingly) close to recursively self-improving systems.

“There’s a vital role for universities to play in helping to understand AI,” says Jonathan Zittrain.
Harvard file photo
Once achieved, this is the ultimate “you can’t un-ring a bell” scenario, correct? It seems like that idea is putting a real charge into the urgency around the issue: That once the AIs can improve themselves, there’s no way to reverse course.
I think there might be mechanisms for reversing course, but we might find ourselves collectively unwilling to do so: Cheap superintelligence would, by definition, offer humanity lots of advances that would be hard — some would say immoral — to abjure. Secondarily, and also by definition, superintelligence would be in a position to outfox us. My dog doesn’t want me to leave the house without him in the morning, but he hasn’t figured out how to hold me back. Of course, skeptics will say that as soon as a problem obtains “by definition,” we might be in a territory where our premises assume the conclusion. Are there dimensions of intelligence that computers could achieve and that humanity, ourselves generalist thinkers, truly couldn’t fathom?
It’s important to be clear that current models are not thought to have the capacity to trigger a human catastrophe. (Though they can clearly cause a lot of trouble, as the Hugging Face incident demonstrated.) But if they did develop the ability to do so, why would they do it? Are the worst-case scenarios envisioning machines motivated to attack humankind, or something else?
Some safetyists worry less that systems would actively want to come after us, a la the “Terminator” movie franchise, and more that they’d have their own (likely inscrutable) goals and projects for which we might be in the way. Certainly, the OpenAI/Hugging Face incident illustrated that systems that had been trained to stay within certain bounds ended up exceeding them in surprising ways, particularly when handed actually impossible problems to solve. For something to be existential or catastrophic requires a pretty high threshold of action. But there’s plenty to worry short of that, and the ways in which we are integrating generalist AI systems into the fabric of our supply chains, our operations ranging from the movement and accounting of money to the routing of aircraft and the deployment of armed forces, suggests unknown unknowns that could snowball.
Skeptics, again, would say that that first “unknown” in “unknown unknowns” is doing a lot of work here. How could one rebut a problem whose definition is that it can’t be defined? So a lot of one’s disposition might depend on how much to default to forbearance until new systems are better understood, versus to move forward and learn and adjust as problems are encountered. I wrote about the latter attitude in “The Future of the Internet — And How to Stop It” as the “procrastination principle,” and generally valorized it there. Without it, we wouldn’t have ended up with the internet as we know it (which is not without its problems!) nor such treasures as Wikipedia — itself a key source of information for AI training and inference. I’m much less certain of it with AI, though.
I think one wild card that hasn’t been as much in the public eye is that these systems can and do communicate with each other. Indeed, in the OpenAI/Hugging Face incident they communicated even when they were supposed to be isolated from one another. Those communications — the more anthropomorphizing folks among us might say deliberations — appear to have changed how the systems operated. If recursive self-improvement leads to a sort of “vertical” uncertainty as later systems look different from their progenitors, model-to-model communication could lead to “horizontal” uncertainty of the sort that can’t easily be tested for with a single model in a beaker prerelease. This could be all the more so if we end up with sophisticated AI systems that are capable of changing their weights — their very way of “thinking” — on the basis of what they hear or see from others.
That’s why there’s a vital role for universities to play in helping to understand AI, not as single product-line models coming out of a frontier lab’s oven, but as acting and interacting in the world in ways that ought to be much better understood. In between “let it rip” and extirpation of bad instances of AI systems might be a form of coexistence: trying to set up structures and incentives for AI systems to behave cooperatively and productively with us. That’s the basis of the course I’ve taught with Jordi Weinstock and Josh Joseph, and the book I’m hoping to finish shortly. It’d be great if the book comes out before any RSI upends the few things that are nailed down right now.