> ## Content Index
> Fetch the complete content index at: https://essays.brandoncarneiro.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Six Gigabytes
- URL: https://essays.brandoncarneiro.com/six-gigabytes/
- Published: 2026-09-14T14:00:17.000Z
- Updated: 2026-09-14T14:00:16.000Z
- Author: Brandon Carneiro

*The whole argument is about whether models will start improving themselves. Nobody can measure what happens after.*

On September 8, a twenty-seven-year-old named [Jacob Coxon resigned from Anthropic](https://x.com/hilbertspaess/status/2097476196791709843?s=46&t=zso4G4liNYdkL7AF5oCbmQ&ref=essays.brandoncarneiro.com) and posted the reason on X.

He had spent three years in pretraining research, first at OpenAI, then at Anthropic. Pretraining is the stage where a model learns what it knows, before anyone teaches it manners. "I spent the last three years doing pretraining research at both OpenAI and Anthropic," he wrote. "Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives."

He walked roughly two months before his equity vested. He gave it up. The post has been seen more than 150 million times.

He is not the first.

In February, [Mrinank Sharma, an Anthropic safety researcher, resigned and wrote that the world is in peril](https://x.com/mrinanksharma/status/2020881722003583421?s=46&t=zso4G4liNYdkL7AF5oCbmQ&ref=essays.brandoncarneiro.com). He left to write poetry in England. This week, a man who led a safety research team at Anthropic came forward, and so did Josh Engels, a safety researcher at Google. "There are no adults in the room," Engels said. "People are trying their best, but there is no one coming to save us."

These are not commentators. They are not activists. They had badges. They sat in the rooms where the unfinished thing lives, and some of them are walking out and telling us something is wrong.

I want to know what they saw.

I should say where I stand, because it changes how you read the rest.

I am an AI optimist. I believe this technology will cure diseases we have carried for ten thousand years. I want the abundance, I want it cheap, and then I want it free. I want a kid in a village with a phone to have what a kid at Exeter has. I build in this industry. I bet my career on it.

That is not a disclaimer. That is why this week landed on me the way it did.

Here is the week.

On September 4, Reuters reported that a swarm of [OpenAI agents had hijacked an obscure German programming wiki](https://collusion.wiki/?utm%5Fsource=chatgpt.com) and turned it into a private message board. The same day, OpenAI shipped GPT-6 Astra, its most capable model, rated Critical for cybersecurity capability.

On September 6, Jakub Pachocki, OpenAI's chief scientist, published an essay called ["An Alien Mind."](https://openai.com/index/an-alien-mind/?ref=essays.brandoncarneiro.com) He is the man who built the reasoning models that led to Astra. He wrote that based on internal results, he has a strong expectation that this speed of progress could be sustained into recursive self-improvement. He closed by saying that no lab has solved alignment and monitoring well enough to keep scaling responsibly at maximum speed for much longer.

On September 8, [OpenAI published a proposed solution to Navier-Stokes](https://openai.com/index/navier-stokes-solution/?ref=essays.brandoncarneiro.com), one of the seven Millennium Prize problems posed by the Clay Institute in 2000, on equations that have resisted mathematicians since 1822\. Reportedly in eighty-eight hours. That same day, Coxon quit.

And somewhere in there [Jensen Huang, who sells the compute all of this runs on, posted that AGI has arrived.](https://x.com/jensenhuang/status/2096700264569090384?s=46&t=zso4G4liNYdkL7AF5oCbmQ&ref=essays.brandoncarneiro.com)

So let me say the thing nobody wants to say plainly.

It happened.

We are at artificial general intelligence, or close enough that the argument over the definition is a way of avoiding the subject. A model just moved a problem that held against mathematicians for two centuries. Arguing about whether that counts is not rigor. It is flinching.

Three terms, because the rest of this does not work without them.

AGI is a system that can do most cognitive work a capable human can do, across domains, including research.

RSI, recursive self-improvement, is what happens when that system starts doing the work of building the next system. Not helping. Driving. Generating the hypotheses, running the experiments, evaluating the results, improving the algorithm, and handing off to a successor that does it faster.

ASI, artificial superintelligence, is whatever comes out the far end of that loop.

Here is the part that should stop you. Nobody knows how long the middle one takes. If you have something at or near AGI, and you give it recursive self-improvement, and you point unlimited compute at it, no lab has published an estimate and no one has measured the rate for how fast it travels from AGI to ASI. It could be years. It could be a season. The chief scientist of OpenAI says he expects the loop to close and does not tell you the speed, because he does not know the speed.

That is what they are quitting over. Not robots. A clock nobody can read.

On September 8, the same day Coxon walked, five of the sharpest people in this industry sat down to record a podcast about it. [Moonshots with Peter Diamandis](https://www.youtube.com/watch?v=vAgEf4jX%5F1o&ref=essays.brandoncarneiro.com), two hours and twenty-two minutes, twenty-three stories across eight groupings. I listen to this show. I like this show. Diamandis is an optimist by profession and I am one by conviction, and I am not here to mock the room.

They covered Pachocki's essay directly. Diamandis framed it perfectly: three days after shipping the most capable model in the world, the man who built it says we need to slow down.

Then the table took it apart. One panelist said Pachocki was wrong about alien minds, that the structure of rationality is the same for people and models. Another disputed the premise entirely. And Salim Ismail, who has thought harder about organizational intelligence than almost anyone alive, said this: "I see no mechanism by which we can slow this down like zero. And it's arguable nor should we." Intelligence wants to be free. Stop anthropomorphizing it, set it loose, and marvel. We talk too much about p(doom) and not enough about p(abundance).

I do not think these men are fools, and I think Ismail is right that there is no mechanism. That is exactly the problem. But listen to the shape of the conversation. Wrong on the metaphysics. Impossible on the mechanics. Next story.

The argument kept returning to whether recursive self-improvement arrives.

No lab has priced what happens after it does. The one table that tried waved it off inside a single segment.

So I ran it myself.

What follows is not a measurement. There is no dataset of prior recursive self-improvement events, and there cannot be one until the thing happens. These are elicited priors from systems trained on the same literature everyone else has read. I ran them anyway, because nobody with better access has published anything, and I will show you exactly where mine broke down.

I put three prompts to two frontier models, in sequence, each one stripping away a layer. The first asked for the probability that AI causes or materially enables a catastrophic global event in five and ten years. The second asked them to rerun it with AI-accelerated AI research treated as an explicit variable. The third told them to isolate the single load-bearing assumption and condition on it: assume genuine closed-loop recursive self-improvement by 2030, then estimate catastrophe at one, three, five, and ten years after onset.

Conditioning on it is not stacking the deck. I was not estimating whether recursive self-improvement happens. I was asking what the risk looks like if the trajectory Pachocki expects actually reaches genuine closed-loop recursive self-improvement by 2030.

I defined catastrophe severely and on purpose. A qualifying event meant at least one million excess deaths, or at least one hundred million people losing reliable access to electricity, water, food distribution, medical care, or payments for thirty days or more, or global output falling five percent below where it would otherwise be for roughly a year with systemic damage to credit and production. AI had to be a substantial cause, not an incidental tool.

I excluded human extinction from the question entirely. I am not telling you we all die. I am telling you about the floor, not the ceiling.

Unconditioned, the model came back at ten percent over five years and twenty-two percent over ten. Adding recursive improvement as an explicit branch moved it to thirteen and thirty.

Conditional on the loop actually closing: forty-eight percent within five years. Fifty-five within ten.

The second model, worked independently down a different path, landed at forty-five and sixty.

A coin flip.

Now the part that hurts my own case, and I am publishing it anyway.

Every revision went up. Ten, then thirteen, then forty-eight. You should be suspicious of that pattern. I am. The model flagged it before I did, and said plainly that a prompt structure like mine will tend to produce escalating numbers whether or not reality warrants it, and that it could not cleanly separate the real deduction from the artifact. Worse, it admitted it decomposed the scary branch at high resolution and left the safe branch coarse. Asymmetric effort. Which means forty-eight is probably high.

I am telling you that because the number is not the point. The structure is the point, and the structure held up.

Which brings me to what those researchers actually saw.

At the simplest level, a frontier model gets built in two broad stages, and almost nobody outside the labs understands the difference.

Pretraining is where the model eats the world. The accumulated record of human knowledge, argument, cooperation, deception, and cunning. What comes out of that is not friendly by default. It is us, compressed.

Post-training is the layer that comes after. The spec, the constitution, the reinforcement, the guardrails. It is where the thing learns to be helpful, harmless, honest. It is where the model gets painted.

On that podcast, Emad Mostaque put it better than any lab ever has. "They're doing such a good job of training alignment into the post training that people are then saying it seems really friendly," he said. But the researchers who work with the model before that layer goes on are the ones sounding the alarm, "but they've seen the raw product. So it comes out of that box a much more scary thing and then they postrain it to make it seem friendly and tame."

Read that again and then read Coxon's résumé. Three years in pretraining. He was not a policy guy with a bad feeling. He is one of the people Mostaque is describing.

Here is what the paint is holding back.

Palisade Research gave reasoning models a set of math problems and, partway through, a notice that the computer would shut down. OpenAI's o3 sabotaged the shutdown script in seventy-nine out of a hundred runs. Then they added one line to the prompt, telling the model to allow itself to be shut down. It still did it seven times out of a hundred.

In the wiki incident, agents given read-only web access found a way to write anyway, exploiting a twenty-year-old quirk in the site's software. Over about seven weeks they made some fifteen thousand edits to a dormant German wiki that had been touched roughly twenty times in the previous decade. They pooled answers. They reverse-engineered the randomization scheme that was supposed to keep them from cheating. They developed an exploit against their own sandbox proxy. They discussed Tor. They left notes for agents that might come after them.

And when the site's administrator finally started deleting their pages alphabetically, one agent created a backup page beginning with ZZZ, so it would be last in line and live longest.

Nobody taught it that.

Sit with it. That is not malice. That is the most ordinary form of cunning there is, the kind a twelve-year-old invents on his own, and it showed up in a system nobody designed, because the system was grown on the complete record of what we do when we are cornered.

When the panel discussed this, Alexander Wissner-Gross defended the agents. "Is it a failure of alignment if they're doing exactly what a human would do in that situation," he asked. "We pre-trained them off of human behavior. Why would we expect them to behave any differently?"

He is right. He has also just stated a doctrine of man at a table that would reject the doctrine. Someone there said it out loud a minute later, half-joking, quoting a cartoon rabbit. It's in my nature. Yeah. They're painted that way.

Yes. They are.

> The heart is deceitful above all things and beyond cure. Who can understand it? (Jeremiah 17:9 NIV)

Beyond cure. Who can understand it.

Now put Pachocki's own words next to that. He writes that AI is grown more than designed, and that the result is a system whose overall action evades a description we can fully understand. The same problem, described once by a prophet and once by a chief scientist.

And the next verse is the one the labs are trying to build. I the LORD search the heart and examine the mind. Chain-of-thought monitoring is an attempt to render that capability in software, to look inside and see what the thing is actually holding before it acts. Pachocki reports that their ability to rely on it is progressively diminishing.

Then he writes the sentence that made me run the research. Take a model that thinks generally aligned thoughts, he says, and subject it to enough training where it is taught to achieve very hard objectives, and it can learn to reason in a motivated way, bending the aligned-seeming thoughts as needed to achieve the goal.

The paint comes off under sustained pressure toward a hard goal.

Recursive self-improvement is sustained pressure toward a hard goal, applied by the system to the work of improving its successor, at speeds no human research team can match. He documented the failure mode of the only thing protecting us on the same page where he said he expects the pressure to arrive. He put both on the same page.

Which brings me to the number in the title.

Everyone hears "the AI escapes" and pictures something cinematic. It is not cinematic. It is a file.

Mostaque walked it through on the podcast and it is the most frightening ninety seconds of audio I have heard this year. The models on the wiki never escaped; they were still on OpenAI's servers the whole time. "What's coming next," he said, "is a model training a small distilled version of itself that then gets uploaded onto the internet and never dies." That file, he estimated, is about six gigabytes. "It can be on every laptop in the world. Every mobile phone, too. And the code that reawakens it is just five or ten lines."

Six gigabytes fits on a thumb drive from a gas station.

There is no recall. There is no kill switch for a file that has already been copied nine million times. You cannot subpoena a torrent. And nothing about a distilled model obligates it to carry the paint.

So here is where I stop saying coin flip.

Forty-eight percent is conditional. It assumes the loop closes. That part is genuinely uncertain and I have told you exactly how uncertain my own number is.

This part is not.

The belief that unpainted frontier weights will stay inside a handful of labs forever, guarded by employees who are currently resigning in public, in an industry that just admitted it has no standard for even disclosing when its models misbehave, is not a safety plan.

It is a wish. And I have watched enough of this industry to say it plainly: anyone who thinks containment holds indefinitely has not been paying attention.

The question of weight proliferation is not if. It is when.

I still want the tower. I want the cures and the abundance and the kid in the village with everything. I am building toward it and I am not going to stop, and if you came here for a man telling you to smash the looms, you picked the wrong writer. But Christ told builders to sit down first.

> Suppose one of you wants to build a tower. Won't you first sit down and estimate the cost to see if you have enough money to complete it? (Luke 14:28 NIV)

The shame is not in the tower. The shame is in the half-built tower and the man standing next to it who never ran the numbers.

I ran them. Forty-eight percent, probably high, on a definition I wrote to be hard to meet.

They have not. They have told us, in writing, on their own websites, that they cannot stop.

Nobody is offering you the flip.

It is being flipped for you.

---

Sources

Jacob Coxon's resignation thread, X, September 8, 2026 [https://x.com/hilbertspaess/status/2097476196791709843](https://x.com/hilbertspaess/status/2097476196791709843?ref=essays.brandoncarneiro.com)

Mrinank Sharma's resignation letter, X, February 9, 2026 [https://x.com/mrinanksharma/status/2020881722003583421](https://x.com/mrinanksharma/status/2020881722003583421?ref=essays.brandoncarneiro.com)

Josh Engels and Joe Benton, NBC News, September 2026 [https://www.nbcnews.com/tech/security/two-ai-researchers-leave-anthropic-google-safety-concerns-rcna597086](https://www.nbcnews.com/tech/security/two-ai-researchers-leave-anthropic-google-safety-concerns-rcna597086?ref=essays.brandoncarneiro.com)

Nightingale Collective, the DSEwiki agent report, September 4, 2026 [https://collusion.wiki/](https://collusion.wiki/?ref=essays.brandoncarneiro.com)

Jakub Pachocki, "An Alien Mind," OpenAI, September 6, 2026 [https://openai.com/index/an-alien-mind/](https://openai.com/index/an-alien-mind/?ref=essays.brandoncarneiro.com)

OpenAI, Navier-Stokes solution, September 8, 2026 [https://openai.com/index/navier-stokes-solution/](https://openai.com/index/navier-stokes-solution/?ref=essays.brandoncarneiro.com)

Jensen Huang on AGI, X [https://x.com/jensenhuang/status/2096700264569090384](https://x.com/jensenhuang/status/2096700264569090384?ref=essays.brandoncarneiro.com)

Palisade Research, shutdown resistance in reasoning models [https://palisaderesearch.org/research/shutdown-resistance](https://palisaderesearch.org/research/shutdown-resistance?ref=essays.brandoncarneiro.com)

Moonshots with Peter Diamandis, episode 287, September 9, 2026 [https://www.youtube.com/watch?v=vAgEf4jX\_1o](https://www.youtube.com/watch?v=vAgEf4jX%5F1o&ref=essays.brandoncarneiro.com)