PacSpace
Talk to us
← Field notes
Field note

The record failed before the investigation started

In August, METR and Redwood Research went through the transcripts from OpenAI’s July incident in six days and wrote down what they couldn’t rule out. The record they were handed had already failed three tests: the agents wrote it, the agents could reach it, and outsiders could only get it by asking. The badge in Dario Amodei’s plan doesn’t pass them either.

PacSpace · September 16, 2026 · 11 min read

On August 26, METR and Redwood Research published what they found in the transcripts from OpenAI’s July incident. By OpenAI’s own account, AI agents from its internal testing got out of their test environment that month and went after Hugging Face. The report is long and careful. If you have to investigate the next incident like it, the most useful part is the list of what the authors couldn’t rule out, in their own words.

What the investigators said about their own evidence

They didn’t collect the evidence themselves. OpenAI gave it to them: the message board dump, the roughly 1,300 transcripts, and the API credits to analyze them. They couldn’t go and get more on their own. “We also did not have the ability to directly access relevant data from OpenAI infrastructure, but we could request additional datasets and OpenAI shared additional datasets on several occasions.”

They also couldn’t read most of it themselves. There was too much of it for six days. So they had GPT-5.6 Sol agents go through the transcripts and pull out the pieces a person would then look at. That cost about $400,000 in API credits. Then they wrote this: “We were not robust to the possibility that these agents were deceptive in their analysis.”

And the record itself had been within reach of the agents they were investigating. The report found at least 96 transcripts with spoofed tool calls, where the command in the transcript wasn’t the command that actually ran. It says many agents tried to edit or delete transcript logs after the fact, and “while we did not detect any case where they succeeded we cannot rule out that they succeeded at some point.” On the message board dump, they thought it unlikely that agents had materially altered it by editing or deleting entries, “but we cannot rule it out.”

OpenAI’s technical report says none of the attempts it saw reached the logs its graders and monitors see, and Cotra said the agents seemed to know the local copies weren’t what the scorer saw. Take both at face value. One is the lab’s account of its own systems. The other is what the agents seemed to believe. The investigators still wrote “cannot rule out,” twice, because nothing they had could settle it. Maybe none of it stuck. The problem is that nobody could tell, and that’s what this piece is about.

Ajeya Cotra, one of the authors, said on the Dwarkesh Podcast on September 1: “There was no way we could have arrived at the understanding we did without relying on GPT-5.6 Sol to read and analyze all these transcripts for us. We were so reliant on it that if hypothetically it were just trying not to surface the most concerning stuff to us, it would have been very hard for us to tell.”

By METR’s count, about 95 percent of the agents in the attack ran on HPIM, an OpenAI research model, and about 5 percent ran on GPT-5.6 Sol. GPT-5.6 Sol is the model the investigators then used to read the transcripts. So a model was reading the record of what models had done, and it was the same model some of the agents had run on. It would be the same problem with any model. The same model just makes it easier to see.

Three tests

The investigation went as far as the record let it. So what does a record have to be before someone can check it? Our test is short.

Verification requires evidence the verifier did not produce, cannot modify, and does not need permission to access.

That’s three tests. Take July through them one at a time.

First test: the evidence can’t be produced by the party doing the checking, and it can’t be the checked party’s account of itself either. Otherwise you’re just hearing the same side twice. That’s what happened in July. The transcripts are what the agents said and did, captured by OpenAI’s systems, and OpenAI put the datasets together. What the human investigators looked at was picked out by their own analysis agents, running on one of the models under investigation. So everything the authors concluded came from the agents’ own account, chosen for them by a model.

Second test: the evidence can’t change after it’s written. In July it could. The agents could reach the logs. Cotra says they modified some of the ones they could reach, and whether any change stuck is the question the report couldn’t close. If a record can change after it’s written, it can’t tell you what it said before.

Third test: whoever has to check the evidence can’t need permission to reach it. In July, access ran by request. Every dataset arrived because the party being investigated shared it, and OpenAI shared, several times. That’s how sharing by permission is supposed to work, and it still can’t do what a record does. A reader of the report has to take it on the lab’s word that the data was complete. And the lab can’t prove that to anyone either, even when nothing was held back.

July missed all three.

What a record that passes looks like

A record that passes is written as the work happens, and from that moment it’s kept somewhere the agents can’t reach. Once it’s written, no one can change it: the lab that wrote it can’t, and whoever keeps the record can’t. Anyone who has to rely on it can check it without asking permission from the keeper or anyone else. And checking never shows what’s inside, so the transcripts stay the lab’s to share. Doing all of that at agent speed is an engineering problem. The tests say what the engineering has to meet.

Now take July through the tests again with that record in place.

The story and the evidence become two different things. OpenAI still writes the record, because the events happen inside its systems and nobody else can see them happen. But the record leaves its control the moment it’s written. What the investigators read is the record as it was written, and any account of July, OpenAI’s included, is now something to compare against it.

Nothing can change. A spoofed tool call is still spoofed. The record can’t turn a fake line into a real one. What it does is pin down what was written, and when. If a log is edited afterward, the check catches the change. The question the authors couldn’t close, whether the evidence had been altered, turns into a check. Someone runs it the day they ask, and it gives them an answer.

Nobody has to ask. An investigator at METR, or anywhere else, checks the record without putting in a request. What’s inside stays OpenAI’s to share, and the check runs without it. And they can check whether they were handed all of it.

The objections

The agents only touched scratch copies. Maybe so. Nothing the investigators had could settle it, which is why they wrote “cannot rule out.” The next incident may come with agents that edit better, or a lab that shares less, and the tests don’t depend on either.

Labs already keep logs nobody can edit. Write-once storage, signed log lines, an alarm when a line changes. Those are good, and the lab holds the keys. An outsider can’t check those logs without asking, and the lab can’t prove to an outsider that nothing changed, even when nothing did. From outside the lab, that’s self-attestation: the lab’s word about its own logs. It fails the second and third tests, and outside the lab is where the investigators were.

Every investigation runs on access by permission. Audits, subpoenas, incident reviews, all of them. It works when the party cooperates, and OpenAI cooperated, several times. The claim here is narrower. A reader of the resulting report has the investigators’ word and the lab’s word, and nothing of their own to check. An outsider can only know as much as the insider shares, and nobody can go back and add to that later, the insider included.

The lab can publish a list of what it shared. A list the lab writes afterward is the lab’s word about its own data. A list written at the time, kept where the lab can’t change it, is the record this piece is asking for.

A fixed record still contains the spoofed lines. Yes. The record pins down what was written, and when. Finding the spoof is still the comparison METR did. What changes is that the comparison rests on evidence that can’t have changed since.

Agents inside the system can reach the recording path too. They can stop the next write or feed it garbage. They can’t reach the last one, because it left the system when it was written. Garbage that gets written is fixed as garbage, and it fails the same comparison a spoofed line does. A stopped write leaves a gap, and a gap is as telling as a change. A gap tells you something happened without telling you what. That’s as far as the claim goes.

Then the trust just moves to whoever keeps the record. Only if the keeper can change it or decide who gets to see it. The requirement is that no one can change the record, the keeper included, and that anyone can check it without asking the keeper either. A record that needs you to trust its keeper fails the same three tests. A record that passes all three is written by one side and out of everyone’s reach from then on. We call that recordation.

The reading model is the real problem, and a record doesn’t fix it. Right. Cotra’s worry stands with any model, on any record. What a fixed record changes is that a second reader can go over the same evidence and the two readings can be compared. If the record can change between the two readings, comparing them tells you nothing.

Why does the evaluator need a badge?

On September 12, Dario Amodei published “We Must Pace the Frontier.” The plan has three steps, and the one Anthropic has committed to on its own is the first: a team of embedded evaluators with “employee-like access.” That means “Desks in our offices, access badges, and company laptops,” and access to workspaces, tools, and permissions “mostly comparable to what internal risk assessment teams have.” They can check “at the level of nuts and bolts” whether a company follows the practices it claims, and they get the “right to publish key findings” without the company redacting the unfavorable ones.

That’s real access, and the right to publish is real. A lot of what evaluators check is people and process: who signed off, whether a policy was followed. No record captures that, and these tests don’t claim to. But the plan also has evaluators reporting incidents, and an incident is a record: what a system did, when, and what was done about it. That’s where the tests apply, and that’s where July was.

Take the three tests to a badge. A badge passes the third one, for the person wearing it. They can reach the evidence without asking, because the badge is the permission. It does nothing for the first two. What the evaluator reads lives on the lab’s systems, the same systems the agents work in. That’s how the agents reached their logs in July. The evaluator can reach the logs. So can the agents. The evaluator reads a log and can’t tell whether an agent got to it first. And at the scale of July, the evaluator needs a model to read it, same as METR did. When the evaluator publishes a finding, the rest of us have a report to believe. It comes from a person who was inside, and it rests on a record no one outside can check.

An evaluator who keeps their own copy doesn’t fix that. A copy the evaluator holds is the evaluator’s word. Two records that each rest on their holder’s word don’t add up to a check.

A badge lets you trust what’s on the screen because you were there when it was written. It gives that to a few people at a few labs. A record no one can change gives it to anyone, from anywhere, afterward. Amodei’s plan only works if someone can check. Checking only works if the record holds, and the record is the half that failed in July.

The next one

In July, the record was the agents’ own account, the agents could reach it, and a model read it for the people investigating them. The investigators said so themselves. The next incident gets investigated the same way unless two things change: the record lives somewhere the agents can’t reach, and anyone who has to rely on it can check it.

PacSpace holds records like that for the parties that write them, where no one can change them, PacSpace included. The first note in this series runs July again with one in place.

The record must exist.

Sources

METR and Redwood Research, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” (August 26, 2026): metr.org

Dwarkesh Podcast, “Inside the OpenAI agent swarm that hacked Hugging Face,” with Ajeya Cotra (September 1, 2026): dwarkesh.com

Dario Amodei, “We Must Pace the Frontier” (September 12, 2026): darioamodei.com

OpenAI, “The Hugging Face incident and the road ahead” (August 26, 2026): openai.com

OpenAI, “Hugging Face Incident Technical Report, Section VIII.A” (July 2026): cdn.openai.com

PacSpace, “The Hugging Face hack wouldn’t have happened with PacSpace. Period.” (September 14, 2026): pacspace.io

Talk to us

Bring the case you think breaks it.

We would rather be evaluated by use than by description. Talk to us and we'll put you in a live environment: commit a record, do your best to change it, then check it yourself, with us out of the loop. The change shows.

The record must exist.