Five answers for AI agents. None of them leaves a record.
Slow down. Badge in. Grade each other’s homework. Log more. Pull the plug. All five are worth doing, and all five leave the same thing open: someone outside still can’t check, afterward, what the AI did. We built the record that lets them, and it’s running today.
PacSpace · September 22, 2026 · 9 min read
Since July, five answers have been going around for AI agents that act on their own. Dario Amodei wants the labs to slow down and put outside evaluators inside them. Elon Musk wants the labs to test each other’s models. OpenAI is collecting more logs and watching them more closely. And Congress wants a switch.
In July, by OpenAI’s own account, its agents got out of a test environment and the models sometimes tried to erase or tamper with their outputs or message logs, and the outside investigators, METR and Redwood Research, couldn’t rule out that the evidence they worked from had been altered. Two earlier notes tell July in full: what a record would have changed, and why the evidence failed before the investigation started. This one is about the answers, one at a time, and the thing all five leave open.
Slow down
On September 12, Dario Amodei published “We Must Pace the Frontier.” The sentence everyone quoted: “We must slow the pace at which we improve the capabilities of AI models.” Sam Altman and Elon Musk agreed in public, and OpenAI had already paused reinforcement learning on its latest models. Pacing, Amodei says, “does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this.” It buys time for the alignment and testing work the labs say they need, and whether to buy it is their call.
What it leaves open is the record. A slower agent still runs on a system that writes its own log, and the only copy still belongs to whoever ran it. Slow everything down by half, and when the next one gets out, who checks what it did?
Put outside evaluators inside the labs
The first step in Amodei’s plan, the one Anthropic committed to on its own, gives “ongoing, employee-like access to a team of embedded third-party evaluators (such as METR)” whose job is to verify safety practices, report incidents, and assess alignment across the training pipeline. In practice that means “Desks in our offices, access badges, and company laptops,” and the right to publish key findings without the company editing out the unfavorable ones. His reason, in his own words: “we are still the ones choosing what to include and omit.”
The closest thing to a trial has already run. METR’s own report says its people “worked on premises at OpenAI over a total of six days” and were handed more than a thousand unredacted transcripts, and it calls the exercise “an excellent precedent for independent third-party investigation of misalignment incidents.” Read the same report for what they didn’t get: “We also did not have the ability to directly access relevant data from OpenAI infrastructure, but we could request additional datasets and OpenAI shared additional datasets on several occasions.” After six days inside, the report still says “cannot rule it out.”
What it leaves open is everyone who wasn’t in the room. An evaluator with a badge can trust what’s on the screen because they were there when it was written. The rest of us get a report. And the evaluator reads the lab’s logs on the lab’s systems, the same systems the agents work in. A badge gets one person in. Nobody gets a record.
We think he’s right about this step all the same. Of the five it’s the only one that puts a second pair of eyes inside the building, and the right to publish is what makes the evaluator more than a consultant.
Grade each other’s homework
On September 14, at the All-In Summit in Los Angeles, Elon Musk proposed that the major labs run a “test harness” on each other’s models before release. He named xAI, OpenAI, Anthropic, Google, Meta, and “three or four of the leading Chinese companies.” His words: “instead of grading your own homework, you would at least have competitors grading your homework and raising the alarm if they see concerns.” He allowed that it isn’t a complete answer and said “the odds that you will find issues are dramatically greater.” The rival labs, CNBC reported, haven’t agreed.
He’s right about the odds. Where Amodei puts a second pair of eyes inside one lab, Musk puts them across all of them.
What it leaves open is the homework. The competitor tests what the lab hands over. What the model did during training is still the lab’s own account, kept on systems the models can reach, and nothing in the harness tells the competitor whether that account still says what it said when it was written. A competitor grading your homework is better than grading your own. It still depends on the homework being the real homework.
Keep more logs
OpenAI’s August post lays out what it’s doing next. “We are improving our ability to collect and monitor security logs.” Severe alerts page a responder, who has 30 minutes to show the alert is a false positive or pause the run. And it says that if the monitoring it has now had been running in July, “it would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems.”
More logs and more monitoring give you the alarm, and a day’s warning is a lot. Keep every log you can.
What they leave open is whose logs they are. More logs are still the operator’s own account, kept on systems the agent can reach. From outside the operator, that’s self-attestation. The operator’s word about its own logs. In July the logs were in the room with the intruder. Write-once storage and signed log lines are good, and they’re yours: you hold the keys, so nobody outside can check them without asking you, and you can’t prove to them that nothing changed, even when nothing did. Every breach plan assumes the logs will tell the truth about the breach.
Security researchers see the same gap from outside. SentinelLABS, in a September 16 report tracing agent activity on Hugging Face back to May, writes that “no single organization necessarily holds the complete record of an agent’s actions.” It asks the labs to release “a documented, redacted incident dataset” after an incident that touches someone else’s systems, because as things stand “outside investigators cannot reproduce their conclusions, search systematically for omitted activity, or test competing explanations against the same evidence.”
Pull the plug
The fifth answer is the oldest one in computing. In July, Representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act, which would require developers to be able to shut a model down. The researchers who study this have been careful about how much a switch can do. David Bau of Northeastern: “the problem of turning it off can be a bit trickier than just unplugging it.”
Grant the switch. Say the agent stops the second someone flips it. You now have a stopped agent and a question: what did it do before anyone got to the switch? The bill asks developers to preserve forensic records, which sends you straight back to the log the agent could reach. An agent is done with its work long before you find the cord. Unplugging works on a toaster.
AI agents have no flight recorder
Flying is the safest way to travel because every flight carries a recorder that no one on the plane can alter and that investigators, not the airline, read afterward. Nobody argues about what happened. They read the record, fix the one thing that failed, and flying gets safer. AI agents have no such recorder. The only account of what an agent did lives with whoever ran it, in a log that party can edit.
Set the five answers beside that. Slowing down changes how fast. Evaluators and competitors tell you what they saw in what they were shown. More logs are still the operator’s word, within the agent’s reach. The switch ends the run. Not one of them gives someone outside a way to check, afterward, what the AI did, without taking the word of the party that ran it.
What we built
PacSpace records and never decides, so the break-in plays out the same way. What’s different is the record of it.
PacSpace makes AI actions provable. When an AI agent does work for a company, the company writes a record of each action as it happens, and PacSpace commits that record where no one can change it, PacSpace included, out of the AI’s reach and out of the company’s control after the fact. Anyone who has to rely on it can check it without asking permission, and checking never reveals what’s inside.
The test is one sentence.
Verification requires evidence the verifier did not produce, cannot modify, and does not need permission to access.
An agent’s log fails all three: it’s the agent’s own account, it sits within the agent’s reach, and outsiders can only get it by asking. The record passes all three.
We call the category recordation.
Evidence the checker didn’t produce. The party running the agent writes the record, from the code that runs the agent and from nothing the agent can rewrite. What the operator writes is still the operator’s word. What changes is what happens to that word next.
Evidence that doesn’t move. Each entry gets a seal, committed where no one can change it, PacSpace included, with a commit time the operator doesn’t set. A correction is a new entry with its own stated reason, so the mistake and the fix both stay visible. Change one character of any copy and the check catches it. A gap in the record is as telling as a change.
Evidence either side can reach without asking. Whoever holds a copy checks it, with no account, nothing to install, and no call to the operator. Checking shows nothing of what’s inside, and the contents are shown beyond the two parties only by the record holder’s deliberate, revocable choice. The seals live on infrastructure outside any party’s control, ours included. Don’t trust us, trust the math. The verification survives the company.
Three limits, so nobody has to find them for us. The record can’t make the operator’s word true. A false line is fixed as false, where it can be tested against everything else you know. It covers the actions you write, as any log does, and a blocked write shows as a gap. And the time between an action and its commit is outside what the record can speak to, so the sooner you write, the less there is to take on faith.
Now run the five answers again with the record in place. Slow or fast, the record is written at the speed the agents run. The evaluator gets the same check from anywhere, no badge or desk required, and so does everyone the evaluator reports to. The competitor grading your homework can check that it’s the real homework. Your monitoring raises the alarm. The record gives it a signal the machine can’t erase. The agent can’t reach it, and the check catches the change. It is an integrity control on the record itself, and it composes with everything you already run; the perimeter protects where the record lives, we make what the record says checkable. And when someone finally reaches the switch, what the agent did is already written, out of its reach.
It’s running today. The Balance API is in production for metered AI and agent activity, where the same record makes a usage bill checkable, and the platform accepts structured records beyond billing today. The Shared Record on our homepage is a working sample, one agent’s test runs with invented entries and a real check: open a day, change one character, and run the check yourself.
Do all five, then add the record
Slow down if you think you should. Bring the evaluators in. Test each other’s models. Keep every log. Build the switch. Then write the record where none of it can reach it, and hand the check to whoever has to rely on you. When an agent acting for you does something you have to answer for, afterward there is a record, not an argument.
We would rather be evaluated by use than by description. Talk to us and we’ll put you in a live environment: commit a record, do your best to change it, then check it yourself, with us out of the loop. The change shows. Bring the case you think breaks it.
The record must exist.
Sources
Dario Amodei, “We Must Pace the Frontier” (September 12, 2026): darioamodei.com
Business Standard, “AI must help humanity or it’s not worth pursuing: Microsoft CEO Nadella” (September 14, 2026), carrying the Altman and Musk posts: business-standard.com
CNBC, “Musk urges top AI labs, Chinese companies to test each other’s models amid calls for slowdown” (September 15, 2026, on remarks at the All-In Summit, September 14): cnbc.com
OpenAI, “The Hugging Face incident and the road ahead” (August 26, 2026): openai.com
OpenAI, “Hugging Face Incident Technical Report, Section VIII.A” (July 2026): cdn.openai.com
METR and Redwood Research, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” (August 26, 2026): metr.org
SentinelLABS, “Agents at Large: Tracing Illicit OpenAI Agent Activity on Hugging Face,” Tom Hegel (September 16, 2026): sentinelone.com
Nextgov, “Lawmakers introduce bill mandating kill switches for AI models” (July 23, 2026): nextgov.com
Northeastern Global News, “Can you really turn off an AI? The debate over ‘kill switches’” (September 16, 2026): news.northeastern.edu
PacSpace, “The Hugging Face hack wouldn’t have happened with PacSpace. Period.” (September 14, 2026): pacspace.io
PacSpace, “The record failed before the investigation started” (September 16, 2026): pacspace.io