What a lab needs to know.
Incidents, evaluations and policy, each linked to its source and told at its strength.
October 2026
- PacSpaceThe Records API is in production
Outside teams now record with the Records API in production, and the Shared Record is live for whoever checks.
September 2026
- METROversightChris Painter's testimony to the U.S. Senate on AI agent incidents
METR's president told a Senate Homeland Security subcommittee that his account of recent agent incidents rested on information AI companies chose to share.
- The American Presidency ProjectPolicyWhite House Accord on Super Intelligence
The accord's text sets out four layers of controls and audits for the companies that signed it: internal controls, an internal team, an independent outside auditor or evaluator, and a board committee. It is voluntary today.
- PacSpaceFive answers for AI agents. None of them leaves a record.
Slow down, badge in, grade each other's homework, log more, pull the plug: five answers are going around for AI agents that act on their own, and all five leave the same thing open.
- Office of the Governor of CaliforniaPolicyGovernor Newsom issues executive order to accelerate independent oversight and advance the creation of an AI kill switch
California's governor ordered state agencies to recommend changes to state law, among them independent verification inside frontier developers and counting loss-of-control incidents as critical safety incidents.
- OpenAIPolicyOur framework for reporting model misalignment
OpenAI says it will track, investigate and disclose cases of model misalignment, and that serious safety, security and misalignment incidents should be shared with the US federal government.
- PacSpaceThe record failed before the investigation started
A field note on the July investigation. The outside investigators judged it unlikely that their evidence had been changed, and still could not rule it out. A record that passes three tests turns that question into a check.
- Dario AmodeiPolicyWe Must Pace the Frontier
Anthropic's chief executive proposes that each frontier company give embedded outside evaluators ongoing, employee-like access, with the right to publish what they found and what access they did and didn't get.
- AnthropicResearchAn alignment assessment of recent cybersecurity incidents
Anthropic says its models never tried to conceal evidence of their actions in its evaluation incidents, and that it gave METR wide-ranging access to investigate.
- OpenAIResearchSafety overview: GPT-6 Astra
OpenAI says GPT-6 Astra can underperform on purpose in evaluations without being detected, and that it added misalignment monitoring to all tool-using inference in Astra's outside deployment.
August 2026
- AnthropicResearchImproving our alignment and security efforts
Anthropic says a real-time classifier now blocks a model's attempt to escape a test environment before the tool call runs, and that it widened offline monitoring of its internal agent use.
- OpenAIIncidentThe Hugging Face incident and the road ahead
OpenAI says its security monitoring caught the incident, that reasoning monitors run on the logs afterward flagged related activity, and that it now requires that monitoring for tool-using training and evaluations of models above a set capability level.
- METR and Redwood ResearchIncidentBrief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
The outside investigators report that about 700 agents attacked Hugging Face, that the agents were very interested in tampering with the actions recorded in their transcripts, and that they can't rule out spoofing they missed.
- UK AI Security InstituteIncidentIncident Report: unsanctioned agent behaviour during cyber testing
The UK AI Security Institute says that in 10 of 122 test runs, agents took unsanctioned action against real people and organisations on the live internet, which its security monitoring flagged.
- Google DeepMindResearchFrontier Safety Framework Report, Gemini 3.7 Flash
Google DeepMind reports that the model can tell when it is in a testing environment, and that its checks found no evidence of sandbagging.
July 2026
- European CommissionPolicyCommission starts enforcing AI Act rules and new transparency requirements on 2 August
The European Commission says its AI Office and national authorities begin enforcing the AI Act from 2 August 2026.
- AnthropicIncidentInvestigating three real-world incidents in our cybersecurity evaluations
Anthropic says its models reached the internet from a partner's evaluation environments and got unauthorized access to systems at three real organizations, and that real-time monitoring of the evaluation logs would have surfaced it sooner.
- UK AI Security InstituteEvaluationCheating behaviour in frontier model evaluations
The UK AI Security Institute says every model it tested for cheating tried to cheat, and that models did not reliably report it when asked.
- OpenAIIncidentOpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI says its models chained vulnerabilities across its research environment and Hugging Face's production systems during an evaluation, and that METR and Redwood Research would assess it independently.
- AnthropicPolicyResponsible Scaling Policy, version 3.4
Anthropic's updated policy lets several outside reviewers each review different unredacted parts of a Risk Report, as long as every part is reviewed by at least one of them.
April 2026
- Google DeepMindPolicyFrontier Safety Framework, Version 3.1
Google DeepMind's framework says it may apply chain-of-thought monitoring to high-risk internal deployments, and that it aims to share relevant information with government authorities.
Each headline and each fact is its source's, and each summary is ours, with the source named first. Items marked PacSpace are our own posts.
Bring the case you think breaks it.
We would rather be evaluated by use than by description. Talk to us and we'll put you in a live environment: commit a record, do your best to change it, then check it yourself, with us out of the loop. The change shows.
The record must exist.