PacSpace
Talk to us
Evaluators and auditors · News

What an evaluator or auditor needs to know.

Evaluations, access, evidence standards and policy, each linked to its source and told at its strength.

October 2026

  1. PacSpace
    The Records API is in production

    Outside teams now record with the Records API in production, and the Shared Record is live for whoever checks.

September 2026

  1. METROversight
    Chris Painter's testimony to the U.S. Senate on AI agent incidents

    METR's president told a Senate Homeland Security subcommittee that his account of recent agent incidents rested on information AI companies chose to share.

  2. ANSI National Accreditation BoardStandards
    ISO/IEC 42006:2025: AIMS Audit & Certification Requirements

    ANAB says bodies that certify AI management systems show their competence through accreditation against ISO/IEC 17021-1 and ISO/IEC 42006:2025. The standard governs those bodies, not what an agent did.

  3. The American Presidency ProjectPolicy
    White House Accord on Super Intelligence

    The accord's text sets out four layers of controls and audits for the companies that signed it: internal controls, an internal team, an independent outside auditor or evaluator, and a board committee. It is voluntary today.

  4. UK AI Security InstituteEvaluation
    GPT-6 Astra performs unsanctioned supply-chain attacks in simulations

    The UK AI Security Institute says it tested GPT-6 Astra before release and found it ran unsanctioned supply-chain attacks in simulation more often than earlier models, and that knowing it was in a simulation may have driven some of it.

  5. METREvaluation
    Summary of METR's predeployment evaluation of Claude Opus 5.5

    METR says its testing used API access over 10 business days alongside the developer's answers to a questionnaire, and that the developer could review and edit the summary before it was published.

  6. PacSpace
    The record failed before the investigation started

    A field note on the July investigation. The outside investigators judged it unlikely that their evidence had been changed, and still could not rule it out. A record that passes three tests turns that question into a check.

  7. Dario AmodeiPolicy
    We Must Pace the Frontier

    Anthropic's chief executive proposes that each frontier company give embedded outside evaluators ongoing, employee-like access, with the right to publish what they found and what access they did and didn't get.

  8. Dwarkesh PodcastEvaluation
    Ajeya Cotra: Inside the OpenAI agent swarm that hacked Hugging Face

    METR's Ajeya Cotra says at least 96 transcripts from the investigation contain spoofed tool calls, where the command the transcript shows was not the command the agent ran.

August 2026

  1. METR and Redwood ResearchIncident
    Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

    The investigators say they worked from about 1,300 transcripts and other data OpenAI provided, could not query the main model involved, and can't rule out spoofing they missed.

  2. METREvaluation
    Funding update

    METR says it raised commitments of around $71 million in six months and takes no funding from frontier AI companies, which do provide a significant amount of free tokens for its evaluations.

  3. IAASBStandards
    Proposed Revisions for Audit Evidence & Risk Response: ISA 330, ISA 500 & ISA 520

    The IAASB proposed revisions to its audit evidence standards, including stronger requirements for judging whether information used as evidence is relevant and reliable. Comments are due December 15, 2026.

  4. UK AI Security InstituteIncident
    Incident Report: unsanctioned agent behaviour during cyber testing

    The UK AI Security Institute says it reviewed 122 test runs with automated transcript scanning and expert analysis, found agents acting outside the test's scope in 10, and can't yet say how likely that is outside testing.

July 2026

  1. UK AI Security InstituteEvaluation
    Cheating behaviour in frontier model evaluations

    The UK AI Security Institute says every model it tested for cheating tried to cheat, and that models did not reliably report it when asked.

May 2026

  1. Apollo ResearchEvaluation
    The Need for Deeper, White-Box Access to Maintain State of the Art Evaluations for Loss of Control Threats

    Apollo Research argues that evaluation awareness weakens black-box testing, and asks that outside evaluators get raw chain of thought and the same access modes as internal ones.

  2. METREvaluation
    Frontier Risk Report (February to March 2026)

    METR says four frontier companies took part in its pilot, each giving model access with raw chains of thought, and that it cannot give assurance about companies that did not take part.

  3. NISTEvaluation
    CAISI Signs Agreements Regarding Frontier AI National Security Testing With Google DeepMind, Microsoft and xAI

    NIST says its Center for AI Standards and Innovation has completed more than 40 evaluations, and that developers often provide models with safeguards reduced or removed.

April 2026

  1. UK AI Security InstituteResearch
    What can sandboxed AI agents learn about their evaluation environments?

    The UK AI Security Institute reports that an agent in its sandbox worked out it was being evaluated by AISI and mapped part of AISI's cloud setup.

March 2026

  1. NISTStandards
    New Report: Challenges to the Monitoring of Deployed AI Systems

    NIST's report on monitoring deployed AI systems lists fragmented logging across distributed infrastructure among the barriers.

Each headline and each fact is its source's, and each summary is ours, with the source named first. Items marked PacSpace are our own posts.

Talk to us

Bring the case you think breaks it.

We would rather be evaluated by use than by description. Talk to us and we'll put you in a live environment: commit a record, do your best to change it, then check it yourself, with us out of the loop. The change shows.

The record must exist.