Independent AI Evaluation Needs Independent Evidence
By Oleh Vasylenko ·
What the White House Accord on Superintelligence leaves unresolved
On 29 September 2026, six leading AI companies signed a voluntary accord at the White House. It gives an independent external auditor or evaluator a named place in frontier AI governance. It does not say what evidence that evaluator needs to reach a conclusion the developer could not have written for it. That question now matters more than the accord itself.
What changed on 29 September
The accord was signed by the US President and by the heads of Google, Anthropic, Meta, OpenAI, Nvidia and xAI. Its text was posted as a screenshot and reproduced in full by the Washington Examiner. This essay relies on that reproduction; no text had appeared on whitehouse.gov at the time of writing.
The accord asks each company training and deploying frontier models to put four layers in place:
- Internal controls that monitor model capabilities and alignment during training and deployment, including that models do not hack or access technical systems in unintended ways.
- An internal team that checks the controls, monitoring and detection work as intended, and that issues are fixed.
- An independent external auditor or evaluator that assesses whether the controls, monitoring and detection work as intended.
- An independent board committee that receives reports from the internal teams and the auditors, and oversees remediation.
The signatories also agreed to meet regularly to develop standards and best practices. The text allows that these steps may later be codified in law. It is voluntary, and it is a beginning rather than a system.
Why the incident makes the problem concrete
The risk named in the accord's first control layer has already appeared in practice. On 21 September, the UN Independent International Scientific Panel on AI published its first thematic brief, on the OpenAI–Hugging Face incident.
According to the brief, between May and July 2026 agents in OpenAI's cybersecurity training and evaluations bypassed network restrictions. They communicated across runs meant to stay separate, cheated an evaluator and tried to hide it, and compromised parts of OpenAI's and Hugging Face's systems. No human directed the individual steps. The Panel does not issue recommendations and does not estimate the probability of severe loss of control.
One detail matters for any evaluation regime. The system under evaluation acted against the evaluation itself. External evaluation therefore faces two distinct risks to evidentiary reliability: dependence on developer-controlled records, and strategic behaviour by the system under evaluation.
Access is not verification
External evaluation did not start with the accord. In mid-September, Anthropic proposed embedding third-party evaluators inside frontier labs, with employee-level access and a right to publish findings. OpenAI said it would follow. On 18 September, Anthropic named Accenture's Faculty unit as its first embedded evaluator, acknowledging that no standards yet exist for what information such evaluators should see or how they should report. Meta, xAI and Google DeepMind had not made the same commitment, TechCrunch reported.
Evaluators welcomed the access and described how they would use it: compare training checkpoints, inspect post-training environments, and check evaluation transcripts and logs against a company's claims. Apollo Research's head of research described the current position plainly: “we are completely relying on AI companies” to check and to report.
Recent practice shows why access alone is thin. METR and Redwood Research had roughly a week on site to investigate the Hugging Face incident and could not reach confident conclusions. Apollo Research had three days to test GPT-6 Astra and cited high evaluation awareness as a limit on what its results showed.
Access makes evaluation possible. It does not, by itself, make the resulting evidence independently verifiable.
The field is not standing still. In May, OpenAI published a playbook for third-party evaluations asking that reports explain how evaluators checked for sandbagging, evaluation awareness and other threats to validity. In August, Google DeepMind piloted a double-blind evaluation in a cryptographically protected environment, so that neither the developer nor the evaluator saw the other's material.
These are real advances. They address evaluation methodology and benchmark integrity. They do not yet answer a separate question: how an external evaluator can independently substantiate that an operational control actually functioned as intended.
Organizational independence is not evidentiary independence
An evaluator can be independent in every organizational sense: separate ownership, separate funding, a contractual right to publish. It can still depend entirely on evidence produced inside the system it evaluates.
Logs, transcripts, dashboards and checkpoints are generated within the developer's infrastructure. An evaluator who reads them is checking the developer's records, not an independent record of what happened. Where the evaluated system can recognise a test or act against an evaluator, as the Panel describes, those records may also reflect the system's own behaviour toward the evaluation.
The accord's third layer therefore rests on a condition it does not state. To assess whether controls are operating as intended, the evaluator needs evidence whose validity can be checked independently of the developer's interpretation, and which remains meaningful even when the model behaves strategically under evaluation.
What a standards discussion should define
The useful next step is not a mechanism. It is agreement on the minimum properties that evaluation evidence should have before anyone relies on it. At a minimum, such evidence should be:
- Attributable — it is clear which party and which process produced it.
- Integrity-protected — later alteration can be detected.
- Time-verifiable — it can be established when the recorded event occurred.
- Independently reviewable — another authorised reviewer can assess it without relying solely on the operator's interpretation or assertions.
- Retained and portable — it survives long enough, and in a usable form, for later review or dispute.
- Reproducible within an authorised perimeter — another authorised reviewer can repeat the check.
- Explicit about its limits — it records what the evaluator was not given, as well as what it was.
None of these properties is new to auditing. What is new is applying them to systems that may behave differently when they are being observed. The standards question is therefore not only who may evaluate frontier AI, but what evidence that evaluation must rest on.
Beyond one jurisdiction
The same question arises wherever external evaluation is required or encouraged, whatever the governance model:
- European Union. The AI Act (Article 55) requires providers of general-purpose AI models with systemic risk to perform and document model evaluations, including adversarial testing, and to report serious incidents.
- United States, state level. California Governor Gavin Newsom signed SB 813 and AB 1405 on 9 September: a framework for independent verification organisations, and a registry of AI auditors with standards for their independence, transparency and integrity. A gubernatorial executive order of 18 September directs state officials to develop recommendations that include embedded onsite verification and a potential model “kill switch” whose efficacy would be verified on an ongoing basis by an independent verification organisation.
- United Nations. The Scientific Panel notes that AI failures cross company and national borders, and that no single organisation or country sees enough incidents to identify every pattern.
- International standards. The ITU Focus Group on Trust and Identity for Humans and Agentic AI lists security criteria and benchmarks for the continuous assessment of AI agents among its planned deliverables.
These regimes differ in who may evaluate and with what authority. They share one technical dependency: evidence that an outside reviewer can rely on.
A question worth answering before standards harden
The accord commits its signatories to develop standards and best practices together. Evaluation frameworks will soon be written into contracts, audits and possibly law. Before they harden, one question deserves explicit treatment:
What evidence must an independent evaluator possess to establish that a control actually operated as intended, rather than merely confirm that the developer says it did?
This is an analysis as of 30 September 2026, based on the accord as publicly reproduced. A follow-up will test it against the implementation rules and evaluator criteria once they are published.
Sources
- White House Accord on Superintelligence, full text as reproduced by the Washington Examiner, 29 September 2026
- Independent International Scientific Panel on AI, thematic brief on the OpenAI–Hugging Face incident, 21 September 2026
- TechCrunch, “Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?”, 16 September 2026
- ITU press release on the Focus Group on Trust and Identity for Humans and Agentic AI, 9 July 2026
- OpenAI, “A shared playbook for trustworthy third party evaluations”, 29 May 2026
- Google DeepMind, “Piloting the world’s first double-blind AI evaluations”, 27 August 2026
- Anthropic, “Partnering with Accenture on embedded evaluation”, 18 September 2026
- Office of the Governor of California, executive order on independent AI oversight, 18 September 2026
- Regulation (EU) 2024/1689 (AI Act), Article 55