What the White House Accord Means for AI Assurance

On September 29, the White House released an accord on super intelligence, signed by Google, Anthropic, Meta, OpenAI, xAI, and NVIDIA. The commitments are voluntary, but they put external assessment of frontier AI developers and board-level oversight in writing.

For anyone who works in assurance, that is a significant development in the AI industry’s commitment to strengthening security measures. The accord sets out what should happen, but it leaves open who is qualified to do it, what they assess against, and what independence requires. Those answers will determine whether the accord leads to real assurance or only the appearance of it.

The four layers

The accord describes four layers of accountability: 

  1. Management owns the controls 
  2. An internal function verifies that they are working 
  3. An independent external party assesses them 
  4. A board committee oversees the process 

The structure will look familiar to anyone who has worked in financial or security assurance. It works in other industries because each layer has defined expectations behind it. For frontier AI, the structure is now on paper. The harder work is defining what each layer actually has to demonstrate. 

“Audit” covers more than one activity

Different activities are likely to get blended under the word “audit.” The accord speaks directly to the first two: 

  • Model evaluation tests whether a model has dangerous cyber, biological, or chemical capabilities.
  • Controls assurance checks whether a company’s controls are designed well and working in practice. This is the territory of a management system audit.
  • Impact assessment looks at effects on people and society. It is a separate discipline.

A mature AI assurance regime will need to keep all three distinct, because each calls for different methods and different practitioner competence. 

Who counts as an independent assessor?

Layer three carries the most weight and the least definition. The accord calls for an “independent external auditor or evaluator,” but it does not say who qualifies, what competence they need, what criteria they assess against, or what independence means when an evaluator also sells testing services to the company it assesses.

Until those questions are answered, two companies can both claim independent assessment and mean very different things by it. Independent assessment is only as credible as the rules for who counts as independent.

That gap matters more because the commitments may not stay voluntary. The signatories have now put external assessment and board oversight in writing, and the accord says these steps may eventually be written into law. If that happens, the rules for who can serve as an independent assessor need to be settled before the mandate arrives.

A requirement with no qualified assessor market behind it produces nothing more than checkbox assurance. No consequences and no accountability mean no real assurance.

What “operating as intended” requires, and what boards need to see

The phrase “operating as intended” appears throughout the accord and raises what assurance practitioners call operating effectiveness. Showing that a control exists or was designed correctly is only the beginning. The harder question is whether the control worked when it was supposed to, and whether there is evidence it kept working as the system changed. Frontier models change between training runs and deployments, so that question never stays answered for long.

That is where the fourth layer comes in. A board committee only adds oversight if its members can use what they receive. Directors need reporting that shows whether controls worked, where they failed, and what was fixed, in language they can understand and act on.

Building on existing infrastructure

AI assurance does not have to be invented from scratch. ISO 42001 specifies requirements for an AI management system. ISO 42006 adds requirements for the bodies that audit and certify against ISO 42001, supplementing ISO 17021-1. Backed by accreditation, the two standards already provide a working model for organizational controls, auditor competence, impartiality, and oversight of the auditors themselves.

They are not frontier model evaluation standards, and should not be treated as such. The assurance architecture has to connect management system assurance with qualified technical evaluation, rather than treating either one as the whole answer. Newer frameworks are beginning to add the technical testing side. AIUC-1, for example, certifies specific AI agents in specific deployments through third-party testing, and it is designed to complement ISO 42001 rather than replace it.

PACT AI was formed to close the gap at the frontier, and A-LIGN is one of its founding member organizations. Through PACT AI, A-LIGN is working with enterprises and technical assurance providers to help formalize the standards of practice, independence rules, and competence criteria that a credible independent verification organization (IVO) model needs.

Connecticut has already enacted an IVO pilot, and the White House accord makes the case for that same kind of infrastructure at the frontier. Mature assurance markets already know how to establish auditor competence, independence, and oversight, and AI should borrow that infrastructure.

The takeaway

The accord builds the structure for AI assurance. The next job is defining who is qualified to fill it. That means settling the rules for independence, competence, and oversight of the assessors themselves, before any mandate arrives. Much of the infrastructure to do that already exists, and the work now is connecting it to frontier AI.

Organizations preparing for stronger AI oversight can start by building an AI management system. Reach out to learn how A-LIGN supports ISO 42001 readiness and certification.