Certified Where? The Missing Layer in AI Conformity Assessment

Analysis ·

A developer builds a general-purpose model and wants to deploy it in three markets. Each market has a defined process for establishing that the model may be deployed. The developer completes the first process, receives whatever documentation it produces, and presents it in the second market.

It counts for nothing. Not because it was rejected — because there is no procedure by which it could be considered.

This is the current state of AI conformity assessment. Four distinct regimes now specify what must be demonstrated about a model before it is placed on a market or used in a regulated context. Between them there is no mutual recognition mechanism of any kind.

What each regime requires

The requirements are not variations on a theme. They ask about different properties, using different evidence, assessed by different parties.

The European Union operates under Regulation 2024/1689. High-risk systems require a risk management system, data governance covering training and validation sets, technical documentation, record-keeping sufficient for traceability, transparency toward deployers, and human oversight provisions. General-purpose model obligations took effect on 2 August 2025; the transparency requirements of Article 50 apply from 2 August 2026.

Assessment is meant to run through harmonised standards, which confer a presumption of conformity. These are being developed by CEN-CENELEC JTC 21 — risk management, dataset governance, quality management, bias, cybersecurity, trustworthiness. As of mid-2026 none had been published in the Official Journal. The quality management standard reached formal vote first. Until publication, no harmonised standard confers the presumption, while the high-risk obligations themselves arrive on schedule.

The United States has no comparable pre-market process. Executive Order 14179 removed the prior framework in January 2025; the AI Action Plan followed in July 2025. The Center for AI Standards and Innovation, at NIST, evaluates frontier models for national security risk. The 2026 cybersecurity executive order is voluntary, with agencies finalising the framework by 1 August 2026 and covered-model definitions set through classified benchmarking. The NIST AI Risk Management Framework remains voluntary.

China operates a two-tier system predating both. Under the Interim Measures for Generative AI Services, in force since August 2023, services with public opinion or social mobilisation properties require both a security assessment and an algorithm filing. As of late 2025, 748 generative services had completed the national filing. Standards work runs through TC260, including a content-labelling standard effective September 2025.

Russia introduced its first framework law, 243-FZ, in July 2026, taking effect in stages from September 2026 with the substantive provisions from March 2027. It applies to foundation models above one billion parameters and creates two statuses — sovereign and national — with the first requiring domestic control of the full lifecycle, technical reproducibility, and domestic data storage. The verification procedure itself is deferred to implementing acts not yet issued.

Against these sits the Council of Europe Framework Convention, in force since 1 November 2025 and ratified by the European Union in May 2026. It establishes principles across the AI lifecycle. It is not self-executing and specifies no verification technique.

Four processes, no bridge

The relevant point is not that the requirements differ. Requirements differ across jurisdictions in every regulated sector — pharmaceuticals, aviation, medical devices, electrical equipment. What those sectors have that AI does not is machinery for handling the difference.

That machinery has a standard shape: an agreed testing methodology, accredited bodies operating to a common competence standard, and an agreement under which a test performed by one party is accepted by another. It took decades to build in each sector, and it is what allows a device tested once to be sold in many places.

For AI models, none of the three components is in place.

The closest institutional approximation is the International Network of AI Safety Institutes, launched in November 2024, whose members include the European Commission, the United Kingdom, the United States, Japan, Korea, Canada, France, Australia, Kenya and Singapore. It has run joint testing exercises — on a large open-weight model in 2024, in Paris in 2025, and on agentic systems subsequently.

Its own mission statement is explicit that members retain the flexibility to conduct and adapt evaluations independently. Joint exercises produce shared understanding. They do not produce a result that one member is obliged to accept from another. Analysts at CSIS have argued the necessary first step is a common evidentiary methodology — the observation implies the obvious: it does not yet exist.

What does travel across jurisdictions

Some technical instruments are jurisdiction-neutral, and it is worth being precise about what they establish.

ISO/IEC 42001specifies an AI management system, and ISO/IEC 42006, published in 2025, sets requirements for bodies certifying against it. Together these enable third-party certification recognised on the strength of the standard rather than any single regulator. This is genuine interoperability infrastructure — but it certifies an organisation's management processes, not the properties of a specific model.

C2PA Content Credentials provides cryptographic provenance for media assets. The conformance programme and public trust list launched in 2025; adoption spans camera manufacturers, major model providers and platforms. It is referenced in the EU transparency requirements and in Californian legislation. Its limits are equally documented: metadata is stripped by re-encoding and screenshots, and a signature establishes which device or software produced an asset — not whether the asset depicts anything real.

Model cards and system cardsare the most widely adopted instruments and the weakest in evidentiary terms. They are assertions by a developer about a developer's own system, in a format that varies between developers.

The pattern across all three is consistent. What travels between jurisdictions today certifies organisations, or documents assets, or describes intentions. Nothing establishes a verifiable property of a specific model to a party that does not operate it.

Why this compounds rather than resolves

Two dynamics run against convergence.

The institutional environment is consolidating into distinct groupings. The United States has assembled partners around the Pax Silica framework, launched in December 2025 and expanded at a second summit in June 2026, with roughly two dozen signatories to the declaration and a wider set to an associated statement of intent. In July 2026, twenty-nine countries signed an agreement in Shanghai establishing the World Artificial Intelligence Cooperation Organization, headquartered there, with no G7 or EU member among the signatories. Reuters reported in August 2026 on a draft State Department letter framing participation in the two as incompatible; the State Department declined to comment on what it described as purportedly leaked internal documents.

Whatever one concludes about the merits, the practical consequence for conformity assessment is unambiguous: an assessment methodology developed within one grouping becomes less likely, not more, to be adopted in the other.

The second dynamic is that the technical requirements are arriving faster than the technical means. The EU transparency obligations apply from August 2026 while the harmonised standards that would establish how to satisfy them remain unpublished. The Russian framework defines a sovereign model status while deferring its verification procedure. Requirements are specified in terms of properties — traceability, reproducibility, oversight — for which no standardised demonstration method exists in any regime.

The gap stated precisely

Conformity assessment in mature sectors rests on a specific idea: that a property of an artefact can be established by a competent third party, recorded in a form others can rely on, and re-verified independently if disputed.

For AI models, the first part is contested, the second is non-standardised, and the third is generally impossible without access held only by the operator.

The instruments that do cross borders — management system certification, content provenance credentials — work because each attaches to something inspectable. A management system has documents and processes. A media asset has a signature bound to bytes.

A model's behaviour has neither. It is observed through outputs, and outputs are produced by a process that leaves no independently verifiable record of how it arrived at them. This connects directly to the measurement problem set out in the first analysis in this series: behaviour varies by conditions that are not disclosed, and no available instrument produces evidence about a specific output.

Standards work in this field is currently occupied with specifying what should be true of AI systems. The harder and less-addressed question is what would make any such claim independently checkable — because until that exists, four regimes will keep producing assessments that cannot travel, and the developer in the opening paragraph will keep completing the same work three times.


Sources: Regulation (EU) 2024/1689 · CEN-CENELEC JTC 21 work programme · Executive Order 14179; US AI Action Plan, July 2025 · Interim Measures for the Management of Generative AI Services (CAC, 2023) · Federal Law 243-FZ of 26 July 2026 · Council of Europe CETS 225 · International Network of AI Safety Institutes, mission statement and joint testing exercises · ISO/IEC 42001:2023, ISO/IEC 42006:2025 · C2PA specification and conformance programme · US Department of State, Pax Silica summit outcomes, June 2026 · Xinhua, 16 July 2026 · Reuters, 14 August 2026 · CSIS analysis of the AI Safety Institute network