skip to content

The Evidence Boundary Between EU GPAI Code Adherence and a Model-Release Opinion


What boards, customers and independent reviewers can—and cannot—conclude from public frontier-model evidence

A board is deciding whether to approve reliance on a frontier model. The provider points to its EU General-Purpose AI Code of Practice signature, public governance framework and system card. What has been established?

Two answers must be kept separate. The public record may support findings about documented governance design and specific disclosed test results. It does not, by itself, transfer an independent opinion that the exact release has acceptable residual risk for the planned deployment.

The regulatory timing makes that distinction consequential. EU AI Act obligations for general-purpose AI providers have applied since 2 August 2025. The European Commission gained enforcement powers over those obligations, including the power to impose fines, on 2 August 2026. Providers of models placed on the market before 2 August 2025 have until 2 August 2027 to comply.

The Commission and AI Board have confirmed that the Code is an adequate voluntary route for providers to demonstrate compliance with the Act. Its Safety and Security Chapter is substantive. Signatories providing models with systemic risk commit to identify and analyse risks, set acceptance criteria, apply safety margins, decide whether to proceed, document their reasons, monitor after release and revisit the case when material conditions change.

A signature therefore matters: it records a commitment to that framework. It is not evidence that every commitment operated as intended, a third-party attestation, or an independent opinion on a particular model release.

Evidence of actual adherence would be stronger than a signature. It would still answer a different question: whether the relevant Code commitments were followed, not whether an outside reviewer has independently justified the residual-risk decision for a defined release.

That is the narrow claim here. It does not imply that a signatory is non-compliant, that a named model is unsafe, or that public evidence should reproduce the regulator’s file. It asks what a board, customer, counterparty or independent reviewer can responsibly conclude from evidence it has actually examined.

Here, a “release opinion” means a non-statutory technical judgement on whether the evidence supports releasing a defined model into a defined deployment envelope. It is not a legal opinion or an audit under a universal AI-assurance standard. Its credibility comes from competence, independence, access, challenge and a conclusion bounded to the object reviewed.

There is one regime, but two evidence surfaces.

The Code requires a Safety and Security Model Report before a covered model is placed on the market. It must explain why systemic risks are acceptable, what evidence supports that conclusion, what mitigations are in place, how external input affected the decision and what conditions would invalidate it. The AI Office normally receives the report without redactions, subject to a narrow national-security exception and a limited timing provision.

The public may see something different. A published summary of the Model Report can omit security-sensitive or commercially sensitive material, and publication is not required where a model is treated as similarly safe or safer than a reference model.

This asymmetry is often the right design. Publishing attack methods, safeguard weaknesses or model-security controls could make systems less safe. Confidential regulatory supervision may be more probative than open publication.

The error is to treat access on one surface as if it transfers a conclusion to another. An outside reviewer that has not examined the regulator-facing evidence cannot borrow the regulator’s possible conclusion. Public evidence can still establish useful facts. But a document’s title—“Code signatory”, “system card” or “safety case”—does not settle its scope, depth or assurance level.

The distinction is established; this application is current.

The compliance-versus-assurance distinction is not new. Gamut’s public method separates a compliance answer from assurance depth. Vorp Labs’ frontier-model checklist says that a framework is not a release result and calls for evidence tied to the exact model and workflow. Research on third-party compliance reviews examines who reviews, what evidence they see and what may be disclosed. A broader proposal for frontier-AI auditing grounds rigorous assessment in deep, secure access to non-public information.

The narrower contribution here is to apply that established discipline to the final EU Code after enforcement began, then test it against the current public frameworks and reports of three signatories. Those materials expose three different boundaries.

There are three providers and three boundaries.

OpenAI’s Frontier Governance Framework is the public summary of its Safety and Security Framework under the Code. It treats one-time capability elicitation as a lower bound, applies safety margins and says threshold determinations reflect a holistic judgement over the evidence. Its GPT-5.6 System Card contains extensive evaluation results and names external testing by the UK AI Security Institute. It also says its safeguards section summarises a more detailed internal report that informed the Safety Advisory Group’s recommendation and the company’s availability decision. The boundary is access: the public card reports meaningful evidence, but not the full record behind the release judgement.

Google DeepMind’s Frontier Safety Framework 3.1 acknowledges subjective analysis and balances safety with innovation through proportionality. It requires a supplemental safety case when a critical capability level is reached, while stating no corresponding requirement for tracked-capability levels. For its CBRN decision, the Gemini 3 Pro report says the full evaluation set was run on a similar earlier model, followed by limited testing of the final model and a decision not to rerun the full set. The boundary is transfer: the reasoning is disclosed, but an independent reviewer relying on it would still need to assess the proxy-to-release mapping.

Anthropic’s Responsible Scaling Policy 3.4 makes review mechanics unusually visible. It assigns the ordinary ultimate decision on Risk Report adequacy and downstream plans to the CEO and Responsible Scaling Officer, requires broader approval when marginal-risk analysis is important, and defines a minimum trigger for full external review. Its February 2026 Risk Report presents risk arguments, monitoring evidence, limitations and redaction reasons, but assesses Anthropic’s activities as a whole rather than one release. The boundary is scope and authority: the public record identifies who decides and how review escalates, but its object is not a model-specific release.

The point is not to rank the providers. It is that the same vocabulary—risk tier, external review, safety case, acceptable residual risk—can sit above different access, transfer assumptions, review triggers, scopes and decision rights. The noun does not settle the assurance question.

The strongest objection is largely right.

The strongest objection is that public reproducibility is the wrong benchmark. The Code gives the AI Office access to richer evidence, makes independent evaluation the default subject to defined exceptions, requires monitoring and incident reporting, and reopens the risk decision when material conditions change. Controlled access can protect security without abandoning scrutiny.

A customer also need not commission an independent model-release opinion merely to make a procurement decision. It must assess its own integration, controls and use case. A Code signature can be useful due-diligence evidence and can support reliance on the regulatory regime.

Those points defeat any claim that incomplete public evidence proves regulatory failure. They do not make a signature sufficient evidence for a model-specific assurance conclusion. Confidential supervision, customer due diligence and independent technical assessment answer different questions. They can complement one another; none should be relabelled as another.

Where an independent conclusion is needed, the practical answer is scoped access rather than universal disclosure. Sensitive evidence can be reviewed in controlled environments. Qualified reviewers can inspect segregated sections. A public conclusion can omit exploit details while stating what was withheld, who reviewed it and how the limitation narrows the conclusion.

Match the conclusion to the evidence.

The opinion should climb no higher than the evidence surface permits.

At the signature-and-framework layer, report the commitment and the documented design: what the public framework says about responsibilities, required artefacts, change control, monitoring and escalation. Do not convert stated design into operating effectiveness.

At the model-disclosure layer, report only the model-specific findings the public evidence supports. Name the tested model and safeguards, the deployment conditions, important negative results, redactions, uncertainty and any transfer from a proxy or earlier version. Do not turn the existence of a system card into an overall release endorsement.

At the independent-opinion layer, the reviewer needs enough scoped access to bind and challenge the decision record. The opinion should identify the exact model and deployment envelope, risk-acceptance criteria, material transfer assumptions, mitigation evidence, decision authority, withheld evidence, conditions of validity and events that reopen the decision. A critical contradiction, materially uncertain evaluation or release outside that envelope stops a positive opinion.

This structure also keeps customer approval in its proper place. A deployer may decide that regulatory reliance, provider evidence and its own application controls are sufficient for a use case without claiming to have independently assured the provider’s model-release judgement.

The EU Code gives frontier-AI governance a stronger compliance spine. Its value should not be understated. Nor should it be asked to carry a conclusion it was never designed to transfer.

The primary and adjacent sources are Regulation (EU) 2024/1689, Articles 55 and 56; the EU General-Purpose AI Code of Practice, Safety and Security Chapter; the OpenAI Frontier Governance Framework and GPT-5.6 System Card; the Google DeepMind Frontier Safety Framework 3.1 and Gemini 3 Pro report; the Anthropic Responsible Scaling Policy 3.4 and Risk Report: February 2026; Gamut’s scoring, assurance and conclusions and evidence, testing and findings; the Vorp Labs frontier model vendor review checklist; Homewood et al. on third-party compliance reviews for frontier AI safety frameworks; and Brundage et al. on frontier AI auditing.

Related by topic
  1. The Judgement Bucket: How to Be Rigorous About the Claims You Can't Measure
  2. The Calibration Gap: Why AI Assurance Needs Experimental Rigour at Consulting Speed
  3. Near-Parity Is Mostly a Claim About Formatting