Special Reports Analysis Medium risk Global

What a Frontier Safety Framework Should Contain to Be Worth Anything

Published frameworks now share a structure. Reading several side by side, the differences that predict whether one constrains anything are narrower than the documents suggest.

Executive summary

Frontier safety frameworks converge on thresholds, evaluations and response commitments. Comparing published examples, four properties separate a document that could constrain a decision from one that could not — and most published frameworks have two of them.

Editorial note. This piece was written to give the section structure before launch. The subject analysis stands, but the specific development in the headline has not yet been verified against the primary document by this desk — the source is linked at the foot of the article. An editor should confirm it and rewrite the framing before this runs as reporting.

Reading published frontier safety frameworks alongside one another is more informative than reading any one of them. The structure is shared: capability thresholds, evaluations intended to detect proximity to them, and commitments about what happens when one is reached. The variation is in properties that are easy to miss and largely determine whether the document could constrain a decision.

One: the threshold has to be evaluable before the decision

A threshold expressed in terms only measurable after deployment cannot inform a deployment decision. This sounds obvious and is violated often, usually where a threshold is defined in terms of real-world consequence — misuse observed in the wild, harm attributable to the system — rather than in terms of measurable capability.

Consequence-based thresholds are more meaningful and less useful. The frameworks that work define capability proxies and are explicit that they are proxies.

Two: the response has to include not shipping

Most frameworks describe response in terms of additional mitigation: more safeguards, restricted deployment, staged release. Fewer state plainly that there is a finding which would result in the model not being released at all.

A framework whose entire response menu consists of things that can be added while still shipping has not described a constraint. It has described a process.

Three: revision has to be governed

Every framework will be revised; the technology moves and early thresholds will prove badly calibrated. A document with no stated revision process will be revised anyway, quietly, and probably in the direction of whatever was inconvenient.

The useful provision is a commitment that revisions are published, dated, and accompanied by the reasoning — so that a threshold moving after an evaluation result is visible as such.

Four: none of it substitutes for access

Self-assessment against self-defined thresholds using self-designed evaluations is what almost all of these frameworks describe, and most say so candidly. External evaluation with genuine access is the only thing that changes the epistemic status of the output, and the capacity to provide it is only now being built.

Until it exists at scale, a framework is a statement of intent. Statements of intent are worth having. They are not evidence, and coverage that treats them as evidence is doing the industry a favour it did not ask for.

References

  1. Bengio, Y. et al. (2025). International AI Safety Report. https://www.gov.uk/government/publications/international-ai-safety-report-2025
  2. UK AI Security Institute. Published evaluations and research. https://www.aisi.gov.uk/

Cite this

Administrator (2026, June 28). What a Frontier Safety Framework Should Contain to Be Worth Anything. AI News Report. https://www.ainewsreport.org.njangi.app/blog/what-a-frontier-safety-framework-should-contain