Reading published frontier safety frameworks alongside one another is more informative than reading any one of them. The structure is shared: capability thresholds, evaluations intended to detect proximity to them, and commitments about what happens when one is reached. The variation is in properties that are easy to miss and largely determine whether the document could constrain a decision.
One: the threshold has to be evaluable before the decision
A threshold expressed in terms only measurable after deployment cannot inform a deployment decision. This sounds obvious and is violated often, usually where a threshold is defined in terms of real-world consequence — misuse observed in the wild, harm attributable to the system — rather than in terms of measurable capability.
Consequence-based thresholds are more meaningful and less useful. The frameworks that work define capability proxies and are explicit that they are proxies.
Two: the response has to include not shipping
Most frameworks describe response in terms of additional mitigation: more safeguards, restricted deployment, staged release. Fewer state plainly that there is a finding which would result in the model not being released at all.
A framework whose entire response menu consists of things that can be added while still shipping has not described a constraint. It has described a process.
Three: revision has to be governed
Every framework will be revised; the technology moves and early thresholds will prove badly calibrated. A document with no stated revision process will be revised anyway, quietly, and probably in the direction of whatever was inconvenient.
The useful provision is a commitment that revisions are published, dated, and accompanied by the reasoning — so that a threshold moving after an evaluation result is visible as such.
Four: none of it substitutes for access
Self-assessment against self-defined thresholds using self-designed evaluations is what almost all of these frameworks describe, and most say so candidly. External evaluation with genuine access is the only thing that changes the epistemic status of the output, and the capacity to provide it is only now being built.
Until it exists at scale, a framework is a statement of intent. Statements of intent are worth having. They are not evidence, and coverage that treats them as evidence is doing the industry a favour it did not ask for.