Logged AI model safety evaluations
Model Evaluated by Risk Date Findings
Agentic systems — authorisation and tool-use boundaries Multiple platform vendors Enterprise security research groups High risk Jul 2026

Security research on deployed agentic systems has converged on a consistent set of weaknesses: over-broad delegated permissions, absent agent inventories, insufficient egress constraint, and audit trails that record actions without the reasoning that produced them.

Recorded here because the findings are reproducible across vendors and deployments rather than specific to one product.

Full report →

Frontier model pre-deployment evaluations Multiple frontier developers UK AI Security Institute Medium risk Jun 2026

The Institute has published a series of pre-deployment evaluations conducted under voluntary access arrangements with frontier developers, covering cyber capability, chemical and biological uplift, autonomy and safeguard robustness.

Read the Institute's own publications for findings. This log records that the evaluations exist and where they are published; it does not restate their conclusions, which are stated with qualifications this summary cannot carry.

Full report →

Open-weight releases — safeguard durability Multiple open-weight developers Independent academic red-teaming groups High risk Apr 2026

A body of published academic work has examined how durable safety training is in models whose weights are released, consistently finding that refusal behaviour can be substantially reduced through fine-tuning at low cost relative to the cost of training the model.

The governance implication recorded here is narrow and well supported: for an open-weight release, the capability being distributed is the capability of the model with its safeguards removed. Individual papers should be consulted for method and magnitude.

Full report →

General-purpose AI — capability and risk synthesis Not model-specific International AI Safety Report expert panel Medium risk Jan 2026

An international, expert-led synthesis of the evidence on general-purpose AI capabilities and risks, produced with the participation of a large number of countries and intergovernmental organisations.

It is a synthesis of existing research rather than an original evaluation, which makes it the appropriate starting point for a reader wanting the state of the evidence and the wrong source for a specific claim about a specific model.

Full report →