Home › AI › UK AISI Adopts EvalEval Platform to St
AI

UK AISI Adopts EvalEval Platform to Standardise AI Benchmarks

The UK AI Security Institute is using EvalEval infrastructure to publish reproducible evaluation data across frontier models, Hugging Face reported.

WHAT YOU NEED TO KNOW
  • The UK AI Security Institute adopted EvalEval's Every Eval Ever schema and Evaluation Cards platform following collaboration that began at NeurIPS 2025.
  • Terminal-Bench 2.0 evaluation data was published for Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4.
  • The release accompanies AISI's paper, "How Inference Compute Shapes Frontier LLM Evaluation," covering five primary benchmarks alongside Cyber CTFs and The Last Ones.

The UK AI Security Institute is using the EvalEval Coalition’s infrastructure to publish model evaluation results and setup configurations, Hugging Face reported. The Institute, a research unit inside the UK Department for Science, Innovation and Technology, adopted the platform to make evaluation science more verifiable and reproducible.

Researchers from the Institute and the coalition first collaborated at a joint workshop alongside NeurIPS 2025. Feedback from AISI helped design the coalition's Every Eval Ever schema, establishing a shared structure for reporting benchmark runs. This work expands on existing AISI tools, including OptStop for evaluation efficiency, HiBayES for hierarchical Bayesian modelling, and standardised procedures for transcript analysis and capability elicitation.

Six frontier models—Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4—feature in the release's Terminal-Bench 2.0 dataset. The Institute also published run data from two cyber assessments, Cyber CTFs and The Last Ones, which test a partially overlapping set of systems. The shared data accompanies AISI's research paper, "How Inference Compute Shapes Frontier LLM Evaluation," which covers five core benchmarks.

Models tested on Humanity's Last Exam solved additional tasks as token use increased, provided they received oracle correctness feedback after each attempt. The published curves illustrate that benchmark performance alters depending on inference-time compute allocations and specific evaluation protocols.

The EvalEval Coalition maintains Evaluation Cards to merge benchmark metadata, evaluation-run records, and model details into unified profiles. This infrastructure enables researchers to identify when identical performance scores arise from meaningfully different experimental conditions.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →