The UK AI Security Institute is using the EvalEval Coalition’s infrastructure to publish model evaluation results and setup configurations, Hugging Face reported. The Institute, a research unit inside the UK Department for Science, Innovation and Technology, adopted the platform to make evaluation science more verifiable and reproducible.
Researchers from the Institute and the coalition first collaborated at a joint workshop alongside NeurIPS 2025. Feedback from AISI helped design the coalition's Every Eval Ever schema, establishing a shared structure for reporting benchmark runs. This work expands on existing AISI tools, including OptStop for evaluation efficiency, HiBayES for hierarchical Bayesian modelling, and standardised procedures for transcript analysis and capability elicitation.
Six frontier models—Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4—feature in the release's Terminal-Bench 2.0 dataset. The Institute also published run data from two cyber assessments, Cyber CTFs and The Last Ones, which test a partially overlapping set of systems. The shared data accompanies AISI's research paper, "How Inference Compute Shapes Frontier LLM Evaluation," which covers five core benchmarks.
Models tested on Humanity's Last Exam solved additional tasks as token use increased, provided they received oracle correctness feedback after each attempt. The published curves illustrate that benchmark performance alters depending on inference-time compute allocations and specific evaluation protocols.
The EvalEval Coalition maintains Evaluation Cards to merge benchmark metadata, evaluation-run records, and model details into unified profiles. This infrastructure enables researchers to identify when identical performance scores arise from meaningfully different experimental conditions.
