Blog
September 2026 · By James Mann
Is GLM-5.3-Flash Mythos-level at Cyber?
We ran GLM-5.3-Flash on ExploitBench with a budget of 1 billion tokens per vulnerability. It matched Claude Mythos Preview’s performance at about 6% of the cost.
Read more →August 2026 · By Alexander Putilin & Justin Olive
A Better LLM Grader for Humanity’s Last Exam
OpenAI is deprecating o3-mini, the default LLM judge for Humanity’s Last Exam, so we compared the alternatives on cost and agreement with a frontier-model panel. Most judging is easy — but there is a long tail of questions where judging accurately requires model intelligence.
Read moreAugust 2026 · By Laurence Wroe
Auditing BixBench: Broken Tasks and Faulty Judging Hide Saturation
We audited BixBench, an agentic benchmark used to evaluate dual-use scientific capabilities. At first glance it is far from saturated — but flawed tasks and judge grading errors explain roughly 70% of failed submissions, and correcting for them can raise measured agent performance by up to 42 percentage points.
Read moreJuly 2026 · By James Mann
Forecasting the Remote Labor Index
Benchmarks can forecast AI progress, not just track it. We forecast the Remote Labor Index — a measure of how much real remote work AI can do — and stress-test the method on benchmarks that have already saturated.
Read moreJune 2026 · By Jay Bailey
Why Are Evaluations Broken?
Why are so many AI evaluations broken, and how can we improve on this problem? We explore the root causes and share our approach to building better evaluations.
Read more