Blog

September 2026 · By James Mann

Is GLM-5.3-Flash Mythos-level at Cyber?

We ran GLM-5.3-Flash on ExploitBench with a budget of 1 billion tokens per vulnerability. It matched Claude Mythos Preview’s performance at about 6% of the cost.

Read more →

August 2026 · By Alexander Putilin & Justin Olive

A Better LLM Grader for Humanity’s Last Exam

OpenAI is deprecating o3-mini, the default LLM judge for Humanity’s Last Exam, so we compared the alternatives on cost and agreement with a frontier-model panel. Most judging is easy — but there is a long tail of questions where judging accurately requires model intelligence.

Read more

August 2026 · By Laurence Wroe

Auditing BixBench: Broken Tasks and Faulty Judging Hide Saturation

We audited BixBench, an agentic benchmark used to evaluate dual-use scientific capabilities. At first glance it is far from saturated — but flawed tasks and judge grading errors explain roughly 70% of failed submissions, and correcting for them can raise measured agent performance by up to 42 percentage points.

Read more

July 2026 · By James Mann

Forecasting the Remote Labor Index

Benchmarks can forecast AI progress, not just track it. We forecast the Remote Labor Index — a measure of how much real remote work AI can do — and stress-test the method on benchmarks that have already saturated.

Read more

June 2026 · By Jay Bailey

Why Are Evaluations Broken?

Why are so many AI evaluations broken, and how can we improve on this problem? We explore the root causes and share our approach to building better evaluations.

Read more