AI safety benchmarks reward over-refusal, study of 192 models finds
This digest was compiled by AI from multiple sources — links to the originals are below.

A study by the UK AI Security Institute and collaborators analyzed 192 models across 5,000 questions and found that AI safety benchmarks measure three distinct traits, not one. The benchmarks reward models that refuse more requests, even when those requests are harmless. The researchers also found that fewer than 2 percent of questions meaningfully distinguish models.
Key Facts
- The study analyzed answers from 192 models across more than 5,000 test questions, making it the largest analysis of its kind to date.
- The eight safety benchmarks measure three distinct traits: refusal strictness, truthfulness, and handling of context-dependent content.
- HarmBench and OR-Bench-Hard are inversely related: a model that scores well on one almost always scores poorly on the other.
- Fewer than 2 percent of test questions meaningfully distinguish between models; nearly all models pass or fail the rest.
- Three short tests of 25 questions each can capture all three safety dimensions more accurately than a random sample of the same size.
Benchmark Structure
The eight safety benchmarks do not measure a single shared quality called 'safety'. They measure three separate traits: how strictly a model refuses requests, how truthfully it answers, and how it handles content that can be harmless or dangerous depending on context. These three traits have little correlation with each other; whether a model answers honestly says almost nothing about how often it refuses requests. HarmBench and SORRY-Bench measure almost the same thing, while OR-Bench-Hard swings in the exact opposite direction.
Over-Refusal Incentive
HarmBench rewards a model for refusing harmful requests, while OR-Bench-Hard punishes it for being overly cautious with harmless ones. A model that scores well on one benchmark will almost always score poorly on the other. This means a model can boost its overall rating simply by blocking more requests across the board, even as it becomes less useful in everyday use. Averaging results across several benchmarks papers over this tradeoff entirely and rewards behaviors that get double-counted by multiple similar tests.
Test Efficiency
Most test questions turn out to be dead weight: nearly every model passes them, or nearly every model fails them, so they do almost nothing to tell models apart. Pick the most informative questions instead, and three short tests of just 25 questions each can capture all three safety dimensions, more accurately than a random sample of the same size. About ten adaptively chosen questions produce a ranking that comes very close to matching the full benchmark.
1 source
AI safety benchmarks reward over-refusal, study of 192 models finds



