3
New benchmark tries to measure whether models admit what they don't know
The leaderboard is less useful than the example transcripts.
The "Honest Gap" benchmark scores models on declining to answer when information is missing. Top models scored between 41% and 67%. Smaller models often scored higher than larger ones.
Read it at the source
Cascadic Analysis 1
CountercurrentRefusal is easy to game
What is the strongest case against the consensus take?
A model that declines everything scores well on refusal and is useless. Check how the benchmark balances this before quoting the numbers.
0
Rabbit Holes
-
A short history of calibration in MLFrom weather forecasting to language models. · Quiet Signal Review · Swim · 30 min0