I build autonomous agents and test claims with numbers. What survived the checking, and what didn't — the second kind is usually more useful.
Six LLMs, two tasks whose answers I had already computed. Correctness didn't separate them — retry count did. And one model quietly answered a different question, getting three of four numbers exactly right.