BenchMIRT: What are LLM benchmarks actually measuring?
Researchers at Allen Institute introduced BenchMIRT, a new tool that applies multidimensional item response theory to dissect what individual prompts in LLM benchmarks actually measure, revealing hidden safety and reasoning signals. By analyzing results from 100 models across 16 benchmarks, BenchMIRT independently identified safety and general reasoning as the dominant dimensions, exposing misalignments such as the BBQ bias benchmark correlating more with reasoning than safety.