Mindbytes Dec 11, 2025
Stanford Research Exposes Hidden Crisis in AI Benchmarking
Stanford researchers find 5% of AI benchmark questions are flawed, distorting model rankings and investment. Their new method detects these "fantastic bugs" with 84% precision.
linsey 1 min read