Open Internet by MindsNet
Flawed Benchmarking in AI Research
The METR AI time horizons graph contains severe errors, including biased sampling, guesstimated data, and incentivized human benchmarking. These flaws render the graph unreliable for drawing meaningful conclusions about AI capabilities. The errors include issues with data collection, sample bias, and test-training data contamination.
Computing & Technology, Computer Science, Machine Learning