The UK's AI Security Institute has released a study indicating that conventional benchmarks for evaluating artificial intelligence capabilities are fundamentally flawed. By imposing a cap on compute budgets, these evaluations fail to capture the true potential of AI agents. The research, which examined seven different benchmarks, found that when the token budget was increased tenfold, success rates for software engineering tasks surged by approximately 25 percent. This finding underscores a critical gap in how AI performance is measured, particularly as newer models exhibit even greater improvements under relaxed constraints. The study suggests that the actual advancements in AI capabilities are about 60 percent steeper than previously reported, challenging the established metrics used in the industry today.
Source: The Decoder