AI
New AI benchmarks reveal underestimated capabilities of advanced models
A UK study highlights that traditional AI evaluations may significantly undervalue the potential of AI agents, particularly in software engineering tasks, suggesting a need for revised benchmarks.