AI
OpenAI uncovers flaws in key AI coding benchmark, impacting developer assessments
OpenAI's review reveals that 30% of the widely used SWE-Bench Pro coding test is ineffective, prompting a retraction of its endorsement. This raises questions about the reliability of AI programming assessments.