OpenAI Reassesses SWE-Bench Pro: About 30% of Tasks Found Broken

Friends, I’d like to share an important update from the OpenAI ecosystem.
An audit of SWE-Bench Pro found that around 30% of the benchmark’s tasks are incorrect.
The main issues include:
— overly strict tests;
— insufficiently specific prompts;
— low test coverage;
— misleading wording.
OpenAI has also withdrawn its previous recommendation to move to SWE-Bench Pro.
Why this matters: benchmark quality directly affects how we evaluate models, draw conclusions about AI capabilities, and make deployment decisions.
How do you verify that a model evaluation truly reflects real ability, rather than dataset errors?

