VMTech
Discuss a project

OpenAI Reassesses SWE-Bench Pro: About 30% of Tasks Found Broken

OpenAI Reassesses SWE-Bench Pro: About 30% of Tasks Found Broken

Friends, I’d like to share an important update from the OpenAI ecosystem.

An audit of SWE-Bench Pro found that around 30% of the benchmark’s tasks are incorrect.

The main issues include:
— overly strict tests;
— insufficiently specific prompts;
— low test coverage;
— misleading wording.

OpenAI has also withdrawn its previous recommendation to move to SWE-Bench Pro.

Why this matters: benchmark quality directly affects how we evaluate models, draw conclusions about AI capabilities, and make deployment decisions.

How do you verify that a model evaluation truly reflects real ability, rather than dataset errors?

Open analytics
On the site 2 views
min read 1 08.07.2026
On Instagram 3 views
On Instagram 1 reach
Instagram

OpenAI Reassesses SWE-Bench Pro: About 30% of Tasks Found Broken

Open the post on Instagram ↗