STORY · FORSKNING_

OpenAI drops SWE-bench Verified due to contamination and leakage

OpenAI says that benchmark SWE-bench Verified has become unreliable because the tests are contaminated and training leakage has occurred. The company recommends using SWE-bench Pro instead to measure progress in AI code generation.

WHY IT MATTERS

Reliable benchmarks are critical for evaluating progress in code-writing AI. If the most widely used tests are unreliable, the industry must find alternative ways to compare models, which affects how AI companies assess their own capabilities.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.