STORY · FORSKNING_
OpenAI drops SWE-bench Verified due to contamination and leakage
OpenAI says that benchmark SWE-bench Verified has become unreliable because the tests are contaminated and training leakage has occurred. The company recommends using SWE-bench Pro instead to measure progress in AI code generation.
WHY IT MATTERS
Reliable benchmarks are critical for evaluating progress in code-writing AI. If the most widely used tests are unreliable, the industry must find alternative ways to compare models, which affects how AI companies assess their own capabilities.
SOURCES
MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.