Skip to content
View original post on X: wh· 38/100AI score38/100

Epoch finds 23 false negatives in DeepSWE, capping its score near 79.6%

AISummary

Epoch AI found 23 false negatives among 131 DeepSWE tasks, implying a ceiling of roughly 79.6%. This suggests the current near-saturated score of about 74% is constrained by artificial benchmark limits rather than model capability.

Post on XView on X
@nrehiew_

Super cool work! This also explains why some benches have an "artificial ceiling".

For example, on DeepSWE, Epoch found 23 false negatives out of 131 tasks. This means a ceiling of roughly 79.6%, which actually matches the current (almost) saturated high score of ~74%

Epoch AI@EpochAIResearch
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
View quoted post on X

Source: wh · x.comPublished · added here