Lucas Beyer praises new coding benchmark for finding bugs in repos
Original titleThis looks like a cool new coding bench! Basically give the agent a repo at one commit in the past, say "find and fix all bugs" and test ...
AISummary
Lucas Beyer calls SWE-sweep a useful new benchmark, where agents must find and fix bugs in a repo checked out at an earlier commit, scored against unit tests from real later bugfixes.
He notes two limitations: a model may find valid bugs that don't match the tested ones, and the construction makes training on the test set easy. He advises not overemphasizing small ranking differences once models score highly.
Source: Lucas Beyer · x.comPublished · added here