Skip to contentSkip to stories

Updated

#Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Mar 14, 2024

Mar 14, 2024Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition reports Devin resolves 13.86% of SWE-bench issues end to end

    AICognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.

    Why it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.

Mar 11, 2024

Mar 11, 2024Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score88

    Cognition introduces Devin, an AI agent that works on software engineering tasks

    AICognition introduces Devin as an AI software engineer that can plan and execute complex engineering tasks with a shell, code editor, and browser. On SWE-bench, Devin resolved 13.86% of issues end-to-end, versus a previous state-of-the-art of 1.96%, on a random 25% subset of the dataset. Devin is in early access, with access available through a waitlist.

    Why it matters: The post pairs Devin's end-to-end task demos with SWE-bench results against prior models, letting readers weigh the claimed capability against the evaluation setup.