Skip to content
Read the original: Cognition Blog (Devin, Windsurf)·PublishedPickAI score60

Cognition tests OpenAI o1 models in Devin's coding agent benchmark

A review of OpenAI’s o1 and how we evaluate coding agents

AISummary

Cognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

AIWhy it matters

The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.

Read the original cognition.com

Source: Cognition Blog (Devin, Windsurf) · cognition.com