Updated
All AI news
Updated
Oct 22, 2025
Chip HuyenAI score22
Oct 14, 2025
Jun 11, 2025
Cognition Blog (Devin, Windsurf)PickAI score62 Cognition argues multi-agent architectures are fragile and proposes context-sharing principles
AICognition argues that parallel multi-agent architectures are fragile because subagents act on conflicting, unshared assumptions. It proposes two principles for reliable agents: share context and full agent traces, and treat actions as carrying implicit decisions. The post recommends simpler single-threaded designs for most cases and notes that context compression and fine-tuned models can extend long-running tasks.
Why it matters: The post explains concrete failure modes of parallel multi-agent setups and offers two context-sharing principles, useful for anyone designing long-running agent systems.
Sep 11, 2024
Cognition Blog (Devin, Windsurf)PickAI score60 Cognition tests OpenAI o1 models in Devin's coding agent benchmark
AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.
Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.