Skip to contentSkip to stories

Updated

#OpenAI

Showing low-relevance items too. Hide low-relevance items

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

Dec 1, 2025

Dec 1, 2025Mon

Nov 14, 2025

Nov 14, 2025Fri

Oct 26, 2025

Oct 26, 2025Sun
  1. Factory NewsAI score36

    AWS and Factory Announce Partnership, Factory Available on AWS Marketplace

    AIFactory has announced a partnership with Amazon Web Services and made its Droids agent platform available on the AWS Marketplace. Enterprise teams can use existing AWS Enterprise Discount Program commitments to buy Factory, with Droids accessible from CLI, Terminal UI, Web, Slack, Linear, and an IDE overlay. The source cites 31× faster feature development, 96.1%+ reduction in migration times, and 95.8% reduction in incident resolution times.

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.