Skip to contentSkip to stories

Updated

#Coding

Jan 18

Jan 18Sun
  1. Hamel HusainAI score40

    Why I Stopped Using nbdev for AI-Assisted Coding

    AIHamel Husain says he stopped using nbdev, a literate programming environment he helped build and maintain, because AI coding tools struggle with its notebook-to-library workflow. He now uses Amp, Cursor, and Claude Code, and reserves notebooks for data analysis, machine learning, and exploratory work. He also favors conventional stacks such as Next.js for web development, arguing that AI performs best on widely used languages with abundant training data.

Jan 13

Jan 13Tue
  1. Tim DettmersAI score36

    Tim Dettmers Argues Agents Should Automate Most Personal Work, Not Just Code

    AITim Dettmers, a professor who has used Claude Code for eight months to automate his own work, argues that more than 90% of code and text should be written by agents. He says the coding-focused hype on Twitter overstates parallel sessions and autonomy, which translate poorly to most non-software tasks. The post offers a balanced guide to what actually works in agent-based automation.

Dec 20, 2025

Dec 20, 2025Sat
  1. MiniMax · new models on Hugging FaceAI score74

    MiniMax-M2.1 open-sources weights for coding and agent tasks

    AIMiniMax has released MiniMax-M2.1 model weights on Hugging Face, with API access on the MiniMax Open Platform and the MiniMax Agent product. The company reports gains over M2 on coding and agent benchmarks such as SWE-bench Verified (74.0) and VIBE average (88.6), and says it outperforms Claude Sonnet 4.5 on multilingual scenarios.

    Why it matters: The release pairs open weights with a broad benchmark table against Claude and GPT models, letting readers compare coding and agent claims directly.

Dec 19, 2025

Dec 19, 2025Fri
  1. Andrej KarpathyAI score75

    Karpathy's 2025 LLM review names RLVR and jagged intelligence as key shifts

    AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.

Dec 16, 2025

Dec 16, 2025Tue
  1. Xiaomi MiMoAI score78

    Xiaomi releases open-source MiMo-V2-Flash MoE model for reasoning and coding

    AIXiaomi released and open-sourced MiMo-V2-Flash, a Mixture-of-Experts model with 309B total and 15B active parameters, under the MIT license. The company reports 73.4% on SWE-Bench Verified, the top score among open-source models, and inference at 150 tokens per second for $0.1 per million input tokens and $0.3 per million output tokens. It supports a hybrid thinking mode and a 256k context window.

    Why it matters: The post gives architecture, speculative decoding speedup, and pricing figures, which help readers judge how the efficiency claims are achieved and what they cost.

Dec 11, 2025

Dec 11, 2025Thu
  1. Nick TurleyAI score78

    OpenAI introduces GPT-5.2 in ChatGPT for professional work

    AIOpenAI is introducing GPT-5.2 in ChatGPT, describing it as its most advanced model series for professional work. GPT-5.2 Thinking is positioned for tasks such as building spreadsheets and presentations, writing and reviewing production code, and analyzing long documents. The post says it beats or ties industry professionals on well-specified knowledge work tasks spanning 44 occupations 70.9% of the time on GDPval, and GPT-5.2 Instant, Thinking, and Pro begin rolling out to all tiers, starting with paid plans.

    Why it matters: The post links the model's professional-work focus to GDPval results across 44 occupations, showing how the claimed capability was measured.

Dec 10, 2025

Dec 10, 2025Wed
  1. Andrej KarpathyAI score34

    Karpathy Uses GPT-5.1 Thinking to Grade December 2015 Hacker News Discussions in Hindsight

    AIAndrej Karpathy built hn-time-capsule, a tool that feeds each December 2015 Hacker News front-page article and its comment thread to GPT-5.1 Thinking for a retrospective analysis. The project, written with Claude Opus 4.5 in about three hours, processes 930 articles at a cost of about $58 and roughly one hour. Results include prescience and wrongness grades for commenters, and the project is hosted on his website with the intermediate data available for download.

Dec 4, 2025

Dec 4, 2025Thu

Nov 25, 2025

Nov 25, 2025Tue

Nov 13, 2025

Nov 13, 2025Thu
  1. Cognition Blog (Devin, Windsurf)AI score65

    Cognition's Devin review says it excels at scoped junior-level engineering work

    AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.

    Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.

Nov 3, 2025

Nov 3, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score47

    Windsurf Codemaps Adds AI-Annotated Code Maps to Help Engineers Understand Code

    AIWindsurf has launched Codemaps, AI-annotated structured maps of a codebase powered by SWE-1.5 and Claude Sonnet 4.5, which users can generate from a task prompt using a Fast (SWE-1.5) or Smart (Sonnet 4.5) model. Codemaps links grouped code sections to exact lines and can be referenced in Cascade with @{codemap} to give agents more specific context.

Oct 30, 2025

Oct 30, 2025Thu
  1. Chip HuyenAI score27

    Chip Huyen's AI product lessons: UX, data, and team structure matter most

    AIChip Huyen argues that many AI product failures stem from user experience, data quality, and organizational structure rather than the AI itself. She cites a chatbot whose traction improved after adding pre-populated questions and a voice option for users whose hands were busy, and a lead scoring model that was broken because marketing wasn't asking the right questions. She also notes that senior engineers gain the most from AI coding while resisting it more, and recommends building small tools for daily frustrations to solve the "idea crisis."

Oct 28, 2025

Oct 28, 2025Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition releases SWE-1.5, a coding agent model served at up to 950 tok/s

    AICognition has released SWE-1.5, a model optimized for software engineering that it says reaches near-frontier coding performance while running at up to 950 tok/s with Cerebras inference. The company reports it is 6x faster than Haiku 4.5 and 13x faster than Sonnet 4.5, and it is available now in Windsurf. The post's SWE-Bench Pro chart places SWE-1.5 at 40.08%, behind Sonnet 4.5 at 43.60%, and it notes that the model was trained with reinforcement learning on the Cascade agent harness.

    Why it matters: The post pairs a benchmark chart with a 950 tok/s speed claim and describes how harness, RL environments, and inference were co-designed, useful context for judging the speed-versus-quality tradeoff.

Oct 27, 2025

Oct 27, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score36

    Devin Automates .NET Framework to .NET Core Migration in Weeks, Not Months

    AICognition says its autonomous coding agent Devin can complete a .NET Framework to .NET Core migration in as little as two weeks, using a Strangler Fig approach adapted from Jimmy Bogard's guide. The post says Devin handles planning via Ask Devin and DeepWiki, dependency sharing, controller and view conversion, and session state adaptation through a remote app.

Oct 22, 2025

Oct 22, 2025Wed

Oct 15, 2025

Oct 15, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score73

    Cognition releases SWE-grep models for fast parallel code context retrieval

    AICognition introduces SWE-grep and SWE-grep-mini, fast agentic models trained with reinforcement learning for multi-turn context retrieval in coding tasks. The company says they match frontier coding models at retrieval while taking an order of magnitude less time, and they power the Fast Context subagent in Windsurf. The models issue up to 8 parallel tool calls per turn within 4 turns, and Cerebras serves SWE-grep-mini at over 2,800 tokens per second and SWE-grep at over 650 tokens per second.

    Why it matters: The post explains the speed-intelligence tradeoff in agentic code search, showing how parallel tool calls and RL training change the cost of retrieving context for coding agents.

Sep 28, 2025

Sep 28, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition rebuilds Devin around Claude Sonnet 4.5 for 2x speed

    AICognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.

    Why it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.

Sep 7, 2025

Sep 7, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score53

    Cognition raises over $400M at $10.2B valuation after Windsurf acquisition

    AICognition, maker of the AI software engineer Devin, raised over $400M at a $10.2B post-money valuation led by Founders Fund. The company says its acquisition of Windsurf more than doubled its ARR, with combined enterprise ARR up over 30% in the seven weeks after the deal. It also reports Devin ARR grew from $1M in September 2024 to $73M in June 2025, with total net burn under $20M.

Jun 26, 2025

Jun 26, 2025Thu

May 18, 2025

May 18, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score22

    Cognition Revives Devin Open Source Initiative With $500 Credits for Projects

    AICognition is bringing back its Devin Open Source Initiative, offering $500 in Devin ACU credits to open-source GitHub projects with over 100 forks. Projects below that threshold will still be considered. Eligible projects must have an OSI-approved license and be actively maintained, and maintainers can apply through a linked form.

May 14, 2025

May 14, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Devin 2.1 adds confidence ratings and built-in codebase intelligence

    AICognition has released Devin 2.1, which reports its confidence in completing tasks using green, yellow, and red ratings. The company says green scores led to twice the likelihood of a merged PR compared with red, and Devin now also answers codebase questions and scores Linear and Jira issues.

    Why it matters: The post explains how Devin now shows confidence scores and asks clarifying questions, which changes how teams can decide which tasks to hand over.

May 5, 2025

May 5, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score39

    Kevin-32B Uses Multi-Turn Reinforcement Learning to Write Faster CUDA Kernels

    AIStanford and Cognition AI researchers introduced Kevin-32B, a 32B-parameter model trained with multi-turn reinforcement learning to write CUDA kernels. On KernelBench, it solves 89% of tasks at best@16 and achieves 65% average correctness over eight refinement steps, versus 53% for o4-mini and 51% for o3. Its best@16 speedup is 1.41x, and multi-turn training outperforms single-turn training as refinement steps increase.

May 4, 2025

May 4, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score42

    DeepWiki launches free AI-generated docs for public GitHub repositories

    AICognition has launched DeepWiki, a free public version of its Devin Wiki and Devin Search tools that helps developers understand codebases. Users can view docs for any repo by replacing github.com with deepwiki.com in the URL, and more than 50,000 top public GitHub repositories are already indexed. Private repositories require a Devin account.

Apr 2, 2025

Apr 2, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score75

    Cognition launches Devin 2.0 with agent-native IDE and new planning tools

    AICognition has released Devin 2.0, a new agent-native IDE experience with a flexible plan starting at $20. The update lets users run multiple parallel Devins, each with its own cloud-based IDE, and adds Interactive Planning, Devin Search, and Devin Wiki.

    Why it matters: The release adds planning, codebase search, and auto-generated wikis to Devin, showing how an agent can prepare work before executing it.

Feb 25, 2025

Feb 25, 2025Tue
  1. Cognition Blog (Devin, Windsurf)AI score40

    Devin Gets Batch Edits, GitLab Support, and Sonnet 3.7 in February Update

    AICognition's February 2025 Devin update adds parallel batch edits, beta GitLab support, and Sonnet 3.7, which Cognition says is the best model it has tested for debugging, codebase search, and agentic planning. Devin is about 2x faster than in October 2024, taking about 7.8 minutes on average to complete junior developer tasks in Cognition's internal evaluations. Other changes include copy-paste in Devin's browser and proactive feedback on suboptimal prompts.

Feb 12, 2025

Feb 12, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score36

    Linktree Uses Devin to Add Social Platforms and Ship About 100 Merged PRs

    AILinktree has used Devin, Cognition's AI software engineer, to merge roughly 100 pull requests in a month, mostly fixing customer-reported bugs and implementing small features. The engineering team also used Devin to add support for new social media platforms, launching five Devins, one per repo and PR, and later used the Devin API with a Playbook script to spawn multiple Devins for multi-repo features. The team says Devin works best on tasks an engineer could finish in a couple of hours.

Jan 20, 2025

Jan 20, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin 101: Automatic PR Reviews with the Devin API

    AICognition's Devin can be triggered through its External API by GitHub Actions to automatically review pull requests, typically within five to ten minutes. The setup involves adding a workflow file, storing a DEVIN_API_KEY secret, and customizing the review prompt to match team conventions. Cognition recommends treating Devin as an extra reviewer rather than a replacement for human oversight, since it does not catch every bug.

Jan 15, 2025

Jan 15, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score40

    Devin's January 2025 Update Adds Repo Context, Enterprise Accounts, and Usage Billing

    AICognition's January 2025 Devin update improves its ability to find relevant files and reuse existing code in repositories, with changes rolling out to all users. It also adds enterprise accounts for centralized management of multiple organizations, audio message support in Slack, and pay-as-you-go billing after monthly ACU capacity is used, starting January 9.

Jan 13, 2025

Jan 13, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score38

    Crossmint Uses Devin to Scale Open-Source Development of GOAT SDK

    AICrossmint said Devin became its top contributor to the open-source GOAT SDK during an initial trial, merging 8 pull requests versus 4 for the next contributor. Examples included a DEXScreener plugin built from a documentation URL and a harder Sui blockchain integration that needed three rounds of feedback and about an hour of human involvement. The company said its results depended on proper training, clear task context, and planned validation, not on treating Devin as superhuman.

Dec 22, 2024

Dec 22, 2024Sun
  1. Cognition Blog (Devin, Windsurf)AI score57

    Devin becomes generally available with faster runs and new customization options

    AICognition has made Devin generally available to all engineering teams, with subscriptions starting at $500 per month. Over the past two weeks, Devin was made about 10% faster and about 10% more cost-efficient, especially for tasks requiring many code edits. The update also adds fixes for stuck or hanging sessions, more options to customize filters and Slack notifications, and larger machine settings for disk, RAM, and CPU.

Dec 11, 2024

Dec 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score38

    Devin Open Source Initiative Gives Maintainers 500 Free ACUs for Repo Work

    AICognition is launching the Devin Open Source Initiative, giving selected open source maintainers 500 free ACUs on a Devin Teams plan as part of Devin's general availability launch. The post shows Devin contributing pull requests to projects including Anthropic's MCP Inspector, Dagger, and nanoGPT, with maintainers still reviewing the results. Devin's GitHub integration forwards PR comments and CI checks to help refine changes, though the company warns a human should still verify final quality.

Dec 9, 2024

Dec 9, 2024Mon
  1. Cognition Blog (Devin, Windsurf)AI score67

    Cognition makes Devin generally available to engineering teams from $500 a month

    AICognition is making Devin generally available to engineering teams starting at $500 a month, with no seat limits and access to its Slack integration, IDE extension, and API. The post recommends starting with small frontend bugs, first-draft PRs for backlog tasks, and targeted refactors, and shares open-source PR sessions where Devin resolved issues for projects including Anthropic MCP, Zod, and nanoGPT.

    Why it matters: The post shows concrete open-source PR examples and the tasks where Devin works best, helping teams judge where an autonomous coding agent fits their workflow.

Dec 2, 2024

Dec 2, 2024Mon
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin's December 2024 update adds Slack collaboration, PR tools, and a REST API

    AICognition's December 2024 Devin update lets users tag Devin in Slack threads and ask it to create PRs, with Devin automatically responding to PR comments and lint failures. The update adds Repo Knowledge that Devin generates by scanning repositories, an Agency setting that makes Devin propose plans before executing complex tasks, and a REST API for structured input and output. Devin also gained faster session startup and enterprise options such as Okta single sign-on.

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.

Sep 4, 2024

Sep 4, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score45

    Devin Adds MultiDevin, Auto PR Replies, and Rollback in September 2024 Update

    AICognition's Devin gained a September 2024 update with MultiDevin, which lets a manager Devin delegate work to up to 10 worker Devins, currently available on the Enterprise plan. Devin now automatically responds to comments on its pull requests, suggests Knowledge additions, and can restore earlier checkpoints, and the company reports up to an 80% reduction in time for common tasks.