Skip to contentSkip to stories

Updated

#Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)AI score49

    Devin Introduces New Self-Serve Plans and Charges for Ask Devin and Devin Review

    AIDevin is retiring its Core and Team plans for a new lineup of Free, Pro at $20/month, Max at $200/month, Teams with usage-based billing and an $80/month minimum, and custom-priced Enterprise. Ask Devin's Deep Mode, Devin Review after a 2-week free trial, and higher-quality DeepWiki generation will move to usage-based billing, with DeepWiki's existing generation and open-source Devin Review remaining free. Self-serve usage beyond included quota will be billed in dollars rather than ACUs.

Apr 8

Apr 8Wed
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases open-weight MiniMax-M2.7 with agent and coding gains

    AIMiniMax has released MiniMax-M2.7 on Hugging Face, describing it as its first model to participate in its own evolution. The source reports 56.22% on SWE-Pro, 46.3% on Toolathon, and 62.7% on MM ClawBench, and says an internal version autonomously optimized a programming scaffold over 100+ rounds for a 30% performance improvement.

    Why it matters: The source ties its benchmark claims to a self-evolution process and a named comparison set, which helps readers weigh how the reported gains were achieved.

Apr 7

Apr 7Tue
  1. Cognition Blog (Devin, Windsurf)AI score70

    How Devin Is Modernizing COBOL at Fortune 500 Companies

    AICognition describes how Devin handles COBOL modernization at several Fortune 500 companies, citing a shortage of COBOL developers and 68% failure rates for such efforts. The post identifies three obstacles for agents: untraceable data across copybooks, little COBOL in model training, and no way to run code on Linux-based VMs. It says Devin succeeds on documentation, batch migrations, and large-scale refactoring, while transactional workloads remain out of reach.

    Why it matters: The post explains why agents struggle with COBOL and which workloads they can migrate, giving a framework for judging where automation fits legacy systems.

Apr 6

Apr 6Mon
  1. Z.ai Release NotesAI score34

    Z.ai's GLM-5.3 and GLM-5.2 Lead Open-Source Coding and Long-Context Models

    AIZ.ai's GLM-5.3 delivers a 50% coding gain over GLM-5.2 on Z.ai Code Bench, reaching open-source state-of-the-art on public benchmarks including Terminal Bench 3.0. GLM-5.3-Flash uses 320B total parameters with 18B activated, combining linear and sparse attention to reduce compute and KV-cache needs. GLM-5.2 supports a 1M lossless context window for long-horizon tasks.

  2. Cognition Blog (Devin, Windsurf)AI score44

    Windsurf releases SWE-1.6, a software engineering model optimized for speed and user experience

    AIWindsurf has made SWE-1.6, its model for software engineering agents, generally available, with the company saying it improves on the SWE-1.6 Preview by reducing overthinking, looping, and sequential tool calls. The model is free for three months, with a free version offered at 200 tok/s through Fireworks and a faster paid version at 950 tok/s through Cerebras.

Apr 3

Apr 3Fri
  1. Z.ai (GLM) · new models on Hugging FaceAI score73

    Z.ai releases GLM-5.1, a flagship model for agentic engineering

    AIZ.ai has released GLM-5.1, its next-generation flagship model for agentic engineering, with stronger coding than GLM-5. The model is described as staying effective over longer agentic tasks, sustaining optimization over hundreds of rounds and thousands of tool calls. The release lists benchmark results including SWE-Bench Pro at 58.4 and Terminal-Bench 2.0 at 63.5, and local deployment is supported through SGLang, vLLM, xLLM, Transformers, and KTransformers.

    Why it matters: The release gives benchmark tables against several rival models, letting readers compare GLM-5.1's coding and agentic results with GLM-5 and frontier systems.

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.

Mar 23

Mar 23Mon
  1. Anthropic EngineeringAI score78

    Anthropic shows a three-agent harness for long-running app development

    AIAnthropic's Labs team describes a three-agent harness with planner, generator, and evaluator agents for building full-stack applications over multi-hour autonomous coding sessions. The evaluator uses Playwright to test the running app against sprint contracts, and a retro game maker built with the harness worked end to end where a single-agent run's core feature did not. The author later removed the sprint construct and kept only the components still needed on Opus 4.6.

    Why it matters: The post shows how a generator-evaluator loop, with explicit grading criteria and a tuned QA agent, turned a solo run's broken output into a working app, and how the harness was pruned as models improved.

Mar 18

Mar 18Wed
  1. Cognition Blog (Devin, Windsurf)AI score72

    Devin can now break tasks down and run a team of managed Devins

    AIDevin can now break large tasks into scoped pieces and delegate them to a team of managed Devins that run in parallel. Each managed Devin runs in its own isolated virtual machine with its own terminal, browser, and development environment, and has its own session link. The main coordinator session monitors progress, resolves conflicts, and compiles results, and managed Devins are available now for all users.

    Why it matters: The post explains how a coordinator session splits work across isolated managed sessions, giving readers a concrete pattern for running agent tasks in parallel.

Mar 17

Mar 17Tue
  1. MiniMax BlogAI score63

    MiniMax M2.7 takes part in its own model and harness evolution

    AIMiniMax says M2.7 is its first model to deeply participate in its own evolution, building agent harnesses and running reinforcement learning experiment workflows. The post reports 56.22% on SWE-Pro, 55.6% on VIBE-Pro, 57.0% on Terminal Bench 2, and a 30% improvement on an internal evaluation set after more than 100 autonomous optimization rounds. It also states that M2.7 handles 30%-50% of its research team's workflow, though human researchers still make critical decisions.

    Why it matters: The post ties M2.7's self-evolution claims to specific benchmark numbers and workflow details, helping readers judge how much of the iteration loop is autonomous.

  2. Xiaomi MiMoAI score80

    Xiaomi MiMo-V2-Pro Flagship Model Targets Agent Workloads With 1M Context

    AIXiaomi announced MiMo-V2-Pro, a flagship foundation model for agent workloads with over 1T total parameters, 42B active, and up to 1M-token context. It ranks 8th worldwide and 2nd among Chinese LLMs on the Artificial Analysis Intelligence Index, and its API is publicly available with usage-tiered pricing.

    Why it matters: The post gives benchmark placements, parameter scale, context length, and tiered API pricing, so readers can compare it against Claude and GPT models on concrete terms.

  3. Apple · new models on Hugging FaceAI score44

    Apple releases SimpleSD-30B-instruct, a self-distilled Qwen code model for research

    AIApple has released apple/SimpleSD-30B-instruct, a research checkpoint built on Qwen that uses Simple Self-Distillation to improve code generation without rewards, verifiers, or teacher models. On LiveCodeBench, the model scores 55.3% pass@1 on LCBv6 versus 42.4% for its base, Qwen3-30B-A3B-Instruct-2507. The checkpoints are for reproducibility, not optimized Qwen releases, and are available under the Apple Machine Learning Research Model License.

  4. Apple · new models on Hugging FaceAI score43

    Apple releases SimpleSD-4B-thinking, a self-distilled Qwen model for code generation

    AIApple has published SimpleSD-4B-thinking on Hugging Face, a research checkpoint built on Qwen that improves code generation through Simple Self-Distillation without rewards, verifiers, teacher models, or reinforcement learning. On LiveCodeBench, it lifts Qwen3-4B-Thinking-2507 from 54.5% to 57.8% pass@1 on LCBv6 and from 59.6% to 63.1% pass@1 on LCBv5. The model is released as a reproducibility checkpoint under the Apple Machine Learning Research Model License, not as an optimized Qwen release.

  5. Apple · new models on Hugging FaceAI score46

    Apple releases SimpleSD-4B-instruct, a self-distilled Qwen code model

    AIApple has released SimpleSD-4B-instruct on Hugging Face, a research checkpoint fine-tuned from Qwen3-4B-Instruct-2507 on its own sampled outputs to improve code generation. On LiveCodeBench, the model scores 41.5% pass@1 on LCBv6, up from the base model's 34.0%, and 45.7% pass@1 on LCBv5, up from 34.3%. The model is released under the Apple Machine Learning Research Model License and is intended for reproducibility rather than as an optimized Qwen release.

Feb 28

Feb 28Sat
  1. Cognition Blog (Devin, Windsurf)AI score36

    Cognition Previews SWE-1.6, Claims 11% Gain Over SWE-1.5 on SWE-Bench Pro

    AICognition previewed its ongoing SWE-1.6 training run, which scores 11% higher than SWE-1.5 on SWE-Bench Pro and runs at 950 tok/s. The model is post-trained on the same pre-trained model as SWE-1.5, and the company is rolling out early access to a small group of users to gather feedback on behavior such as overthinking and excessive self-verification. The company says training steps now run 6x faster than three months ago, with rollouts in NVFP4 precision.

Feb 26

Feb 26Thu
  1. Cognition Blog (Devin, Windsurf)AI score67

    How Cognition Uses Devin to Build Devin Across Slack, Linear, and Code Review

    AICognition reports merging 659 Devin PRs into its own codebase last week, up from 154 in its best week in 2025. The post describes internal workflows across web, Slack, Linear, CLI, and API, including Devin Review for PR diffs and bug catching, a daily design system audit, automated bug triage on Linear, and DANA for data analysis.

    Why it matters: The post shows concrete workflows for using Devin across Slack, Linear, and code review, with specific usage figures that help teams judge fit for their own engineering processes.

Feb 24

Feb 24Tue
  1. Cognition Blog (Devin, Windsurf)AI score46

    Cognition Launches Cognition for Government to Modernize Federal Software With Devin and Windsurf

    AICognition launched Cognition for Government on February 25, 2026, offering its Devin autonomous software engineering agent and Windsurf AI IDE to modernize U.S. government legacy systems. Devin, available in AWS GovCloud with a FedRAMP High version forthcoming, can complete migrations 5-40x faster than human engineers, while Windsurf is the only FedRAMP High AI IDE and holds DoD IL4/5/6 accreditation.

Feb 23

Feb 23Mon
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin 2.2 adds desktop testing, self-review autofix, and 3x faster startup

    AICognition released Devin 2.2, which gives Devin full access to its own Linux desktop so it can launch and test desktop applications, not just browser-based web apps. Devin can also plan, code, review its own output, and fix issues before opening a PR, and it now starts up 3x faster. New users get $10 in free credits, and Desktop support is enabled by default for new sessions as of February 24, 2026.

Feb 12

Feb 12Thu
  1. MiniMax · new models on Hugging FaceAI score88

    MiniMax releases M2.5 model with 80.2% on SWE-Bench Verified

    AIMiniMax has released M2.5, which it says reaches 80.2% on SWE-Bench Verified and 76.3% on BrowseComp with context management. The company reports 37% faster end-to-end runtime than M2.1 on SWE-Bench Verified and prices M2.5 at $1 per hour at 100 tokens per second, with a 50 tokens per second version at $0.30 per hour. Weights are available on Hugging Face, with inference support listed for SGLang, vLLM, Transformers, and KTransformers.

    Why it matters: The source gives benchmark scores against Claude and GPT models plus per-task token and runtime figures, so readers can weigh the cost-speed tradeoff directly.

Feb 11

Feb 11Wed
  1. Z.ai Release NotesAI score49

    Z.ai Releases GLM-5.3-Flash, GLM-5.3 and a Series of Updated GLM Models

    AIZ.ai's release notes list GLM-5.3-Flash, a hybrid-architecture model with 320B total parameters and 18B activated, and GLM-5.3, which the company says achieves a 50% gain over GLM-5.2 on Z.ai Code Bench. Other entries in the notes include GLM-5.2 with 1M lossless context and GLM-5.1, which Z.ai says can work independently for up to 8 hours in a single run.

Feb 10

Feb 10Tue
  1. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5, a 744B-parameter open model for agentic engineering

    AIZ.ai launches GLM-5, scaling from 355B to 744B total parameters with 40B active and pre-training data from 23T to 28.5T tokens. The model integrates DeepSeek Sparse Attention to reduce deployment cost and reports strong results on reasoning, coding, and agentic benchmarks against GLM-4.7, DeepSeek-V3.2, Kimi K2.5, and several frontier models.

    Why it matters: The source gives concrete scale, data, and benchmark comparisons against named frontier models, showing where GLM-5 sits among open-source and proprietary systems.

Feb 9

Feb 9Mon
  1. Cognition Blog (Devin, Windsurf)AI score43

    Devin Can Now Autofix Review Comments from Devin Review and Other Bots

    AICognition has configured Devin to automatically autofix incoming review comments from Devin Review and other PR review bots, as well as lint and CI/CD issues. Devin resolves flagged problems and feeds the fixes back into the pull request without human intervention for mechanical fixes. Users can select which bots Devin responds to in Settings > Customization > Autofix settings.

Feb 4

Feb 4Wed
  1. Anthropic EngineeringAI score72

    Anthropic finds container resource limits can shift agentic coding eval scores

    AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

    Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.

  2. Anthropic EngineeringAI score75

    Anthropic details how parallel Claude agents built a 100,000-line C compiler

    AINicholas Carlini of Anthropic's Safeguards team describes an agent-team setup where 16 Claude instances worked in parallel on a shared codebase without human intervention to write a Rust-based C compiler. Over nearly 2,000 Claude Code sessions costing about $20,000 in API fees, the team produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The post focuses on harness design, including high-quality tests, lock files for task claiming, GCC as a reference oracle for the kernel, and the limits the project reached.

    Why it matters: The post shows concrete harness design choices for long-running agent teams, including test design, locking, and parallel work division, that readers can adapt to their own autonomous projects.

Jan 27

Jan 27Tue
  1. Cognition Blog (Devin, Windsurf)AI score32

    Cognition opens London office to expand Devin autonomous coding for European businesses

    AICognition is opening a London office to expand rollout of Devin, its autonomous software engineering agent, to leading European businesses. The company says finance has emerged as a clear use case, with Goldman Sachs, Santander, Citi, and BNY among partners using Devin for modernization, migration, security remediation, and codebase documentation.

  2. Cognition Blog (Devin, Windsurf)AI score38

    Cognizant Partners with Cognition to Scale Devin and Windsurf Across Its Engineering Teams

    AICognizant is deploying Cognition's Devin autonomous software engineer and Windsurf agentic IDE across its engineering organization and global client base. Engineers already use Windsurf for agent-assisted coding and are exploring Devin for end-to-end tasks such as code migration, refactoring, testing, and maintenance. Cognition will embed forward-deployed AI engineers to support project selection, engineer enablement, and ROI measurement.

Jan 20

Jan 20Tue
  1. Cognition Blog (Devin, Windsurf)AI score54

    Cognition launches Devin Review to help humans review AI-generated code

    AICognition introduced Devin Review, a free early-release code review tool that works on any public or private GitHub PR, with features for organizing diffs, chatting about changes, and flagging AI-detected bugs. The company says code review, not code generation, is now the bottleneck as coding agents increase the volume and size of pull requests.

Jan 19

Jan 19Mon
  1. Factory NewsAI score47

    Factory Introduces Agent Readiness to Score Codebases for Autonomous Coding Agents

    AIFactory's new Agent Readiness tool evaluates repositories across eight technical pillars and five maturity levels, using 60+ binary criteria run via the /readiness-report command. The company says it can also open pull requests to fix foundational gaps such as missing AGENTS.md files, linter configuration, and pre-commit hooks. Factory says scores are now more consistent, with variance dropping from an average of 7% to 0.6%.

Dec 20, 2025

Dec 20, 2025Sat
  1. MiniMax · new models on Hugging FaceAI score74

    MiniMax-M2.1 open-sources weights for coding and agent tasks

    AIMiniMax has released MiniMax-M2.1 model weights on Hugging Face, with API access on the MiniMax Open Platform and the MiniMax Agent product. The company reports gains over M2 on coding and agent benchmarks such as SWE-bench Verified (74.0) and VIBE average (88.6), and says it outperforms Claude Sonnet 4.5 on multilingual scenarios.

    Why it matters: The release pairs open weights with a broad benchmark table against Claude and GPT models, letting readers compare coding and agent claims directly.

Dec 16, 2025

Dec 16, 2025Tue
  1. Xiaomi MiMoAI score78

    Xiaomi releases open-source MiMo-V2-Flash MoE model for reasoning and coding

    AIXiaomi released and open-sourced MiMo-V2-Flash, a Mixture-of-Experts model with 309B total and 15B active parameters, under the MIT license. The company reports 73.4% on SWE-Bench Verified, the top score among open-source models, and inference at 150 tokens per second for $0.1 per million input tokens and $0.3 per million output tokens. It supports a hybrid thinking mode and a 256k context window.

    Why it matters: The post gives architecture, speculative decoding speedup, and pricing figures, which help readers judge how the efficiency claims are achieved and what they cost.

Nov 13, 2025

Nov 13, 2025Thu
  1. Cognition Blog (Devin, Windsurf)AI score65

    Cognition's Devin review says it excels at scoped junior-level engineering work

    AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.

    Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.

Nov 3, 2025

Nov 3, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score47

    Windsurf Codemaps Adds AI-Annotated Code Maps to Help Engineers Understand Code

    AIWindsurf has launched Codemaps, AI-annotated structured maps of a codebase powered by SWE-1.5 and Claude Sonnet 4.5, which users can generate from a task prompt using a Fast (SWE-1.5) or Smart (Sonnet 4.5) model. Codemaps links grouped code sections to exact lines and can be referenced in Cascade with @{codemap} to give agents more specific context.

Oct 28, 2025

Oct 28, 2025Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition releases SWE-1.5, a coding agent model served at up to 950 tok/s

    AICognition has released SWE-1.5, a model optimized for software engineering that it says reaches near-frontier coding performance while running at up to 950 tok/s with Cerebras inference. The company reports it is 6x faster than Haiku 4.5 and 13x faster than Sonnet 4.5, and it is available now in Windsurf. The post's SWE-Bench Pro chart places SWE-1.5 at 40.08%, behind Sonnet 4.5 at 43.60%, and it notes that the model was trained with reinforcement learning on the Cascade agent harness.

    Why it matters: The post pairs a benchmark chart with a 950 tok/s speed claim and describes how harness, RL environments, and inference were co-designed, useful context for judging the speed-versus-quality tradeoff.

Oct 27, 2025

Oct 27, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score36

    Devin Automates .NET Framework to .NET Core Migration in Weeks, Not Months

    AICognition says its autonomous coding agent Devin can complete a .NET Framework to .NET Core migration in as little as two weeks, using a Strangler Fig approach adapted from Jimmy Bogard's guide. The post says Devin handles planning via Ask Devin and DeepWiki, dependency sharing, controller and view conversion, and session state adaptation through a remote app.

Oct 15, 2025

Oct 15, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score73

    Cognition releases SWE-grep models for fast parallel code context retrieval

    AICognition introduces SWE-grep and SWE-grep-mini, fast agentic models trained with reinforcement learning for multi-turn context retrieval in coding tasks. The company says they match frontier coding models at retrieval while taking an order of magnitude less time, and they power the Fast Context subagent in Windsurf. The models issue up to 8 parallel tool calls per turn within 4 turns, and Cerebras serves SWE-grep-mini at over 2,800 tokens per second and SWE-grep at over 650 tokens per second.

    Why it matters: The post explains the speed-intelligence tradeoff in agentic code search, showing how parallel tool calls and RL training change the cost of retrieving context for coding agents.

Sep 28, 2025

Sep 28, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition rebuilds Devin around Claude Sonnet 4.5 for 2x speed

    AICognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.

    Why it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.

Sep 7, 2025

Sep 7, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score53

    Cognition raises over $400M at $10.2B valuation after Windsurf acquisition

    AICognition, maker of the AI software engineer Devin, raised over $400M at a $10.2B post-money valuation led by Founders Fund. The company says its acquisition of Windsurf more than doubled its ARR, with combined enterprise ARR up over 30% in the seven weeks after the deal. It also reports Devin ARR grew from $1M in September 2024 to $73M in June 2025, with total net burn under $20M.

Jun 26, 2025

Jun 26, 2025Thu

May 18, 2025

May 18, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score22

    Cognition Revives Devin Open Source Initiative With $500 Credits for Projects

    AICognition is bringing back its Devin Open Source Initiative, offering $500 in Devin ACU credits to open-source GitHub projects with over 100 forks. Projects below that threshold will still be considered. Eligible projects must have an OSI-approved license and be actively maintained, and maintainers can apply through a linked form.