Skip to contentSkip to stories

Updated

#Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Jul 9

Jul 9Thu
  1. Fidji SimoXAI score62

    OpenAI launches ChatGPT Work, an agent powered by Codex and GPT-5.6

    AIOpenAI introduced ChatGPT Work, a new agent inside ChatGPT powered by Codex and GPT-5.6. The quoted announcement says it can take action across apps and files, stay with a project for hours if needed, and turn a goal into finished work. Fidji Simo's own post adds that the team has worked to make Chat more agentic for a while.

  2. Meta AI BlogOfficialAI score72

    Meta releases Muse Spark 1.1 with agent and coding gains

    AIMeta Superintelligence Labs has introduced Muse Spark 1.1, a multimodal reasoning model aimed at agentic tasks, with gains in tool use, computer use, coding, and multimodal understanding. It supports a 1 million token context window and is available in Thinking mode in the Meta AI app and on meta.ai, with developers able to access it through a public preview of the Meta Model API.

    Why it matters: The post specifies Muse Spark 1.1's agent, coding, and multimodal gains and its Meta Model API preview access, which helps developers judge its fit for their workflows.

Jul 8

Jul 8Wed
  1. Michael TruellXAI score57

    Cursor and SpaceXAI release Grok 4.5, a coding-focused model

    AICursor co-founder Michael Truell announced Grok 4.5, a model trained with SpaceXAI that the post calls Opus-class, fast, and low cost. He says it is a significant step up over Composer 2.5 and has become the daily driver for many on the Cursor team. A benchmark table shows Grok 4.5 at 83.3% on Terminal-Bench 2.1 and 78.0% on SWE-Bench Multilingual, with the post saying more releases will follow.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition releases SWE-1.7, a coding model trained with long-horizon RL

    AICognition launched SWE-1.7, which it says reaches frontier-level coding performance at lower cost, trained from a Kimi K2.7 base. The post describes RL methods including top-p sampling replay to preserve entropy, compressed weight deltas across multi-cluster training, and self-compaction for rollouts up to six hours. SWE-1.7 is available in Devin via Cerebras at 1000 TPS.

    Why it matters: The post details entropy preservation, multi-cluster weight sync, and self-compaction, offering concrete RL training techniques for long-horizon coding agents to compare against one's own pipeline.

  3. Cognition Blog (Devin, Windsurf)OfficialAI score47

    Cognition Tests Trustworthiness of SWE-1.7, Built on Kimi K2.7 Code

    AICognition says its SWE-1.7 model, developed from the open-source Kimi K2.7 Code base, performs as well as or better than leading U.S. frontier models on its new trustworthiness evaluation suite. The suite combines 145 politically sensitive questions, sampled in English and Chinese, with realistic coding scenarios to measure propaganda, censorship, and security behavior. Cognition says SWE-1.7 improves substantially over the base Kimi K2.7 Code model, though the company says the benchmarks are still in development.

Jul 7

Jul 7Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score39

    FrontierCode 1.1 refines its code-quality benchmark to curb unfair internet use

    AICognition released FrontierCode 1.1, an update to its code-quality benchmark that adds a fair internet use prompt and a verifier that zeroes out runs consulting upstream fixes. The company also relaxed 75 of over 1,000 grading criteria, added scores for Sonnet 5 and updated scores for Fable 5, and dropped reporting on the Diamond subset.

Jul 6

Jul 6Mon

Jul 2

Jul 2Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score38

    Cognition launches Devin Security Vulnerability Remediation Program for enterprise backlogs

    AICognition launched the Devin Security Vulnerability Remediation Program, in which its forward-deployed engineers embed with customer teams to deploy Devin to find, validate, and fix vulnerabilities. The program first works through existing scanner backlogs from tools such as Snyk, SonarQube, and Semgrep, shipping validated fixes as pull requests, then adds Devin Security Swarm for continuous discovery of logic flaws. Most engagements run about six weeks, and eligibility is limited to enterprise Devin Cloud customers meeting the program's requirements.

Jul 1

Jul 1Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score57

    Cognition launches Devin Security Swarm to find, verify, and patch vulnerabilities

    AICognition has launched Devin Security Swarm, which uses parallel agents to find vulnerabilities across a codebase, confirms exploitability in isolated sandboxes, and opens remediation PRs. In an evaluation on 50 real-world GitHub Security Advisory vulnerabilities, Devin reached 72% recall at about $90.23 per run, compared with 68% for Claude Security at $131.87 per run. The product is available starting today, with scan profiles and incremental scans that process only changed code after the first full baseline.

  2. Mistral AI · new models on Hugging FaceOfficialAI score54

    Mistral AI releases Leanstral 1.5, an open-source Lean 4 code agent model

    AIMistral AI released Leanstral 1.5 on Hugging Face as an open-source code agent model for Lean 4 proof assistant tasks. The model uses 119B total parameters with 6.5B activated per token, a 256k context length, and accepts text and image input. The source gives setup paths through Mistral Vibe and a local vLLM server, with recommended settings of temperature 1.0 and reasoning effort set to high for complex prompts. The model is licensed under Apache 2.0.

Jun 30

Jun 30Tue
  1. Andrew NgXAI score50

    Andrew Ng outlines three loops for building 0-to-1 AI products

    AIAndrew Ng describes three loops he uses to build 0-to-1 products with AI agents: an agentic coding loop, a developer feedback loop, and an external feedback loop. He says the agentic coding loop runs every few minutes, letting coding agents build, test, and iterate on software for around an hour without human intervention. The developer feedback loop operates over tens of minutes to hours, with humans steering product decisions because they hold a context advantage over AI systems.

    Image from @AndrewYNg's post
  2. Xiaomi MiMoOfficialAI score22

    Xiaomi MiMo praised as developers build on open-weights models

    AIXiaomi MiMo's account celebrated growing developer adoption of its open-weights models, crediting Cline for building on MiMo. Cline's linked post announced a $9.99/month subscription offering 2-5x discounted access to GLM-5.2 and other open-weight models including DeepSeek, Kimi, MiniMax, MiMo, and Qwen, with a $1.99 promo for sign-ups via npm i -g cline.

Jun 29

Jun 29Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition's Devin Fusion routes coding work between two models to cut cost

    AICognition has released a preview of Devin Fusion, a multi-model harness that runs a frontier main agent alongside a cheaper sidekick agent. On FrontierCode 1.1 Extended, the company reports scores near frontier models at up to 60% lower cost per task, and 41% lower cost when paired with Fable 5, which access was suspended from June 12, 2026.

    Why it matters: The post explains a sidekick architecture with cached persistent contexts, which contrasts with advisor-style tools and shows how cost cuts depend on the main model's delegation behavior.

Jun 27

Jun 27Sat
  1. Ahead of AI (Sebastian Raschka)BlogAI score37

    Local Coding Agents: Setting Up Qwen3.6 with Open-Source Harnesses

    AISebastian Raschka's tutorial shows how to build a fully local coding agent by pairing an open-weight LLM served through an inference runtime with an open-source harness that can read files, edit code, and run commands. He recommends Qwen-Code for Qwen3.6, citing Nvidia's Polar paper, which found Qwen models performed best in Qwen-Code. The Qwen3.6 35B-A3B model is about 22 GB to download and needs roughly 30–40 GB of RAM.

Jun 23

Jun 23Tue

Jun 19

Jun 19Fri
  1. Mark ChenXAI score46

    Codex now hands off threads between local and remote hosts

    AICodex can now transfer work threads between a local laptop and a remote host, letting users continue tasks elsewhere after closing the lid. Codex can also orchestrate the handoff automatically. Mark Chen joked that this removes the need to prop the laptop open with physical claws.

Jun 16

Jun 16Tue
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.2 with 1M-token context and MIT open-source license

    AIZ.ai has released GLM-5.2, its flagship model for long-horizon tasks, which it says substantially improves on GLM-5.1 and supports a 1M-token context. The model adds IndexShare, which cuts per-token FLOPs by 2.9× at 1M context, and is released under the MIT open-source license.

    Why it matters: The source gives benchmark tables against named rival models and deployment settings, useful for judging where GLM-5.2 sits among current flagship models.

Jun 15

Jun 15Mon
  1. Z.ai Release NotesOfficialAI score62

    Z.ai Release Notes: GLM-5.2 Adds 1M Lossless Context for Long Tasks

    AIZ.ai's release notes list GLM-5.2 as supporting 1M lossless context, with improved long-horizon task performance and reduced context drift and goal forgetting. The company says GLM-5.2 achieves open-source SOTA performance on coding and long-horizon task benchmarks. The page also includes the newer GLM-5.3 and GLM-5.3-Flash entries, which are listed above GLM-5.2.

    Why it matters: The page lists a dated series of Z.ai model releases, showing how the coding and long-horizon agent line has evolved from GLM-4.5 through GLM-5.2.

  2. Werner VogelsXAI score22

    Werner Vogels says verification is now the key bottleneck

    AIAmazon CTO Werner Vogels argues that the verification bottleneck matters more than ever. The post is a short reply to Yaron Minsky's remarks on Jane Street's ambition to make formal methods as pervasively useful for building software as sophisticated type systems are today.

Jun 11

Jun 11Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score62

    Moonshot AI releases Kimi K2.7 Code, a coding-focused agentic model

    AIMoonshot AI published Kimi-K2.7-Code, a coding-focused agentic model built on Kimi K2.6, with a 1T-parameter MoE architecture and 32B activated parameters. The model card reports about 30% fewer thinking tokens than K2.6 and benchmark results against GPT-5.5 and Claude Opus 4.8, with weights and code released under a Modified MIT License.

    Why it matters: The model card gives benchmark comparisons against GPT-5.5 and Claude Opus 4.8 on coding and agentic tasks, useful for judging its position among current coding models.

Jun 10

Jun 10Wed
  1. AI Snake OilBlogAI score70

    Why AI hasn't replaced software engineers, and why it likely won't

    AIThe essay argues that AI compresses the execution layer of software work while decision-making and accountability remain human, so AI is not yet replacing software engineers. It cites AI-attributed layoffs at Block, Snap, and Intuit that the authors say were not driven by AI, and WARN Act filings in which only one company checked an AI box. A Federal Reserve analysis is cited as finding software engineer employment growing about 3 percentage points per year more slowly after ChatGPT than a no-AI counterfactual.

  2. Factory NewsOfficialAI score58

    Factory launches automated STRIDE-based security review for pull requests in Droid

    AIFactory is rolling out automated security review in Droid, running a STRIDE-based check on every non-draft PR alongside standard code review. Findings include severity, a CWE reference, an explanation, and a suggested fix, posted as inline diff comments. The feature is available today on all plans, and a deeper multi-agent /security-review deep audit is available for full-repository scans.

  3. Zed BlogOfficialAI score48

    Zed Unveils DeltaDB, Version Control Built Around Agent Conversations Instead of Commits

    AIZed is building DeltaDB, a version control system that records every operation as a fine-grained delta, linking agent conversations to the code they produce. The company says a beta will arrive in a few weeks, and it invites users to join a waitlist. The system is designed so teammates can collaborate on work in progress without waiting for commits, pull requests, or pushes.

  4. Xiaomi MiMoOfficialAI score67

    Xiaomi releases open-source MiMo Code V0.1 terminal coding assistant

    AIXiaomi MiMo has released MiMo Code V0.1, an open-source AI coding assistant for the terminal under the MIT license. It ships with MiMo V2.5, a multimodal model offered free for a limited time with a million-token context window. The tool automatically loads existing Claude Code skills, MCP servers and commands, and reuses API configuration, and it supports providers including Anthropic, OpenAI, DeepSeek, Kimi and GLM.

    Why it matters: The post specifies MiMo Code's Claude Code compatibility and MIT license, which bear directly on whether existing coding-agent setups can migrate without rework.

    Image from @XiaomiMiMo's post
  5. Xiaomi MiMoOfficialAI score82

    MiMo Code open-sources a terminal coding agent for long-horizon tasks

    AIXiaomi's MiMo team released MiMo Code, an MIT-licensed terminal coding agent built on OpenCode for long-horizon programming tasks. The design centers on three areas: Max Mode parallel sampling that generates five candidates per turn, Goal-based completion verification, and a memory system that checkpoints session state and rebuilds context. The article reports offline benchmark results and a double-blind A/B test with 1,213 pairs in which MiMo Code's win rate exceeded 65% beyond 200 execution steps.

    Why it matters: The article explains how MiMo Code handles long-horizon coding through computation, checkpointed memory, and cross-session evolution, useful for judging design tradeoffs in coding agents.

Jun 9

Jun 9Tue
  1. One Useful Thing (Ethan Mollick)BlogAI score72

    Ethan Mollick tests Claude 5 Fable and finds it runs long projects with little user input

    AIEthan Mollick, who had early access to Claude 5 Fable, reports that it outperformed other public models in his tests, including an isochrone travel-time map and a nine-and-a-half-hour software build called Concord. He says the model delegated work to other agents and made many design choices he could not see or weigh in on, leaving him closer to a client than a hands-on operator. He also notes high token usage, frequent fallback to Claude 4.8 Opus under security guardrails, and persistent quirks in its writing style.

Jun 8

Jun 8Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score70

    Cognition Introduces FrontierCode, a Benchmark for Mergeable Code Quality

    AICognition introduced FrontierCode, a coding benchmark built with open-source maintainers that measures whether models produce code a maintainer would merge. On FrontierCode Diamond, the hardest 50 tasks, Claude Opus 4.8 scored 13.4%, GPT-5.5 scored 6.3%, and Gemini 3.1 Pro scored 4.7%. The authors report 81% fewer misclassification errors than SWE-Bench Pro, though this figure comes from their own analysis of agent trajectories.

    Why it matters: The benchmark's blocker and rubric design shows how code quality can be measured beyond unit-test correctness, which matters for judging coding agents.

  2. Xiaomi MiMo · new models on Hugging FaceOfficialAI score41

    Xiaomi releases MiMo-V2.5-Pro-FP4-DFlash, an FP4 model with block-diffusion decoding

    AIXiaomi MiMo has released MiMo-V2.5-Pro-FP4-DFlash, the FP4 backbone behind MiMo-V2.5-Pro-UltraSpeed, with MXFP4 quantization applied only to the MoE experts and a BF16 DFlash drafter for block-diffusion speculative decoding. The backbone has 1.02T total and 42B active parameters, and the drafter proposes blocks of up to 8 tokens per forward pass. The release is supported in SGLang, with example launch commands provided.

Jun 4

Jun 4Thu
  1. Cohere · new models on Hugging FaceOfficialAI score60

    Cohere releases North Mini Code 1.0, a 30B-A3B open-weights coding model

    AICohere and Cohere Labs released North Mini Code 1.0, an open-weights 30B-A3B mixture-of-experts model for code generation and agentic terminal tasks, under Apache 2.0. The model has 256K context and 64K max output, and is trained for tool use. Its benchmark table lists Terminal-Bench v2 at 36.0, SWE-Bench Verified at 67.6, and LiveCodeBench v6 at 70.3, below Qwen3.6 on several tasks.

    Why it matters: The card lists benchmark results against Qwen3.6, Gemma4, and other models, showing where North Mini Code trails on some coding and agentic tasks.

  2. One Useful Thing (Ethan Mollick)BlogAI score44

    Ethan Mollick Announces Co-Existence, a Sequel Book on Working Alongside AI

    AIEthan Mollick is releasing Co-Existence on October 20, a new book about working with AI systems that are sometimes, but not always, better than humans. The book follows his 2024 title Co-Intelligence, which he says was written about an era of chatbots rather than autonomous agents. Mollick also reports writing every chapter draft himself while using AI readers and fact-checkers, and building the book's website with Claude Code using Opus 4.8.

Jun 3

Jun 3Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition launches $10M AI Productivity Guarantee for enterprise Devin customers

    AICognition introduced the AI Productivity Guarantee, under which it will issue credits up to $10M if Devin delivers less engineering value than enterprise customers pay for. The company uses an AI estimator to measure hours of productive output, validated against engineers' own estimates of how long the same work would have taken by hand. Value is converted to dollars at a standard global rate and compared against each customer's consumption near the end of the annual contract.

    Why it matters: The post explains how Cognition estimates Devin's output in hours and backs the estimate with a $10M credit commitment, a concrete model for measuring AI vendor value.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition Estimates Engineering Hours Saved by Its Devin Coding Agent

    AICognition built an automated agent that classifies Devin sessions as productive and estimates the human engineering hours each one would have taken. On 233 held-out sessions the estimator reached an rlog of 0.74, with individual errors often 2 to 3 times in either direction but roughly unbiased in aggregate. The system is calibrated to underestimate and is currently running with Devin customers.

    Why it matters: The post shows how the measurement design, from hours-based metrics to conservative calibration, determines whether agent productivity estimates can be trusted in aggregate.

Jun 2

Jun 2Tue