Skip to contentSkip to stories

Updated

#Coding

Showing low-relevance items too. Hide low-relevance items

Jul 28

Jul 28Tue

Jul 27

Jul 27Mon

Jul 26

Jul 26Sun
  1. Fireworks AI BlogAI score60

    Fireworks AI adds open-weight Kimi K3 with US-only serverless endpoints

    AIFireworks AI made the open-weight Kimi K3 available for inference and training on its platform, with US-only serverless endpoints and Zero Data Retention. In its own head-to-head with Opus 5, the post reports K3 at 92.7% accuracy and $0.52 per task on SWE (480) against Opus 5's 94.8% and $1.05, with the vendor claiming up to 5x better cost efficiency per task.

    Why it matters: The post compares Kimi K3 with Opus 5 on accuracy and cost per task, giving readers concrete figures to judge the open model against closed alternatives for their own workloads.

  2. Philipp SchmidAI score62

    EvoCode-Bench Tests Coding Agents Across Multi-Turn Iterative Specification Changes

    AIEvoCode-Bench is a multi-turn coding benchmark with 26 tasks spanning 227 sequential rounds, where agents keep a persistent workspace and must pass cumulative tests after each evolving instruction. The results show that agents perform much worse when building on their own prior work than when starting from a clean, human-completed codebase. Regressions, not failure to implement new features, are the main bottleneck, and agents that maintained a persistent requirements document more than doubled their success rates.

Jul 25

Jul 25Sat

Jul 24

Jul 24Fri
  1. Mike KriegerAI score46

    Mike Krieger says Claude Opus 5 became his daily driver

    AIAnthropic co-founder Mike Krieger says Claude Opus 5 has become his daily driver at work and on weekends. He reports it can work for hours on complex tasks and consistently gets to the bottom of tricky problems, and he has also built some games with it. Anthropic's announcement describes Opus 5 as close to the frontier intelligence of Fable 5 at half the price.

Jul 23

Jul 23Thu
  1. Matei ZahariaAI score36

    Berkeley STAR Lab packages AI research optimizers into one GEPA API

    AIBerkeley's STAR Lab packaged multiple LLM-based "autoresearch" algorithms into a single API within the GEPA package, letting users mix and match them. The optimizers can be applied to tasks including prompt writing, agent design, and code optimization. The quoted thread adds that GEPA, AutoResearch, and Meta-Harness each win on different tasks, and that the new optimize_anything omni meta-optimizer beats every standalone optimizer at a matched budget.

  2. One Useful Thing (Ethan Mollick)AI score67

    Ethan Mollick's guide to choosing AI tools for agentic work

    AIEthan Mollick's guide says ChatGPT and Claude are the main choices for real work, since their agent modes can act on a computer. He separates agent modes that run on the company's computers from those that access the user's own computer. He recommends keeping approval settings on for sending, spending, or deleting, because of prompt injection risk. He also notes that Gemini currently lags for agentic work, though its Notebook and video tools are useful.

Jul 22

Jul 22Wed
  1. Cognition Blog (Devin, Windsurf)AI score41

    Cognition signs MOU with U.S. Department of Energy to join Genesis Mission

    AICognition has signed a memorandum of understanding with the U.S. Department of Energy to join the Genesis Mission, a national AI initiative launched by executive order in November 2025. Cognition will contribute its Devin autonomous AI software engineer in four areas: software and data security, modernizing legacy scientific code, expanding scientific workforce capacity, and cloud modernization. Devin Desktop and CLI are listed as FedRAMP Class D (High) Authorized, and the company has offered in-kind code security scans for national laboratory codebases.

Jul 21

Jul 21Tue
  1. Bryan CatanzaroAI score57

    Poolside releases open-weight Laguna S 2.1 for agentic coding

    AIPoolside released Laguna S 2.1, an open-weight model with 118B total parameters and 8B active per token. The author says it performs strongly on agentic coding and long-horizon tasks, and it can run on a single NVIDIA DGX Spark. Weights are on Hugging Face under the OpenMDW-1.1 license, with access also available through OpenRouter and Poolside's API.

  2. JetBrains AI BlogAI score55

    JetBrains Air adds ACP agents, local models, and Java/Kotlin code intelligence

    AIJetBrains Air now connects to ACP-compatible coding agents, including GitHub Copilot CLI, OpenCode, Pi, and Cline, through the Agent Client Protocol. The release also adds Beta Java and Kotlin navigation and diagnostics powered by the IntelliJ IDEA code engine, local model support through Ollama or LM Studio, and Docker-based agent tasks on Windows.

  3. koray kavukcuogluAI score72

    Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

    AIGoogle introduces Gemini 3.6 Flash as its workhorse model, with better coding, knowledge work, and multimodal performance while reducing token usage. It also launches Gemini 3.5 Flash-Lite, described as the fastest and most cost-effective 3.5-class model for high-throughput applications, and 3.5 Flash Cyber, a version of 3.5 Flash fine-tuned to find and fix cybersecurity vulnerabilities.

    Why it matters: The post lists three distinct models, each aimed at a different job, so readers can map which one fits coding, high-volume, or security workloads.

    Video from @koraykv's post
  4. JetBrains AI BlogAI score62

    JetBrains Context adds repository indexing to coding agents in early access

    AIJetBrains has launched JetBrains Context in early access, a repository intelligence layer that builds a semantic index so coding agents can retrieve relevant code without repeated searching. In tests on 205 SWE-bench tasks, 175 production-monorepo tasks, and 1,953 code-localization tasks, it reduced agent turns by up to 68%, latency by up to 59%, and execution cost by up to 48%. It works with Claude Code, Codex CLI, and Junie CLI at no additional cost for JetBrains AI subscribers, and it does not store source code on JetBrains Context servers.

    Why it matters: The source gives benchmark figures for turns, latency, and cost, showing how repository indexing might change agent workflows on large codebases.

Jul 20

Jul 20Mon

Jul 14

Jul 14Tue
  1. Cognition Blog (Devin, Windsurf)AI score44

    Cognition Marks One Year Since Windsurf Merger With Devin and SWE Model Gains

    AICognition says its one-year-old merger with Windsurf has produced a more capable Devin, which now manages other Devins at a mid-to-senior engineering level, and new SWE-1.7 model, described as its most capable and efficient to date. The company reports growing from 44 to 350 people and revenue run rate from $73M to $500M+ since merging the brands.

Jul 13

Jul 13Mon
  1. Cognition Blog (Devin, Windsurf)AI score62

    Fable 5 with a sidekick costs less than Opus 4.8 on FrontierCode

    AICognition found that Fable 5 led runs cost less than Opus 4.8 led runs on FrontierCode 1.1 when both used the same sidekick, $1.86 versus $2.04 per run. Fable 5 scored 60.7 against 54.6 for Opus 4.8 in those configurations, and it took fewer lead turns, delegated earlier, and rarely edited code itself. The post attributes the difference to delegation style rather than per-token price, and notes that the approach gives little benefit on short or serial debugging tasks.

    Why it matters: The source compares lead-model delegation habits on a coding benchmark, showing how a pricier model can lower total agent cost through fewer turns and better handoffs.

  2. Cognition Blog (Devin, Windsurf)AI score39

    Cognition's Devin Reaches FedRAMP High In-Process for Federal Engineering Teams

    AICognition's entire platform, including Devin Cloud, is now FedRAMP Class D (High) In-Process and listed on the FedRAMP Marketplace, extending FedRAMP High authorization beyond Devin Desktop (formerly Windsurf). Devin Desktop and CLI are already FedRAMP High Authorized for workloads with ITAR and DoW IL4, IL5, and IL6 requirements. The company says Devin Security Swarm can find and validate vulnerabilities and open remediation pull requests, and that fleets of Devins can upgrade legacy software 5-40x faster than humans alone.

Jul 12

Jul 12Sun

Jul 10

Jul 10Fri

Jul 9

Jul 9Thu
  1. Meta AI BlogAI score72

    Meta releases Muse Spark 1.1 with agent and coding gains

    AIMeta Superintelligence Labs has introduced Muse Spark 1.1, a multimodal reasoning model aimed at agentic tasks, with gains in tool use, computer use, coding, and multimodal understanding. It supports a 1 million token context window and is available in Thinking mode in the Meta AI app and on meta.ai, with developers able to access it through a public preview of the Meta Model API.

    Why it matters: The post specifies Muse Spark 1.1's agent, coding, and multimodal gains and its Meta Model API preview access, which helps developers judge its fit for their workflows.

Jul 8

Jul 8Wed
  1. Michael TruellAI score57

    Cursor and SpaceXAI release Grok 4.5, a coding-focused model

    AICursor co-founder Michael Truell announced Grok 4.5, a model trained with SpaceXAI that the post calls Opus-class, fast, and low cost. He says it is a significant step up over Composer 2.5 and has become the daily driver for many on the Cursor team. A benchmark table shows Grok 4.5 at 83.3% on Terminal-Bench 2.1 and 78.0% on SWE-Bench Multilingual, with the post saying more releases will follow.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Cognition releases SWE-1.7, a coding model trained with long-horizon RL

    AICognition launched SWE-1.7, which it says reaches frontier-level coding performance at lower cost, trained from a Kimi K2.7 base. The post describes RL methods including top-p sampling replay to preserve entropy, compressed weight deltas across multi-cluster training, and self-compaction for rollouts up to six hours. SWE-1.7 is available in Devin via Cerebras at 1000 TPS.

    Why it matters: The post details entropy preservation, multi-cluster weight sync, and self-compaction, offering concrete RL training techniques for long-horizon coding agents to compare against one's own pipeline.

  3. Cognition Blog (Devin, Windsurf)AI score47

    Cognition Tests Trustworthiness of SWE-1.7, Built on Kimi K2.7 Code

    AICognition says its SWE-1.7 model, developed from the open-source Kimi K2.7 Code base, performs as well as or better than leading U.S. frontier models on its new trustworthiness evaluation suite. The suite combines 145 politically sensitive questions, sampled in English and Chinese, with realistic coding scenarios to measure propaganda, censorship, and security behavior. Cognition says SWE-1.7 improves substantially over the base Kimi K2.7 Code model, though the company says the benchmarks are still in development.

Jul 7

Jul 7Tue
  1. Cognition Blog (Devin, Windsurf)AI score39

    FrontierCode 1.1 refines its code-quality benchmark to curb unfair internet use

    AICognition released FrontierCode 1.1, an update to its code-quality benchmark that adds a fair internet use prompt and a verifier that zeroes out runs consulting upstream fixes. The company also relaxed 75 of over 1,000 grading criteria, added scores for Sonnet 5 and updated scores for Fable 5, and dropped reporting on the Diamond subset.

Jul 6

Jul 6Mon

Jul 5

Jul 5Sun

Jul 4

Jul 4Sat

Jul 2

Jul 2Thu
  1. Cognition Blog (Devin, Windsurf)AI score38

    Cognition launches Devin Security Vulnerability Remediation Program for enterprise backlogs

    AICognition launched the Devin Security Vulnerability Remediation Program, in which its forward-deployed engineers embed with customer teams to deploy Devin to find, validate, and fix vulnerabilities. The program first works through existing scanner backlogs from tools such as Snyk, SonarQube, and Semgrep, shipping validated fixes as pull requests, then adds Devin Security Swarm for continuous discovery of logic flaws. Most engagements run about six weeks, and eligibility is limited to enterprise Devin Cloud customers meeting the program's requirements.

Jul 1

Jul 1Wed
  1. Cognition Blog (Devin, Windsurf)AI score57

    Cognition launches Devin Security Swarm to find, verify, and patch vulnerabilities

    AICognition has launched Devin Security Swarm, which uses parallel agents to find vulnerabilities across a codebase, confirms exploitability in isolated sandboxes, and opens remediation PRs. In an evaluation on 50 real-world GitHub Security Advisory vulnerabilities, Devin reached 72% recall at about $90.23 per run, compared with 68% for Claude Security at $131.87 per run. The product is available starting today, with scan profiles and incremental scans that process only changed code after the first full baseline.

  2. Mistral AI · new models on Hugging FaceAI score54

    Mistral AI releases Leanstral 1.5, an open-source Lean 4 code agent model

    AIMistral AI released Leanstral 1.5 on Hugging Face as an open-source code agent model for Lean 4 proof assistant tasks. The model uses 119B total parameters with 6.5B activated per token, a 256k context length, and accepts text and image input. The source gives setup paths through Mistral Vibe and a local vLLM server, with recommended settings of temperature 1.0 and reasoning effort set to high for complex prompts. The model is licensed under Apache 2.0.

Jun 30

Jun 30Tue
  1. Andrew NgAI score50

    Andrew Ng outlines three loops for building 0-to-1 AI products

    AIAndrew Ng describes three loops he uses to build 0-to-1 products with AI agents: an agentic coding loop, a developer feedback loop, and an external feedback loop. He says the agentic coding loop runs every few minutes, letting coding agents build, test, and iterate on software for around an hour without human intervention. The developer feedback loop operates over tens of minutes to hours, with humans steering product decisions because they hold a context advantage over AI systems.

    Image from @AndrewYNg's post