Skip to contentSkip to stories

Updated

Coding

Showing low-relevance items too. Hide low-relevance items

Oct 9

TodayOct 9Fri
  1. Rohan PaulXAI score57

    Microsoft paper finds coding agents struggle more with code understanding than editing

    AIMicrosoft researchers introduce CABRA, a framework that generates synthetic coding tasks with one difficulty dimension varied at a time. Across 6,840 tasks, plain LLMs degraded as tasks grew, while agents stayed near-perfect by offloading work to tools such as grep. On SWE-bench Verified, counts of reading and analysis calls correlated with agent failures at -0.200, versus -0.159 for lines edited.

    Image from @rohanpaul_ai's post
  2. Prime Intellect BlogOfficialAI score65

    Prime Agent is rewritten in Rust by a swarm of agents

    AIPrime Intellect says it rewrote its Prime Agent coding tool in Rust, using a swarm of more than 2,000 agents over two weeks. The company reports cold start to typing about 13 times faster than the TypeScript version, and memory use over 80% lower after startup. Prime Agent remains open source and adds native Windows support in beta and Homebrew installation.

    Why it matters: The post shows how a multi-agent swarm rewrote a coding agent with parity checks, giving a concrete case of agent-driven software engineering with measured results.

  3. Prime IntellectOfficialAI score46

    Prime Agent swarm of 2,000+ agents rewrites itself in Rust

    AIPrime Intellect says its Prime Agent orchestrated over 2,000 agents over two weeks to rewrite the agent in Rust. The run used more than 10,000 sandboxes, over 200B GLM-5.3 tokens, and 16,000 agent-to-agent messages. The company says the rewritten agent reaches usable input about 13 times faster and uses 83% less startup memory.

    Video from @PrimeIntellect's post
  4. Prime IntellectOfficialAI score22

    Prime Intellect uses objective parity checks to port TypeScript features safely

    AIPrime Intellect says it preserved its TypeScript version's features and behavior by giving agents objective parity checks. The checks diff terminal frames, compare session transcripts and model requests, check daemon protocol messages, and audit every feature. Agents could see where behavior diverged and fix it before changes merged.

    Video from @PrimeIntellect's post
  5. Prime IntellectOfficialAI score20

    Root agent rewrites code through planner, implementer, reviewer, and verifier pipeline

    AIA root agent splits a rewrite into dependent tasks, with each task run through a Planner, Implementer, Reviewer, and Verifier state machine. Each implementation must pass independent review and verification in a fresh Prime Sandbox before merging, and failed checks send the task back to the implementer. Tasks can run in parallel without skipping these checks.

    Video from @PrimeIntellect's post
  6. Prime IntellectOfficialAI score42

    Prime Intellect plans reusable agent state machines in Prime Agent

    AIPrime Intellect says it is turning the workflow behind a recent rewrite into reusable state machines in Prime Agent, letting users run their own agent teams through implementation, review, and verification. The company also says it is accelerating work on capabilities and evals and connecting Prime Agent with cloud agent swarms and hosted training for autonomous research. This release adds native Windows support in beta and Homebrew installation.

  7. Sherwin WuXAI score38

    OpenAI's Codex now predicts users' next messages in beta

    AISherwin Wu, who owns OpenAI's account context here, says he has been tab-accepting about 40-50% of Codex's next-message suggestions after a week of use. OpenAI Devs says composer predictions, which suggest a user's next message from the conversation and their style, are in beta for Pro users.

  8. Ars Technica · AINewsAI score57

    AI coding agents generate more code but not more software, study finds

    AIA study by Harvard researchers Fiona Chen and James Stratton, using Jellyfish engineering data from over 700 software firms, finds little evidence that AI coding tools increase software output or reduce employment. The authors report that efficiency gained during coding is absorbed by downstream constraints, mainly longer code review, more pull request revisions, and more reviewer comments.

  9. Claude Code · GitHub ReleasesOfficialAI score33

    Claude Code v2.1.296 adds gateway policy controls and fixes hook and permission bugs

    AIAnthropic releases Claude Code v2.1.296, which adds a code key to the Claude apps gateway's managed.policies[] and an allow_large option to the Read tool for reading large text files in one call. The release also adds autoCompactWindow for subagents and CLAUDE_CODE_WORKFLOW_SUBAGENT_MODEL, and fixes many bugs in hooks, MCP servers, permission checks and self-hosted runners.

  10. GitHub Copilot ChangelogOfficialAI score36

    Copilot code review adds organization billing and review request controls

    AIGitHub adds two Copilot code review admin controls. Organization owners can bill code reviews from members with a Copilot license to the owning organization instead of member quotas, which requires AI Credits paid usage and allows an optional budget. Owners and repository admins can also restrict review requests to users whose Copilot license comes from their organization or enterprise.

  11. RadixArkOfficialAI score22

    RadixArk praises Proximal for training coding agents with Miles

    AIRadixArk says Proximal is using Miles to train coding agents and calls it a flexible, scalable foundation for teams running their own training workloads. Proximal says its training framework is built on Miles, with runs on Modal's on-demand GPU clusters and serverless GPUs for inference. Its sandboxing infrastructure runs on Kubernetes and can handle millions of concurrent rollouts.

  12. ChatGPTOfficialAI score28

    ChatGPT's dot can now delegate work to Codex threads

    AIOpenAI says its dot can start work in Codex and follow up on existing threads, drawing on ChatGPT conversations, Codex threads, and automations. The dot also decides whether to continue a thread or start a fresh one, and can review and edit Scheduled Tasks in ChatGPT Work.

    Video from @ChatGPT's post
  13. ZDNet · AINewsAI score38

    Linus Torvalds says AI helps him with tasks outside his expertise

    AILinus Torvalds says he uses AI to do things he is bad at, such as building a user interface for a guitar pedal project he wrote in C. He says AI is a wonderful tool for beginners, but warns that maintainers are stressed by AI-generated Linux kernel patches and bug reports. Torvalds says AI review tools like Sashiko are now appearing on the Linux Kernel Mailing List, with some subsystem maintainers expecting patches to be reviewed before acceptance.

  14. Tibor BlahoXAI score38

    Codex adds composer predictions for Pro users in beta

    AIOpenAI's developer account announces composer predictions in Codex, which suggest a user's next message based on the conversation and how the user writes. During the beta, the feature is included at no additional cost for eligible Pro users.

  15. Lydia Hallie ✨XAI score36

    Claude Code auto-compact summarizes conversations, not the last 1M tokens

    AIAnthropic's Lydia Hallie clarifies that Claude Code's auto-compact replaces the whole conversation with a short summary. On 1M-context models it triggers around 967K tokens, and each message before that point re-reads the full conversation, mostly from cache. Running /autocompact 400k makes compaction trigger at 400K instead.

    Video from @lydiahallie's post
  16. OpenAI DevelopersOfficialAI score33

    Codex adds composer predictions for Pro users in beta

    AIOpenAI says composer predictions in Codex is now in beta for Pro users. The feature suggests a user's next message based on their conversation and how they phrase requests. OpenAI calls it one of the most loved features its team has tested internally.

    Video from @OpenAIDevs's post
  17. ClaudeDevsOfficialAI score60

    Claude Code Projects opens to all Pro and Max users on the waitlist

    AIAnthropic's ClaudeDevs account says it has let in every Pro and Max user from the Claude Code Projects waitlist. The post links a 4-minute walkthrough video for new users getting started with the feature.

    Why it matters: The post shows Claude Code Projects access opening to Pro and Max users from the waitlist, with a walkthrough for new users getting started.

    Video from @ClaudeDevs's post
  18. Codex · GitHub ReleasesOfficialAI score13

    Codex 0.162.1 fixes TUI crash and startup failures

    AIOpenAI releases Codex 0.162.1, a bug-fix update for the Codex CLI. It fixes a TUI crash when asynchronous questions contain multiple lines, preserving line breaks and complete hyperlink destinations. It also fixes startup failures caused by mismatches between a running background server's feature settings and CLI defaults.

  19. elvisXAI score62

    StepFun's Step 5 Preview targets long coding agent runs

    AIElvis Saravia says he has tested StepFun's Step 5 Preview as a coding agent since early access and found that it checks its own work and stops when tasks are done. The post says the model is built for engineering tasks such as bug fixing, multi-file features, and refactoring, plus frontend generation and financial report output.

    Image from @omarsar0's post
  20. dexXAI score40

    Dex Horthy says small tasks should skip heavy planning workflows

    AIDex Horthy says the share of tasks that can be one-shot without strict process has grown, but alignment, grilling, and planning workflows still matter. He argues that heavy planning on small tasks makes developers feel slower, and predicts tools will add escape hatches so humans or models can decide to ship directly. He adds that as model capabilities improve, the "smart zone" has grown to roughly 200k–400k tokens, and HumanLayer is prototyping research-to-implement and research-to-short-design-to-implement workflows.

  21. Tessl BlogOfficialAI score36

    Tessl's agentic code review splits PR checks into standards, lenses, and memory

    AITessl Blog describes an agentic code review workflow built for teams whose coding agents produce pull requests faster than humans can review them. The workflow runs review against a written standard in the repository, applies four parallel perspectives covering correctness, security and privacy, scale and resilience, and maintainability, then records each finding, verdict, and response. Tessl Code Review, which the post says is free to start, runs these perspectives as skills, and the team's memory of past decisions is fed back into the standard.

  22. Kilo (acq. by Anaconda)OfficialAI score60

    StepFun's Step 5 Preview is free in Kilo for one week

    AIKilo says StepFun has announced Step 5 Preview, which is free to use in Kilo for one week. The post lists 600B total parameters with 27B active per token, a 1M-token context window with vision, and highlights strong coding and finance performance at lower cost.

    Image from @kilocode's post
  23. Simon WillisonBlogAI score27

    Simon Willison builds a new blog feature largely by voice with Codex

    AISimon Willison says he built a Newsletters index for his blog almost entirely by voice, using the ChatGPT desktop app's Codex voice mode while cooking dinner. The feature imports weekly Substack posts via RSS and undocumented API, monthly newsletters from a GitHub archive repository, and a private sponsors-only newsletter. He says he switched back to typing for review and fixes before deploying the pull request.

  24. Gergely OroszXAI score28

    CTO says new grads aren't AI-native, lack AI coding tool experience

    AIA CTO hiring new graduates at a larger company reports they are generally unfamiliar with AI coding tools and have little hands-on use of them. Many of those who did internships worked at traditional companies that also did not use these tools, so they are more fluent in pre-AI software development methods than the "AI-native" label suggests.

  25. QbitAINewsAI score67

    TRAE merges Code and Work into one platform with Agent and IDE modes

    AITRAE has merged its TraeCode and TraeWork products into a unified new TRAE with an Agent mode and an IDE mode. In hands-on tests, multiple agents handled planning, design, coding, testing, and fixes within one project, with outputs saved in a shared 'My Artifacts' area. The tests also found that agents working in parallel produced conflicting specifications, so someone had to coordinate them.

  26. QbitAINewsAI score38

    Lenovo's TianxiCode Agent Tops SWE-bench-Live Lite Leaderboard at 71%

    AILenovo's TianxiCode, paired with DeepSeek-v4.1-Flash, ranked first on the SWE-bench-Live Lite leaderboard with a 71% issue resolution rate and passed official Verified review. The framework combines multi-hop retrieval, autonomous planning with multi-turn tool calling, and test-driven self-correction, and will be applied to Lenovo AI hardware products.