Skip to contentSkip to stories

Updated

Coding

Showing low-relevance items too. Hide low-relevance items

Oct 9

TodayOct 9Fri
  1. AI EraNewsAI score36

    Anthropic says AI should test and fix its own code

    AIAnthropic recommends that AI coding tools like Claude Code test and revise their own output, rather than leaving debugging to developers. The source describes a developer who built a small app with Claude Code and added an AI customer-service bot, but the text provided is only the opening scenario.

  2. SGLangOfficialAI score37

    Miles runs full RL loop on Rubin GPUs with SGLang and Megatron

    AIMiles runs the full reinforcement learning loop on Nvidia Rubin, using SGLang for rollout, Megatron for training, and one container image. On a single 4-GPU tray, Qwen3-30B-A3B's GSM8K reward rises from about 45% to about 95% over 50 rollouts, matching the GB300 curve. The post also reports DeepSeek-V4-Flash end-to-end rollout and training, and Qwen3.5-35B-A3B agentic RL with 64 concurrent mini-SWE-agent sandboxes on SWE-bench Verified, where reward holds near 0.6 and median response length falls about 30%.

  3. GitHub Copilot ChangelogOfficialAI score36

    GitHub Copilot adds local sandboxing and separate accounts in weekly releases

    AIGitHub makes local sandboxing generally available in Copilot CLI, the Copilot app, and VS Code sessions using Agent Host, limiting agents' access to files, networks, and credentials at no extra cost. The Copilot app now lets users sign in with separate GitHub accounts for the Copilot license and for repositories. Copilot CLI's /model command lists local models from a running Ollama instance alongside cloud models, and VS Code 1.141 adds a side-by-side agent session grid and worktree cleanup.

  4. Replit ⠕OfficialAI score34

    Replit previews Windows desktop app, cross-project chat, and TikTok Ads MCP

    AIReplit says its Desktop app for Windows is in private preview with Microsoft and NVIDIA, building and running apps in isolated sandboxes on a user's PC. The company also lets users work across projects from one chat, including finding projects, reading their files, and sending them tasks. A new TikTok Ads MCP lets users create, launch, and track TikTok ads from Replit.

    Video from @Replit's post
  5. elvisXAI score44

    Tinker cuts long-context token prices, making agent RL rollouts cheaper

    AITinker has cut prices up to 70% on long-context prefill and sampling, which now cost the same as short context. The cut lowers the cost of agentic RL rollouts, which spend most of their tokens re-reading growing context, and of evaluating trained models on long inputs. Tinker also added GLM-5.3-Flash and DeepSeek-v4.1-Flash for cost-efficient long-context work.

  6. 🚨 AI News | TestingCatalogXAI score45

    Google's Gemini 4 Argon model spotted in Antigravity

    AIHidden references to a Gemini 4 Argon model with low, medium, and high reasoning efforts have appeared recently in Antigravity, according to testingcatalog. Business Insider reported that Google employees are testing an internal Gemini 4 checkpoint called "Carbon," which performs at the Opus 5.5 level on coding tasks.

    Image from @testingcatalog's post
  7. Rohan PaulXAI score57

    Microsoft paper finds coding agents struggle more with code understanding than editing

    AIMicrosoft researchers introduce CABRA, a framework that generates synthetic coding tasks with one difficulty dimension varied at a time. Across 6,840 tasks, plain LLMs degraded as tasks grew, while agents stayed near-perfect by offloading work to tools such as grep. On SWE-bench Verified, counts of reading and analysis calls correlated with agent failures at -0.200, versus -0.159 for lines edited.

    Image from @rohanpaul_ai's post
  8. Prime Intellect BlogOfficialAI score65

    Prime Agent is rewritten in Rust by a swarm of agents

    AIPrime Intellect says it rewrote its Prime Agent coding tool in Rust, using a swarm of more than 2,000 agents over two weeks. The company reports cold start to typing about 13 times faster than the TypeScript version, and memory use over 80% lower after startup. Prime Agent remains open source and adds native Windows support in beta and Homebrew installation.

    Why it matters: The post shows how a multi-agent swarm rewrote a coding agent with parity checks, giving a concrete case of agent-driven software engineering with measured results.

  9. Prime IntellectOfficialAI score46

    Prime Agent swarm of 2,000+ agents rewrites itself in Rust

    AIPrime Intellect says its Prime Agent orchestrated over 2,000 agents over two weeks to rewrite the agent in Rust. The run used more than 10,000 sandboxes, over 200B GLM-5.3 tokens, and 16,000 agent-to-agent messages. The company says the rewritten agent reaches usable input about 13 times faster and uses 83% less startup memory.

    Video from @PrimeIntellect's post
  10. Prime IntellectOfficialAI score22

    Prime Intellect uses objective parity checks to port TypeScript features safely

    AIPrime Intellect says it preserved its TypeScript version's features and behavior by giving agents objective parity checks. The checks diff terminal frames, compare session transcripts and model requests, check daemon protocol messages, and audit every feature. Agents could see where behavior diverged and fix it before changes merged.

    Video from @PrimeIntellect's post
  11. Prime IntellectOfficialAI score20

    Root agent rewrites code through planner, implementer, reviewer, and verifier pipeline

    AIA root agent splits a rewrite into dependent tasks, with each task run through a Planner, Implementer, Reviewer, and Verifier state machine. Each implementation must pass independent review and verification in a fresh Prime Sandbox before merging, and failed checks send the task back to the implementer. Tasks can run in parallel without skipping these checks.

    Video from @PrimeIntellect's post
  12. Prime IntellectOfficialAI score42

    Prime Intellect plans reusable agent state machines in Prime Agent

    AIPrime Intellect says it is turning the workflow behind a recent rewrite into reusable state machines in Prime Agent, letting users run their own agent teams through implementation, review, and verification. The company also says it is accelerating work on capabilities and evals and connecting Prime Agent with cloud agent swarms and hosted training for autonomous research. This release adds native Windows support in beta and Homebrew installation.

  13. Sherwin WuXAI score38

    OpenAI's Codex now predicts users' next messages in beta

    AISherwin Wu, who owns OpenAI's account context here, says he has been tab-accepting about 40-50% of Codex's next-message suggestions after a week of use. OpenAI Devs says composer predictions, which suggest a user's next message from the conversation and their style, are in beta for Pro users.

  14. Ars Technica · AINewsAI score57

    AI coding agents generate more code but not more software, study finds

    AIA study by Harvard researchers Fiona Chen and James Stratton, using Jellyfish engineering data from over 700 software firms, finds little evidence that AI coding tools increase software output or reduce employment. The authors report that efficiency gained during coding is absorbed by downstream constraints, mainly longer code review, more pull request revisions, and more reviewer comments.

  15. Claude Code · GitHub ReleasesOfficialAI score33

    Claude Code v2.1.296 adds gateway policy controls and fixes hook and permission bugs

    AIAnthropic releases Claude Code v2.1.296, which adds a code key to the Claude apps gateway's managed.policies[] and an allow_large option to the Read tool for reading large text files in one call. The release also adds autoCompactWindow for subagents and CLAUDE_CODE_WORKFLOW_SUBAGENT_MODEL, and fixes many bugs in hooks, MCP servers, permission checks and self-hosted runners.

  16. GitHub Copilot ChangelogOfficialAI score36

    Copilot code review adds organization billing and review request controls

    AIGitHub adds two Copilot code review admin controls. Organization owners can bill code reviews from members with a Copilot license to the owning organization instead of member quotas, which requires AI Credits paid usage and allows an optional budget. Owners and repository admins can also restrict review requests to users whose Copilot license comes from their organization or enterprise.

  17. RadixArkOfficialAI score22

    RadixArk praises Proximal for training coding agents with Miles

    AIRadixArk says Proximal is using Miles to train coding agents and calls it a flexible, scalable foundation for teams running their own training workloads. Proximal says its training framework is built on Miles, with runs on Modal's on-demand GPU clusters and serverless GPUs for inference. Its sandboxing infrastructure runs on Kubernetes and can handle millions of concurrent rollouts.

  18. ChatGPTOfficialAI score28

    ChatGPT's dot can now delegate work to Codex threads

    AIOpenAI says its dot can start work in Codex and follow up on existing threads, drawing on ChatGPT conversations, Codex threads, and automations. The dot also decides whether to continue a thread or start a fresh one, and can review and edit Scheduled Tasks in ChatGPT Work.

    Video from @ChatGPT's post
  19. ZDNet · AINewsAI score38

    Linus Torvalds says AI helps him with tasks outside his expertise

    AILinus Torvalds says he uses AI to do things he is bad at, such as building a user interface for a guitar pedal project he wrote in C. He says AI is a wonderful tool for beginners, but warns that maintainers are stressed by AI-generated Linux kernel patches and bug reports. Torvalds says AI review tools like Sashiko are now appearing on the Linux Kernel Mailing List, with some subsystem maintainers expecting patches to be reviewed before acceptance.

  20. Tibor BlahoXAI score38

    Codex adds composer predictions for Pro users in beta

    AIOpenAI's developer account announces composer predictions in Codex, which suggest a user's next message based on the conversation and how the user writes. During the beta, the feature is included at no additional cost for eligible Pro users.

  21. Lydia Hallie ✨XAI score36

    Claude Code auto-compact summarizes conversations, not the last 1M tokens

    AIAnthropic's Lydia Hallie clarifies that Claude Code's auto-compact replaces the whole conversation with a short summary. On 1M-context models it triggers around 967K tokens, and each message before that point re-reads the full conversation, mostly from cache. Running /autocompact 400k makes compaction trigger at 400K instead.

    Video from @lydiahallie's post
  22. OpenAI DevelopersOfficialAI score33

    Codex adds composer predictions for Pro users in beta

    AIOpenAI says composer predictions in Codex is now in beta for Pro users. The feature suggests a user's next message based on their conversation and how they phrase requests. OpenAI calls it one of the most loved features its team has tested internally.

    Video from @OpenAIDevs's post
  23. ClaudeDevsOfficialAI score60

    Claude Code Projects opens to all Pro and Max users on the waitlist

    AIAnthropic's ClaudeDevs account says it has let in every Pro and Max user from the Claude Code Projects waitlist. The post links a 4-minute walkthrough video for new users getting started with the feature.

    Why it matters: The post shows Claude Code Projects access opening to Pro and Max users from the waitlist, with a walkthrough for new users getting started.

    Video from @ClaudeDevs's post
  24. Codex · GitHub ReleasesOfficialAI score13

    Codex 0.162.1 fixes TUI crash and startup failures

    AIOpenAI releases Codex 0.162.1, a bug-fix update for the Codex CLI. It fixes a TUI crash when asynchronous questions contain multiple lines, preserving line breaks and complete hyperlink destinations. It also fixes startup failures caused by mismatches between a running background server's feature settings and CLI defaults.

  25. elvisXAI score62

    StepFun's Step 5 Preview targets long coding agent runs

    AIElvis Saravia says he has tested StepFun's Step 5 Preview as a coding agent since early access and found that it checks its own work and stops when tasks are done. The post says the model is built for engineering tasks such as bug fixing, multi-file features, and refactoring, plus frontend generation and financial report output.

    Image from @omarsar0's post
  26. Alex BraginXAI score29

    Solo CTO rebuilds AQUA wallet in React Native with Claude Code

    AIJAN3 CTO Jan Ceuleers says he rebuilt the AQUA wallet from Flutter into React Native, using Claude Code for most of the coding. He started by binding GDK to a React Native app and showing a wallet balance in about half an hour, then ported the Aqua UI component library from Flutter over roughly a year of prior work.

  27. dexXAI score40

    Dex Horthy says small tasks should skip heavy planning workflows

    AIDex Horthy says the share of tasks that can be one-shot without strict process has grown, but alignment, grilling, and planning workflows still matter. He argues that heavy planning on small tasks makes developers feel slower, and predicts tools will add escape hatches so humans or models can decide to ship directly. He adds that as model capabilities improve, the "smart zone" has grown to roughly 200k–400k tokens, and HumanLayer is prototyping research-to-implement and research-to-short-design-to-implement workflows.