Skip to contentSkip to stories

Updated

#Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 21

Sep 21Mon
  1. ModelScopeOfficialAI score36

    Qwen Launches RecreationBench for Hybrid Computer-Use Agent Evaluation

    AIQwen introduced RecreationBench, a benchmark of 250 application-recreation tasks across Ubuntu, macOS, Windows, Android, and Web. Unlike GUI-only or terminal-only benchmarks, agents must explore a running reference app, recreate it in code, and pass programmatic tests plus VLM-based visual evaluation. The dataset is available on ModelScope.

    Image from @ModelScope2022's post

Sep 20

Sep 20Sun
  1. xAI News (Grok)OfficialAI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

    Image from @StepFun_ai's post

Sep 18

Sep 18Fri
  1. Mike KnoopXAI score30

    Mike Knoop wonders what an underscore.js equivalent for AI looks like

    AIMike Knoop asks what the underscore.js equivalent for AI would look like, noting that such programming primitives feel close. He adds that he barely reads or writes code anymore despite these emerging tools. The quoted post introduces Probably, a toy programming language built around Jev, where constructs like "feels," "match," and "while" let AI make decisions within ordinary code.

  2. Thomas DohmkeXAI score24

    Claude Code adds AGENTS.md support in version 2.1.277

    AIClaude Code version 2.1.277 now reads AGENTS.md when a folder has no CLAUDE.md, according to Anthropic engineer Thariq Shihipar's post. The behavior can be toggled in /config. Thomas Dohmke's main post jokingly says AI is finally aligned, with no further technical detail.

  3. Google for DevelopersOfficialAI score38

    Android Bench 2.0 tests AI models on multi-day engineering workflows

    AIGoogle has released Android Bench 2.0, an updated benchmark that evaluates AI models on long-horizon tasks such as building apps from scratch, migrating cross-platform codebases to Android, and making complex architectural transitions. The benchmark uses continuous completion scoring to show which tasks each model performs well on.

  4. GitHub Blog · AI & MLOfficialAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.

Sep 17

Sep 17Thu
  1. Together AI BlogOfficialAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  2. Sherwin WuXAI score44

    ChatGPT for Word launches, bringing ChatGPT and Codex into Microsoft Word

    AIOpenAI has released ChatGPT for Word, letting users access ChatGPT and Codex directly inside Microsoft Word. The post says ChatGPT for Excel and PowerPoint has been growing rapidly, and Word completes that set. The quoted ChatGPT post adds that the tool can draft from notes, rewrite paragraphs, proofread, suggest edits, and flag formatting issues.

  3. Boris ChernyXAI score45

    Claude Code adds Projects for parallel cloud coding sessions

    AIBoris Cherny says Projects in Claude Code have changed how he codes: he sends thoughts as they come, and Claude splits them into threads that the project remembers. The quoted ClaudeDevs post says Projects is rolling out on desktop and web in beta for select users, running work as parallel cloud sessions that pass context between them.

    Image from @bcherny's post
  4. Google AI StudioOfficialAI score58

    Google AI Studio open-sources Speakeasy's OpenAPI SDK generator suite

    AIGoogle AI Studio announced that Speakeasy is open sourcing its OpenAPI client generation suite under AGPLv3, following a May 2026 vendor shutdown that disrupted Google's SDK pipeline. The suite covers SDK generation for 7 languages, an agent-native CLI generator, and a documentation MCP server generator. Google says its pipeline now serves six targets with roughly one engineer maintaining it.

  5. catXAI score62

    Claude Code adds Projects that coordinate multiple parallel sessions

    AIAnthropic's Claude Code is rolling out Projects on desktop and web, in beta for select users. A project splits work into threads, runs them as parallel cloud sessions, passes context between them, and keeps running after the user leaves. The author says Claude keeps context across tasks and can give an aggregated status update on request.

  6. JetBrains AI BlogOfficialAI score44

    Building a RAG Pipeline for Semantic Code Search: A Developer Diary

    AIJetBrains describes building Air Context, a RAG pipeline that gives LLM agents semantic code search over real repositories instead of grep. The first installment covers parsing, chunking, and vectorization, arguing that fixed-size line chunks split related code and that structure-aware chunking using language grammar produces better retrieval units.

Sep 16

Sep 16Wed
  1. Amp NewsOfficialAI score50

    Amp Runners Now Serve Multiple Directories and Update Themselves

    AIAmp runners can now serve multiple directories, specified with repeated --dir flags or found automatically with --discover-dirs, which scans Git checkouts up to two levels deep by default. Runners also check for new releases about once an hour, install them, and restart into the new version once no thread is running, at most once every 12 hours, with auto-update disabled via amp.runner.autoUpdate.enabled: false.

  2. Greg BrockmanXAI score62

    Databricks rolls out Astra to all engineers, reports 60% higher coding spend

    AIDatabricks rolled out Astra to every engineer, about 3,500 people, after a pilot with around 200 users. Engineers given Astra increased coding spend by roughly 60% compared to baseline. The company reports Astra outperforms Opus 5 and Sol 5.6 on highly complex system design tasks, but sees no clear gain on medium or low complexity coding. Astra gets a separate sub-budget in Unity Gateway to encourage selective use.

  3. Perplexity DevelopersOfficialAI score34

    Perplexity's Search SDK extracts query-relevant passages from URLs for agents

    AIPerplexity says its Search SDK extracts passages relevant to a query from user-provided URLs. Agents can use those passages instead of full pages, keeping unrelated content out of the model context. The company also points to an Agent Skill for installing the Search SDK in coding agents.

  4. TinkerOfficialAI score32

    Sundial trains Inkling-Small to fix LaTeX errors in under a second

    AISundial fine-tuned Thinking Machines' Inkling-Small with RLVR on 3,978 verified TeX.StackExchange fixes, using rewards for compilation and PDF match and penalties for removed content. The trained model fixes 83.7% of LaTeX errors in under one second at $0.0013 per fix, according to the post. Sundial says it is rolling out the model in its editor, applying fixes as suggestions and rebuilding the PDF.

  5. Kilo (acq. by Anaconda)OfficialAI score40

    Kilo Mobile lets users run full AI agent loops from their phone

    AIKilo Mobile now lets users spawn Cloud Agents, start sessions on remote machines, and dictate prompts by voice from a phone. Users can also review and comment on pull requests and approve Security Agent remediations without a laptop. On iPhone, Live Activities show session status on the Lock Screen when an agent needs input.

    Image from @kilocode's post
  6. Matei ZahariaXAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

  7. Greg BrockmanXAI score22

    Codex voice coming to CarPlay for hands-free road-trip coding

    AIGreg Brockman shared a demo of Codex voice running in CarPlay, letting users build software by speaking while driving. A quoted post from Jonathan Roomer says the setup runs through a third-party iOS app that connects ChatGPT Voice to Codex and his Mac.

Sep 15

Sep 15Tue
  1. Zed BlogOfficialAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  2. xAI News (Grok)OfficialAI score43

    Grok Build Adds Memory That Saves Project Notes Between Sessions

    AIGrok Build now has memory, which records conventions, decisions, and project facts after each completed turn and reads them in later sessions. Notes are stored per project plus a global set, and the /dream command organizes them into topic files while /memory opens a read-only browser. The feature is available now and applies to new sessions.

  3. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition and AWS sign multi-year deal to deploy Devin for enterprise modernization

    AICognition and AWS have entered a multi-year Strategic Collaboration Agreement to help enterprises deploy the Devin autonomous engineer in production. Devin can be purchased through AWS Marketplace, and the companies are exploring deeper engineering integrations within customers' AWS environments. Mercedes-Benz reportedly used Devin to analyze more than 200,000 lines of COBOL, reducing an estimated eight-month modernization project to eight days.

    Why it matters: The collaboration shows how an autonomous coding agent is being packaged for enterprise legacy modernization inside existing AWS environments, with concrete customer migration figures.

  4. Kilo (acq. by Anaconda)OfficialAI score22

    Kilo App launches on Product Hunt for iOS and Android

    AIKilo announces that its Kilo App is live on Product Hunt, letting users start coding agents, check sessions, and review pull requests from iOS and Android. The company asks supporters to upvote or comment on its Product Hunt listing.

    Image from @kilocode's post

Sep 14

Sep 14Mon
  1. Factory NewsOfficialAI score40

    Factory raises $200M at $5B valuation to scale self-improving enterprise software development

    AIFactory has raised $200M at a $5B valuation from investors including Blackstone, Khosla Ventures, and Sequoia Capital, bringing its total funding to over $400 million. The company says it will use the capital to accelerate research, product, and global go-to-market efforts. Factory says hundreds of thousands of developers use its platform, with customers including Nvidia, Blackstone, and T-Mobile.

Sep 13

Sep 13Sun
  1. Fireworks AI BlogOfficialAI score52

    Fireworks adds DeepSeek-V4.1-Flash, matching GPT-6 Astra coding accuracy at 1/15th the cost

    AIFireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.

Sep 11

Sep 11Fri
  1. Augment Code BlogOfficialAI score80

    Augment Code details how its software factory raised output per developer 4.5×

    AIAugment Code reports that size-adjusted output per active developer rose from 12.3 to 55.7 between November 2025 and July 2026, while median time to merge fell from 11.2 to 3.1 hours. The post says the company added specialized agents wherever work was piling up, across planning, review, verification, feedback, and incident response, and kept engineers responsible for product decisions, architecture, and production risk.

    Why it matters: The post pairs internal productivity and quality metrics with the order in which agents were added, showing how review and verification bottlenecks shaped a software delivery pipeline.

  2. Baseten BlogOfficialAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

  3. Andrew NgXAI score22

    AI engineers now shape product direction, not just implement specs

    AIAndrew Ng argues that skilled AI engineers increasingly drive the build loop and make product decisions rather than merely implementing specifications from product managers and designers. He identifies four key skills for shaping the build: driving the build loop, making product decisions, communicating and leading, and high-agency ownership.

  4. Cognition Blog (Devin, Windsurf)OfficialAI score51

    Cognition introduces Fusion in Devin Desktop and CLI for lower-cost coding

    AICognition is making Fusion available in Devin Desktop and CLI, a harness where a frontier lead model plans and reviews while a cheaper sidekick executes. Across listed coding benchmarks, Cognition reports Fusion cuts cost per task by about 11% to 46% versus the lead model alone, while the sidekick does the implementation work. The post recommends pairing Fable 5.1 with SWE-2, and argues price per task matters more than price per token.

  5. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score72

    Shanghai AI Lab releases Atria Dawn Preview, an agentic model built on GLM-5.2

    AIShanghai Artificial Intelligence Laboratory has released Atria Dawn Preview, an agentic model built on the 744B-parameter MoE GLM-5.2 foundation model, with a 256K context window. The release page reports benchmark results across search, coding, tool use, productivity, and cybersecurity, and describes text-only setup for Codex and Claude Code.

    Why it matters: The release page gives a full benchmark table against named rivals and setup steps for Codex and Claude Code, useful for anyone evaluating agentic models.

Sep 10

Sep 10Thu