Skip to contentSkip to stories

Updated

#Coding

Oct 2

Oct 2Fri
  1. Claude Code · GitHub ReleasesAI score38

    Claude Code v2.1.288 is released with fixes and new controls

    AIAnthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh api for cloud sessions without the GitHub CLI, and --max-findings for /code-review. The release also fixes many issues, including mid-response API timeouts, resume and compaction bugs, and auto mode denials and model switching on Bedrock and Mantle.

  2. GitHub Copilot ChangelogAI score34

    Copilot code review gains API access and Balanced default effort level

    AIGitHub Copilot code review can now be requested through the REST and GraphQL APIs, with an optional review effort level set per request. Balanced became the default review effort level for new and existing repositories and organizations as of September 28, 2026, while users who explicitly selected Lite keep that setting. The changes are generally available to Copilot Pro, Pro+, Max, Business, and Enterprise plans.

  3. GitHub Blog · AI & MLAI score23

    Three Skills Developers Need as AI Changes Their Work

    AIAI is changing developer work, and the article recommends three skills: directing AI agents, reviewing AI output instead of trusting the first answer, and using saved time for judgment-heavy problems such as customer needs and tradeoffs. It cites GitHub Copilot's built-in Rubber Duck agent, which uses a second model to critique plans, code, and tests. The author argues that developers remain responsible for outcomes while AI handles more implementation.

  4. GitHub Copilot ChangelogAI score53

    GitHub Copilot adds new models, dynamic workflows, and desktop app automation

    AIGitHub Copilot's weekly release adds Claude Sonnet 5.5 and GPT-6.1 Sol for specified plan tiers, plus HydraFusion, a research preview that lets Copilot select and coordinate models for a task. It also introduces dynamic workflows in public preview, which let users save and reuse multi-step processes, and computer use in public preview on macOS and Windows for automating desktop apps.

Oct 1

Oct 1Thu
  1. NVIDIA BlogAI score62

    NVIDIA Blackwell GPUs power OpenAI's GPT-6 Astra Ultrafast mode in API

    AIGPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is now available in the OpenAI API and to eligible ChatGPT Work and Codex users. The source says Ultrafast offers up to 8x faster token generation than Astra Standard mode, which can shorten coding agents' response times between tool calls. OpenAI also says it uses its own models to keep optimizing inference software on NVIDIA GPUs after deployment.

    Why it matters: The source ties a specific speed claim to coding agents' edit-test-debug loops, showing where faster token generation changes developer workflows.

  2. Comfy BlogAI score47

    852話 Hakoniwa Creates YUI Short Film Using Comfy Agent

    AIArtist 852話 Hakoniwa made YUI, the first animated short film created with Comfy Agent, in three days of focused work for about 200,000 yen excluding labor. She used the Agent to regenerate shots, compare video models such as Wan, MiniMax H3 and Seedance, and check outputs against her character sheet. The main film was made mostly with Seedance 2.5, and the making-of video with MiniMax H3.

  3. JetBrains AI BlogAI score75

    JetBrains Air enters early access as an agent system inside its IDEs

    AIJetBrains has opened the Early Access Program for Air, an agentic development experience available as a plugin on JetBrains Marketplace or in the 2026.3 EAP builds of its IDEs. Air works with existing agents such as Codex, GitHub Copilot, Junie, and Cursor, and it ships with no agents installed. Free Junie Lite runs are offered, while cloud runs require a JetBrains AI subscription.

    Why it matters: The post explains how Air brings existing agents into the IDE, showing a concrete workflow for managing parallel agent sessions alongside code review tools.

  4. Anthropic NewsroomAI score38

    Barclays expands Claude across operations, targeting 50% developer adoption by end-2026

    AIBarclays is expanding its collaboration with Anthropic to roll Claude out across its global operations, with Claude Code expected to reach 50% of its developer population by the end of 2026. Its Colleague Knowledge Assistant, powered by Claude through retrieval-augmented generation, has been used by more than 16,000 colleagues and handled over one million searches. In Global Markets, Claude models classify and route roughly 120,000 client emails daily.

  5. Manus BlogAI score47

    Manus 2.0 Adds Game Dev for Building Multiplayer Games Without Coding

    AIManus has launched Game Dev in Manus 2.0, a feature that lets users with no coding experience build games with a real-time tweak panel, asset management, and multiplayer servers. The tweak panel lets users adjust settings such as speed, gravity, damage, and spawn rate while playing. Manus also handles much of the multiplayer infrastructure, including server deployment and networking, so games can be shared and played with friends.

Sep 30

Sep 30Wed
  1. Apple Machine Learning ResearchAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  2. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  3. Google · Gemini appAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

  4. Cloudflare Blog · AIAI score72

    Cloudflare launches Auto Router in AI Gateway to cut AI token spend

    AICloudflare has released Auto Router in public beta through AI Gateway, where setting the model to cloudflare/auto routes each request to a model judged capable enough for the task. Internal tests showed up to 30% cost savings against frontier models, and on a 97-task internal benchmark cloudflare/auto scored 86.6% at $0.0084 per success versus 96.6% at $0.0210 for Claude Opus 5.5. The router is free during beta.

    Why it matters: The source gives a benchmark table of success rates and costs per trial, showing how routing trades quality against price for a gateway deployment.

Sep 29

Sep 29Tue
  1. Factory NewsAI score42

    Factory Launches Generally Available Automations to Run Recurring Engineering Workflows

    AIFactory's Automations, now generally available, let users describe a recurring workflow, set a schedule or event trigger, and have its Droid run it, with the model chosen per task. Templates cover ticket-to-PR, code review, security audits, PR babysitting, and morning Slack briefs. Among enterprise organizations using Automations in the past 30 days, 52% used automated code review, 48% used security review, and 35% used AutoWiki.

  2. Baseten BlogAI score38

    Baseten Partners With OpenAI to Offer Open Models to OpenAI Customers

    AIBaseten announced a partnership with OpenAI that makes open models powered by Baseten available to OpenAI customers for multi-model agentic coding. The company says organizations can route each task to the best-fit open or closed model, with Codex and GPT models among the options, and that Baseten's day-zero access to new open models lets teams evaluate them quickly. Baseten also cites US-based infrastructure with zero data retention for all prompts and capacity across more than 90 clusters in 20+ clouds.

  3. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    AIBAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    Why it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

  4. Replit BlogAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.

Sep 28

Sep 28Mon
  1. Google · Gemini appAI score38

    See what 4 builders are making with Gemini 3.8 Flash

    AIGoogle says Gemini 3.8 Flash, its most intelligent workhorse model, improves on 3.7 Flash in software engineering, agentic tasks, and multistep reasoning by running extra reasoning steps and calling tools iteratively. The post highlights four community builds, including a model rocket simulation, an animated ink-painting effect, a 3D dinosaur skeleton, and an interactive automatic transmission simulation. Developers can try the model through Google Antigravity and Google AI Studio.

Sep 27

Sep 27Sun
  1. Amp NewsAI score67

    Amp switches its default medium mode to Claude Opus 5.5

    AIAmp now uses Claude Opus 5.5 for its medium mode by default, replacing GPT-5.6 Sol, while ChatGPT subscribers can keep medium pinned to GPT-5.6 Sol. In Amp's internal evals, Opus 5.5 solved 65% of tasks versus 61% for GPT-5.6 Sol and 56% for Opus 5, at lower cost, and it runs at high reasoning effort because xhigh and max cost more without scoring better.

    Why it matters: The source reports internal eval scores, cost comparisons, and usage guidance for choosing reasoning effort, helping developers decide which model and setting to run.

Sep 25

Sep 25Fri
  1. GitHub Blog · AI & MLAI score33

    How to build custom workflows with canvases in the GitHub Copilot app

    AICanvases in the GitHub Copilot app are customizable interfaces that you and the agent share, such as kanban boards, dashboards, or checklists. You create one by running /create-canvas and describing the workflow, what you can do in the interface, and what the agent can do. Changes made by either you or the agent appear immediately in the shared canvas, and completed canvases can be saved as reusable extensions.

Sep 24

Sep 24Thu
  1. GitHub Blog · AI & MLAI score46

    GitHub Copilot app's canvases argue chat is the wrong AI interface

    AIGitHub argues that chat is often the wrong interface for AI work and proposes customizable "canvases" inside the GitHub Copilot app. Canvases are full-stack applications running without browser chrome that can communicate bi-directionally with the Copilot agent and execute code locally. The post cites examples including a Connect 4 game, a Winget package manager UI, and a SQLite database interface.

  2. GitHub Blog · AI & MLAI score66

    GitHub Security Lab shows an LLM agent running AI-driven fuzzing for C/C++ projects

    AIGitHub Security Lab describes the Fuzzing Taskflow, an LLM agent pipeline that identifies entrypoints, writes harnesses, runs AFL++, reads coverage reports, and triages crashes for C/C++ repositories. The agent makes decisions while MCP tools handle execution, and state is stored in a SQLite database. The post also warns that the taskflow runs AFL and build commands directly on the host, so it should be used only in disposable environments without elevated privileges.

    Why it matters: The post explains how an LLM agent automates fuzzing steps like harness writing, coverage gap chasing, and crash triage, with a runnable workflow and design tradeoffs.

Sep 23

Sep 23Wed
  1. Amp NewsAI score42

    Amp Lets Teams Share a Runner Across Their Workspace

    AIAmp users can now share a runner with their workspace by starting it with --share, letting everyone spawn threads on that machine from ampcode.com. Shared runners appear under Shared Runners in the picker, and --amp-env gives them workspace and project Secrets & Env Vars but never personal ones. Amp warns that collaborators run code as the owner with their files and credentials, so sharing should be limited to trusted people, and workspace admins can disable runner sharing in Member Settings.

  2. Amp NewsAI score34

    Amp's macOS app now runs threads on your Mac without a terminal

    AIThe Amp macOS app now starts a runner automatically, so threads can run on your Mac without keeping amp --no-tui open in a terminal. Users add folders or projects under Runner in App Settings, then select "This Mac" when starting threads from ampcode.com, a phone, or Puck. A Keep This Mac Awake option prevents sleep while the runner is on and the Mac is plugged in, though the screen still turns off and locks.

Sep 22

Sep 22Tue

Sep 21

Sep 21Mon
  1. Amp NewsAI score36

    Amp Runners Add Git Worktree Creation and Secret Injection

    AIAmp runners can now create Git worktrees from the directory picker, with each new worktree placed as a sibling folder on a new branch off the current HEAD while uncommitted changes stay in the original checkout. Runners can also opt in with --amp-env to inject Secrets & Env Vars configured on ampcode.com into thread shell commands, MCP servers, and plugins, with changes applied to the next thread without a restart.

  2. Xiaomi MiMo · new models on Hugging FaceAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.

Sep 20

Sep 20Sun
  1. xAI News (Grok)AI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

Sep 18

Sep 18Fri
  1. GitHub Blog · AI & MLAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.

Sep 17

Sep 17Thu
  1. Together AI BlogAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

Sep 16

Sep 16Wed
  1. Amp NewsAI score50

    Amp Runners Now Serve Multiple Directories and Update Themselves

    AIAmp runners can now serve multiple directories, specified with repeated --dir flags or found automatically with --discover-dirs, which scans Git checkouts up to two levels deep by default. Runners also check for new releases about once an hour, install them, and restart into the new version once no thread is running, at most once every 12 hours, with auto-update disabled via amp.runner.autoUpdate.enabled: false.

Sep 15

Sep 15Tue
  1. Zed BlogAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  2. Cognition Blog (Devin, Windsurf)AI score60

    Cognition and AWS sign multi-year deal to deploy Devin for enterprise modernization

    AICognition and AWS have entered a multi-year Strategic Collaboration Agreement to help enterprises deploy the Devin autonomous engineer in production. Devin can be purchased through AWS Marketplace, and the companies are exploring deeper engineering integrations within customers' AWS environments. Mercedes-Benz reportedly used Devin to analyze more than 200,000 lines of COBOL, reducing an estimated eight-month modernization project to eight days.

    Why it matters: The collaboration shows how an autonomous coding agent is being packaged for enterprise legacy modernization inside existing AWS environments, with concrete customer migration figures.

Sep 14

Sep 14Mon
  1. Factory NewsAI score40

    Factory raises $200M at $5B valuation to scale self-improving enterprise software development

    AIFactory has raised $200M at a $5B valuation from investors including Blackstone, Khosla Ventures, and Sequoia Capital, bringing its total funding to over $400 million. The company says it will use the capital to accelerate research, product, and global go-to-market efforts. Factory says hundreds of thousands of developers use its platform, with customers including Nvidia, Blackstone, and T-Mobile.