Skip to contentSkip to stories

Updated

#Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 1

Oct 1Thu
  1. Manus BlogOfficialAI score47

    Manus 2.0 Adds Game Dev for Building Multiplayer Games Without Coding

    AIManus has launched Game Dev in Manus 2.0, a feature that lets users with no coding experience build games with a real-time tweak panel, asset management, and multiplayer servers. The tweak panel lets users adjust settings such as speed, gravity, damage, and spawn rate while playing. Manus also handles much of the multiplayer infrastructure, including server deployment and networking, so games can be shared and played with friends.

Sep 30

Sep 30Wed
  1. Hamel HusainXAI score38

    Hamel Husain Reviews Claude's New Auto Eval Plugin for Evaluations

    AIHamel Husain has published a longer review of a new Claude Auto Eval plugin after many users asked about it. He invites readers to share their experiences using the plugin and how it went for them. The plugin is part of Claude's ability to help build evaluations and hillclimb on them, as described by @ClaudeDevs.

    Image from @HamelHusain's post
  2. indigoXAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

    Image from @indigox's post
  3. Apple Machine Learning ResearchOfficialAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  4. Varun MohanXAI score40

    Google announces Gemini 4 Argon, a new frontier model for software tasks

    AIGoogle announced Gemini 4 Argon, a new frontier model that delivers frontier performance across complex software tasks, according to Varun Mohan. Thousands of Googlers have been using it internally in Antigravity, and it is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

  5. Google AIOfficialAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    AIGoogle AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    Why it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

    Image from @GoogleAI's post
  6. Google DeepMindOfficialAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  7. Microsoft CopilotOfficialAI score23

    Microsoft's new Copilot combines Home, Code, and Autopilot in one app

    AIMicrosoft's new Copilot brings Home, Code, and Autopilot together in one place for creating custom apps, building decks, automating workflows, and resuming work. Users can start using the Copilot app now and try new features as they become available in Frontier.

    Image from @MSFTCopilot's post
  8. Google · Gemini appOfficialAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

  9. eric zakariassonXAI score22

    Cursor promotes engineering bots in Grok bot marketplace

    AICursor's Eric Zakariasson announced engineering bots available in the Grok bot marketplace on x.ai. The bots can hand off coding tasks to Cursor, manage pull requests through GitHub and Origin plugins, and share video demos of what they build.

    Image from @ericzakariasson's post
  10. FireworksOfficialAI score34

    GLM 5.3 Flash now available for training on Fireworks' Serverless API

    AIFireworks AI has made GLM 5.3 Flash available for training through its Serverless Training API, open to all users. The model supports both vision and text inputs. Fireworks says it performs well on its benchmarks for agentic coding, document analysis, and tool use while remaining cost-efficient to serve.

  11. ClineOfficialAI score34

    Cline desktop app can run agents on remote Linux servers over SSH

    AICline says its desktop app can keep running on a laptop while the agent executes on any Linux machine reachable via SSH, configured under Settings → Remote. The setup requires no root access, no npm, and no public port, uploading a self-contained helper and tunneling only the authenticated Cline protocol.

  12. GitHubOfficialAI score42

    Project HydraFusion Now Available in GitHub Copilot App and VS Code

    AIProject HydraFusion is now available in the GitHub Copilot app and VS Code, where users can select it like any other model. Behind the scenes, it routes a task across multiple models to draft, critique, revise, or escalate, then returns one combined result.

    Video from @github's post
  13. Ant LingOfficialAI score31

    Ant Ling model turns plain-language prompts into interactive Three.js pages

    AIAnt Ling can convert plain-language prompts into standalone, interactive Three.js pages covering topics such as an internal combustion engine, an optical-disc reader, paramecium organelles, and vector-field divergence. The output is runnable code rather than just an explanation.

    Video from @AntLingAGI's post
  14. Ant LingOfficialAI score38

    Ling-3.1-flash ports C image library to Rust with 8.015× speedup

    AIAnt Ling reports that its Ling-3.1-flash model completed a roughly 20-hour Rust port of a C image library. After a performance regression caused by busy-waiting workers and a parallelism adjustment, the model recovered and reached an 8.015× speedup. All 30 correctness checks passed.

    Image from @AntLingAGI's post
  15. DeepSeek HarnessXAI score62

    DeepSeek Harness v0.2 preview launches as a desktop app for macOS and Windows

    AIDeepSeek releases the DeepSeek Harness v0.2 preview with a desktop app for macOS and Windows. The release adds a plugin manager for installing, disabling, and uninstalling plugins without terminal commands, plus an experimental creator mode that generates plugins from user descriptions. The company says DeepSeek Harness is now the most widely used coding agent among users of the official DeepSeek API by DAU and daily sessions.

  16. Google GeminiOfficialAI score45

    Gemini skills now proactive, stackable, and support reference files

    AIGemini can build custom skills from chats and apply a saved skill automatically when a prompt matches it. Multiple skills can be stacked for larger tasks, such as combining a personal writing style skill with a brand guidelines skill. Starting today, skills can include reference files such as plain text documents, PDFs, or images, with sharing and Google Drive file support coming soon.

  17. howie.seriousXAI score22

    Why local AI agents like Claude Code and Codex rely on shell access

    AILocal and desktop agents such as Claude Code and Codex are powerful largely because they can use the shell, which connects them to the whole CLI ecosystem. The post lists tools including git, ffmpeg, curl, pandoc, gh, cron, and ssh as examples. It also says the video itself was produced by an agent operating the shell.

    Video from @howie_serious's post
  18. Cloudflare Blog · AIOfficialAI score72

    Cloudflare launches Auto Router in AI Gateway to cut AI token spend

    AICloudflare has released Auto Router in public beta through AI Gateway, where setting the model to cloudflare/auto routes each request to a model judged capable enough for the task. Internal tests showed up to 30% cost savings against frontier models, and on a 97-task internal benchmark cloudflare/auto scored 86.6% at $0.0084 per success versus 96.6% at $0.0210 for Claude Opus 5.5. The router is free during beta.

    Why it matters: The source gives a benchmark table of success rates and costs per trial, showing how routing trades quality against price for a gateway deployment.

  19. Karl's AI WattsXAI score38

    Can you keep your session after switching models in magpie?

    AIKarl's AI Watts asks whether a menu-bar tool can switch models while preserving the existing conversation, so users avoid re-explaining their project each time. The post frames this as the reason they want to keep the menu bar tool, which the quoted post describes as magpie, a menu-bar switcher for 20+ agents including Claude Code and Codex that also offers a local gateway.

Sep 29

Sep 29Tue
  1. Factory NewsOfficialAI score42

    Factory Launches Generally Available Automations to Run Recurring Engineering Workflows

    AIFactory's Automations, now generally available, let users describe a recurring workflow, set a schedule or event trigger, and have its Droid run it, with the model chosen per task. Templates cover ticket-to-PR, code review, security audits, PR babysitting, and morning Slack briefs. Among enterprise organizations using Automations in the past 30 days, 52% used automated code review, 48% used security review, and 35% used AutoWiki.

  2. v0OfficialAI score42

    GPT-6.1 Sol now available in v0

    AIGPT-6.1 Sol is now live in v0, with access via the v0 app link provided in the post. The quoted Vercel post says it is also on AI Gateway and improves on GPT-6 Sol for coding, computer use, multi-step workflows, and complex document analysis.

  3. Tibor BlahoXAI score78

    OpenAI's DevDay 2026 brings dots agents, GPT-6.1 Sol, and Ultrafast speed tier

    AIOpenAI announced more than 20 updates at DevDay 2026, including dots always-on agents, GPT-6.1 Sol, Ultrafast token generation, ChatGPT Space, and a $500/month Pro 500 plan. GPT-6.1 Sol is priced at $2 input and $10 output per 1M tokens and is available in the API as gpt-6.1-sol. Ultrafast generates tokens up to 8x faster in Codex and up to 6x faster in the API.

    Why it matters: The post lists dozens of OpenAI DevDay 2026 changes across models, agents, plans, and APIs, useful for scanning what shipped and who gets access.

    Image from @btibor91's post
  4. Sherwin WuXAI score38

    OpenAI's Ultrafast Astra runs up to 8x faster in Codex

    AIOpenAI's Ultrafast mode for Astra runs up to 8x faster than Astra Standard and 4x faster than Astra Fast in Codex. Sherwin Wu says inference now feels nearly instant, shifting the bottleneck to users' next complaints.

  5. Baseten BlogOfficialAI score38

    Baseten Partners With OpenAI to Offer Open Models to OpenAI Customers

    AIBaseten announced a partnership with OpenAI that makes open models powered by Baseten available to OpenAI customers for multi-model agentic coding. The company says organizations can route each task to the best-fit open or closed model, with Codex and GPT models among the options, and that Baseten's day-zero access to new open models lets teams evaluate them quickly. Baseten also cites US-based infrastructure with zero data retention for all prompts and capacity across more than 90 clusters in 20+ clouds.

  6. CognitionOfficialAI score34

    Devin now uses ChatGPT plan quota for OpenAI models

    AICognition says usage of Devin will count toward the Codex and ChatGPT Work usage already included in a user's plan, with an option to set a lower Devin limit. OpenAI models in Devin Cloud's Lite, Normal, Ultra, and Fusion modes will be paid for by the ChatGPT plan once the account is connected, while other models use the Devin quota.

  7. Kilo (acq. by Anaconda)OfficialAI score32

    Kilo adds Sign in with ChatGPT across all its surfaces

    AIKilo now supports Sign in with ChatGPT, letting users apply their ChatGPT plan across VS Code, JetBrains, the CLI, Cloud Agents, Code Reviews, and the mobile app. The integration enables running 10 agents in parallel, and Cloud Agents continue working after the user closes their laptop.

    Image from @kilocode's post
  8. BAAI · new models on Hugging FaceOfficialAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    AIBAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    Why it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.