Skip to contentSkip to stories

Updated

#Coding

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 5

Sep 5Sat
  1. AI at MetaOfficialAI score38

    AIRA₃ ensemble places 8th with gold-medal results in live competition

    AIMeta's AIRA₃ entered the live competition with an ensemble of models, and the 8th-ranked gold-medal entry combined GPT 5.5 (w/ OpenCode) and Claude 4.8 (w/ ClaudeCode). Post-hoc testing found Muse Spark 1.2 (w/ MuseCode) also reached gold-medal level, while Muse Spark 1.1 (w/ OpenCode) and GLM 5.2 (w/ OpenCode) reached silver-medal level, all graded on the same private test set.

    Image from @AIatMeta's post

Sep 4

Sep 4Fri
  1. Matei ZahariaXAI score46

    Qwen3.8-Flash-Next runs at 68.3 tok/s on a single RTX 5090

    AIA Berkeley Sky Lab researcher says stronger open models and new inference systems will make powerful local AI practical. The linked post reports Qwen3.8-Flash-Next running at 68.3 tok/s on a single RTX 5090 using an NVFP4 checkpoint, with 63GB host RAM and a 51GB n-gram table stored on NVMe at about 0.5% throughput cost.

  2. Andrew NgXAI score42

    Andrew Ng maps key skills for using AI coding agents effectively

    AIAndrew Ng presented an AI Engineering Skills Map for using coding agents such as Claude Code, Codex, Cursor, OpenCode, and Pi. The workflow he describes covers planning, execution, and deployment with monitoring, and he identifies five key skills: directing the workflow, enabling agent autonomy, reviewing the work, customizing the agent and its environment, and coding agent foundations. The source says these skills matter more as the agents evolve quickly.

Sep 3

Sep 3Thu
  1. TinkerOfficialAI score51

    Bespoke Labs post-trains Inkling on one code repo and reports broader coding gains

    AIBespoke Labs post-trained the Inkling base model on a single GitHub repository using supervised fine-tuning and GRPO reinforcement learning. The post reports a 57-point improvement on the held-out fontTools evaluation over the base model, along with gains on Terminal-Bench 2.1 and SWE-bench Lite. It also says the post-trained model uses about 40% fewer tokens.

    Image from @tinkerapi's post
  2. Mark ChenXAI score80

    Mark Chen announces GPT-6 Astra with computer use and agent oversight

    AIOpenAI researcher Mark Chen announced GPT-6 Astra, which he described as the company's most capable and aligned model yet. He said it can build and test software, work across apps on a computer, and help with open scientific problems. The post also highlights improved computer use compared with Operator and stronger monitoring that can stop potentially unauthorized agent actions.

    Why it matters: The post links a named model release to specific capabilities like computer use and aligned agent behavior, giving readers concrete claims to check against the model.

  3. Thomas DohmkeXAI score38

    Copilot's new search tool understands codebase intent and decision context

    AIMicrosoft and Copilot's Thomas Dohmke announced a search tool that understands a codebase beyond literal phrases, returning results based on intent, semantic reasoning, and the context behind key decisions. The post's quoted Entire context describes Agentic Search, an API across accessible repos that returns the code, session, transcript, and prompt behind a change.

Sep 2

Sep 2Wed
  1. Daniel HanXAI score34

    Stanford's Modern Software Developer course adds AI-native engineering curriculum

    AIMihail Eric announced the 2026 edition of his Stanford course "The Modern Software Developer," with 85% of the Fall 2025 material replaced by AI-native topics such as agent skills, context engineering, and agentic code review. Students will ship pull requests to real open-source AI repositories, with partners including Browserbase, HeyGen, and CopilotKit offering mentorship.

  2. Google AI StudioOfficialAI score78

    Google releases Gemini 3.8 Flash and restricted 3.8 Flash Cyber model

    AIGoogle introduces Gemini 3.8 Flash for coding, agentic tasks, and multi-step reasoning, priced at $0.75 per million input tokens and $3.75 per million output tokens during the introductory period. Gemini 3.8 Flash Cyber targets vulnerability detection and automated patching and is available only to trusted defenders through the new Fairwind Program. The introductory price expires December 31, 2026, after which $1.50 and $7.50 per million tokens apply.

    Why it matters: The post separates a general coding and agent model from a restricted cyber variant, showing how one shared core is deployed under different access and safety tiers.

  3. Varun MohanXAI score57

    Gemini 3.8 Flash released with gains in agentic coding and knowledge work

    AIGoogle's Gemini 3.8 Flash is out, and Varun Mohan says it substantially improves on 3.7 Flash for agentic coding and general knowledge work. It is now available to everyone on Antigravity. The attached benchmark table lists Gemini 3.8 Flash at $0.75 per 1M input tokens and $3.75 per 1M output tokens, with introductory pricing through December 31, 2026.

    Image from @_mohansolo's post
  4. Google AI DevelopersOfficialAI score32

    Gemini 3.8 Flash builds interactive 3D hardware teardown visualizers with Three.js

    AIGoogle AI Developers says Gemini 3.8 Flash, built for complex reasoning, generated an interactive 3D visualizer using Three.js in Google AI Studio. The visualizer produces physically proportioned teardowns of hardware devices, automatically splitting each device into layers that users can explode and inspect with a deconstruction slider.

    Video from @googleaidevs's post
  5. Logan KilpatrickXAI score42

    Gemini 3.8 Flash scores 73.7% on DeepSWE 1.1 benchmark

    AIGoogle's Gemini 3.8 Flash reached 73.7% on the DeepSWE 1.1 benchmark, according to a post from Logan Kilpatrick. The post gives no further details on methodology, comparisons, or pricing.

    Image from @OfficialLoganK's post
  6. koray kavukcuogluXAI score62

    Gemini 3.8 Flash claims stronger engineering results at lower cost than larger models

    AIGoogle's Koray Kavukcuoglu says Gemini 3.8 Flash is a major step up from Gemini 3.7 Flash and outperforms most larger frontier models on complex engineering problems at a fraction of the cost. The attached DeepSWE V1.1 chart, sourced to Datacurve AI, plots average cost per task against score for Gemini 3.8 Flash and other models. A link to Google's blog post with more details is included.

    Image from @koraykv's post
  7. Logan KilpatrickXAI score62

    Google releases Gemini 3.8 Flash with gains in agentic and coding tasks

    AIGoogle announced Gemini 3.8 Flash, its third updated Flash model in six weeks, citing improvements in agentic and coding capabilities. The benchmark table lists input at $0.75 and output at $3.75 per 1M tokens, with introductory pricing of $1.50 and $7.50 expiring December 31, 2026. Terminal-bench 2.1 shows 89.4% for Gemini 3.8 Flash against 85.8% for Gemini 3.7 Flash.

    Image from @OfficialLoganK's post
  8. Google AI StudioOfficialAI score62

    Google releases Gemini 3.8 Flash with improved coding, agent, and reasoning

    AIGoogle AI Studio announced Gemini 3.8 Flash, which it calls its most intelligent workhorse model. The company says it brings significant improvements over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning in specialized domains. It is available at the same introductory price as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, through the Gemini API and AI Studio.

    Image from @GoogleAIStudio's post

Sep 1

Sep 1Tue
  1. Meituan LongCatOfficialAI score46

    LongCat-2.0 Now Free to Try in Cline

    AIMeituan's LongCat-2.0, a 1.6T open-weights MoE model with a 1M context window, is now free to use in Cline. Cline's post says it scores similarly to Claude Opus 4.7 and Gemini 3.1 Pro. Users can select it under free models via /model after installing Cline with npm i -g cline.

  2. catXAI score50

    Anthropic's Claude Fable 5.1 enables more ambitious, months-long projects

    AIAnthropic's team says Claude Fable 5.1 has let them take on projects that previously would have taken months, and invites users to try it in Claude Code, Claude Cowork, and Claude Tag. The post asks what big bets users want to make, and it builds on Anthropic's announcement of Claude Fable 5.1 and Claude Mythos 5.1 as its most advanced models for coding and knowledge work.

  3. Anthropic · YouTubeOfficialAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

Aug 31

Aug 31Mon
  1. Claude Apps Release NotesOfficialAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

  2. Philipp SchmidBlogAI score60

    Frontier models now compose Bash workflows that replace dedicated coding tools

    AIThe author rebuilt an agent harness with only a bash tool and a media viewer, and task completion stayed in the same range. Three example workflows show multi-file edits, bisecting a flaky test, and correlating compressed logs in SQLite, with the intermediate data kept out of the model context. In a comparison against separate file, edit, and search tools on the same coding tasks, the shell-centered setup performed on par or better, though the author notes images still need a multimodal channel.

Aug 29

Aug 29Sat
  1. Tencent HyOfficialAI score47

    Tencent Hunyuan open-sources Hy4 preview, a 770B MoE model

    AITencent Hunyuan has open-sourced Hy4 preview under Apache 2.0, a flagship mixture-of-experts model with 770B total parameters, 49B active per token, and a 1M context window. Blind evaluation by 163 internal experts across 203 engineering tasks gave it an average score of 2.99, narrowly ahead of GLM 5.3 at 2.92 and Kimi K3 at 2.94. The model includes a native MTP layer for speculative decoding and is trained on production workflows spanning software engineering, data analysis, game development, and scientific research.

Aug 28

Aug 28Fri
  1. Thomas DohmkeXAI score25

    Entire launches one API for code and coding sessions

    AIEntire positions itself as a unified API for code and coding sessions, working across any agent, repo, and session as a coding system of record. The post frames this as a single interface layer for coding work, comparing it to unified-interface products in payments, models, and banking.

  2. Andrew NgXAI score20

    Andrew Ng maps software engineering fundamentals for agentic coding era

    AIAndrew Ng published an AI Engineering Skills map covering the software engineering fundamentals developers need when working with coding agents. He argues that understanding full-stack architecture, data management, system design, security, reliability, and production scaling lets developers steer agents toward the right tradeoffs in latency, availability, consistency, and cost. Without these fundamentals, vibe-coded applications often end up with poor tradeoffs the developer never anticipated.

Aug 27

Aug 27Thu
  1. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.

Aug 26

Aug 26Wed
  1. Cursor ChangelogOfficialAI score46

    Cursor Cloud Agents now let you start projects from scratch without a repo

    AICursor Cloud Agents no longer require a connected GitHub or other third-party SCM provider to begin work. Users select "Start from scratch" in the repo picker, and Cursor creates an Origin repo in the background that can be saved as a private or internal repo via "Create repo." Cursor also now port-forwards the cloud agent's live environment to the browser for previews, and a connected Vercel account lets users publish a live URL.

  2. LMSYS OrgOfficialAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogOfficialAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Fireworks AI BlogOfficialAI score46

    DeepSeek V4 Pro Solves Security Tasks at Half the Cost Per Success

    AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.

  3. Z.ai Release NotesOfficialAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.