Skip to contentSkip to stories

Updated

#Coding

Sep 22

Sep 22Tue
  1. Alex AlbertAI score37

    Claude prompt recreates 1906 Market Street in Blender for video

    AIA prompt shared by Alex Albert asks Claude to recreate San Francisco's Market Street as it stood on April 17, 1906, before the earthquake, using Blender. It requires building a source file from Sanborn fire insurance maps, the Miles Brothers film, period photos, and USGS topography, with reusable Blender Python generators for facades, street lamps, and vehicles, ending in a 10-second video up the street.

  2. Mike KriegerAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  3. Boris ChernyAI score62

    Claude Opus 5.5 ports HAProxy to Rust faster and cheaper than Fable 5.1

    AIAnthropic introduced Claude Opus 5.5 as the first model in its Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 for most tasks at 40% lower run cost than Opus 5. Boris Cherny reports that Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of its tests, with Opus 5.5 finishing in 9.5 hours versus 12 hours and at 51% less cost.

  4. StepFunAI score52

    StepFun releases Step Code v0.1.0 as an open-source coding CLI

    AIStepFun has released Step Code v0.1.0, an open-source command-line tool under the MIT License that covers reading and editing code, running tests, and shipping from one CLI. The post reports 80.9% on Terminal-Bench 2.1 and 73.3% on Multi-Frame, a 150-task long-horizon benchmark from StepFun. It also includes one-command static site publishing with StepPage and links the GitHub repository.

  5. WorkBuddyAI score18

    HKUST students build two AI workbenches with WorkBuddy, win Game Track

    AIHKUST's Anchor team used WorkBuddy to build two production-ready workbenches and won the Game Track championship. Kaiwu Producer creates a complete FPS game in 8 hours through full-pipeline 3D generation with an AI-driven narrative memory engine and zero human intervention. Anchor is a de-labeling narrative engine that automatically detects stereotypical dependencies.

  6. Tencent HunyuanAI score44

    WebCraftBench Scores AI-Built Websites by Live Use and Human Preference

    AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.

Sep 21

Sep 21Mon
  1. Amp NewsAI score36

    Amp Runners Add Git Worktree Creation and Secret Injection

    AIAmp runners can now create Git worktrees from the directory picker, with each new worktree placed as a sibling folder on a new branch off the current HEAD while uncommitted changes stay in the original checkout. Runners can also opt in with --amp-env to inject Secrets & Env Vars configured on ampcode.com into thread shell commands, MCP servers, and plugins, with changes applied to the next thread without a restart.

  2. Latent.SpaceAI score37

    TypeSafe CEO Jev on reliable System One Models beyond chat-first AI

    AITypeSafe CEO Jev argues AI can solve extremely hard problems yet still fail at basic automation, so his company builds reliable decision-making models inside software rather than chat interfaces. He says the company rejects public benchmarks and API-layer refusals, and that data and task fit matter more than brute-force compute. He also says System One Models could reshape coding agents and software, and that he would not pre-train a model from scratch even with $1 billion.

  3. Xiaomi MiMoAI score31

    Xiaomi MiMo-V2.6-Pro climbs to Code Arena WebDev top 10

    AI🔥 Quoted post (context): MiMo-V2.6-Pro has returned to the top 10 on Code Arena WebDev, debuting around #10 overall and about #3 among open-weights models under an MIT license. It scores 1628 points on AutoEval, tying Claude Fable 5 (High) and narrowly ahead of Hy4-preview (1624). That is a +153 point gain over the previous MiMo-V2.5-Pro (1475). Among open-weights models it trails Kimi K3 Max (#1) by 46 points and Qwen3.8 Flash Next (#2) by 7 points. These are early AutoEval scores from a Reward Model trained on Arena's human preference data, not live human votes; scores are expected to converge as more live votes arrive.

  4. Xiaomi MiMoAI score36

    Xiaomi MiMo-V2.6 unifies code, design, and tool use across creative outputs

    AIXiaomi's MiMo-V2.6 combines code, design, and tool use to build frontend interfaces, presentations, Figma-linked visual assets, and videos. The post says MiMo-V2.5-TTS supports narration in video production, and that the model can compose music, including an orchestral piece for around ten instruments that can be converted to MIDI. On Design Arena, the Pro version reportedly performs comparably to Claude Opus 5 and GPT-5.6 Sol.

  5. Xiaomi MiMoAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

  6. Xiaomi MiMo · new models on Hugging FaceAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.

Sep 20

Sep 20Sun
  1. xAI News (Grok)AI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

Sep 19

Sep 19Sat
  1. StepFunAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

Sep 18

Sep 18Fri
  1. Mike KnoopAI score30

    Mike Knoop wonders what an underscore.js equivalent for AI looks like

    AIMike Knoop asks what the underscore.js equivalent for AI would look like, noting that such programming primitives feel close. He adds that he barely reads or writes code anymore despite these emerging tools. The quoted post introduces Probably, a toy programming language built around Jev, where constructs like "feels," "match," and "while" let AI make decisions within ordinary code.

  2. GitHub Blog · AI & MLAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.

Sep 17

Sep 17Thu
  1. Together AI BlogAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  2. Sherwin WuAI score44

    ChatGPT for Word launches, bringing ChatGPT and Codex into Microsoft Word

    AIOpenAI has released ChatGPT for Word, letting users access ChatGPT and Codex directly inside Microsoft Word. The post says ChatGPT for Excel and PowerPoint has been growing rapidly, and Word completes that set. The quoted ChatGPT post adds that the tool can draft from notes, rewrite paragraphs, proofread, suggest edits, and flag formatting issues.