Skip to contentSkip to stories

Updated

#Coding

Showing low-relevance items too. Hide low-relevance items

Sep 25

Sep 25Fri
  1. François CholletAI score32

    Chollet: Software engineering difficulty stays constant across abstraction levels

    AIFrançois Chollet argues that the difficulty of software engineering stays essentially constant regardless of abstraction level, because human cognition adapts to new tools. He says tools are affordances rather than magic wands that eliminate work, and that great software engineering remains immensely challenging despite changed workflows. Simon Willison's background post similarly argues that coding agents make software engineering harder, requiring extraordinary discipline and knowledge.

Sep 24

Sep 24Thu
  1. GitHub Blog · AI & MLAI score46

    GitHub Copilot app's canvases argue chat is the wrong AI interface

    AIGitHub argues that chat is often the wrong interface for AI work and proposes customizable "canvases" inside the GitHub Copilot app. Canvases are full-stack applications running without browser chrome that can communicate bi-directionally with the Copilot agent and execute code locally. The post cites examples including a Connect 4 game, a Winget package manager UI, and a SQLite database interface.

  2. Lewis Tunstall @ COLM 🌉AI score42

    Hugging Face releases over 5,000 RL environments for data science tasks

    AIHugging Face released SmolDataEnvs, more than 5,000 open-source RL environments aimed at real-world data science tasks. They target the gap between simple educational games and frontier-level benchmarks, especially for improving coding in models under 10B parameters. The environments are designed as a testbed for developing new RL methods such as GRPO or OPSD.

  3. GitHub Blog · AI & MLAI score66

    GitHub Security Lab shows an LLM agent running AI-driven fuzzing for C/C++ projects

    AIGitHub Security Lab describes the Fuzzing Taskflow, an LLM agent pipeline that identifies entrypoints, writes harnesses, runs AFL++, reads coverage reports, and triages crashes for C/C++ repositories. The agent makes decisions while MCP tools handle execution, and state is stored in a SQLite database. The post also warns that the taskflow runs AFL and build commands directly on the host, so it should be used only in disposable environments without elevated privileges.

    Why it matters: The post explains how an LLM agent automates fuzzing steps like harness writing, coverage gap chasing, and crash triage, with a runnable workflow and design tradeoffs.

Sep 23

Sep 23Wed
  1. Amp NewsAI score42

    Amp Lets Teams Share a Runner Across Their Workspace

    AIAmp users can now share a runner with their workspace by starting it with --share, letting everyone spawn threads on that machine from ampcode.com. Shared runners appear under Shared Runners in the picker, and --amp-env gives them workspace and project Secrets & Env Vars but never personal ones. Amp warns that collaborators run code as the owner with their files and credentials, so sharing should be limited to trusted people, and workspace admins can disable runner sharing in Member Settings.

  2. Amp NewsAI score34

    Amp's macOS app now runs threads on your Mac without a terminal

    AIThe Amp macOS app now starts a runner automatically, so threads can run on your Mac without keeping amp --no-tui open in a terminal. Users add folders or projects under Runner in App Settings, then select "This Mac" when starting threads from ampcode.com, a phone, or Puck. A Keep This Mac Awake option prevents sleep while the runner is on and the Mac is plugged in, though the screen still turns off and locks.

  3. Karl's AI WattsAI score22

    Notch admits he is enjoying vibe coding after earlier opposing AI coding

    AIMinecraft creator Notch says on X that he is enjoying vibe coding and admits he may have been slightly wrong. Months earlier he had publicly rejected AI-written code, but he later began having AI build internal tools such as a map editor and node graph tools. The main post adds that he has accumulated a set of small tools for himself before much game development has happened.

  4. Mike KnoopAI score25

    Formal verification gains ground, but human understanding remains an alignment gap

    AIMike Knoop argues that formal verification is becoming feasible and is important for security. He adds that it does not automatically build human understanding, which he calls an even bigger alignment problem. The post is framed as a reply to Boris Cherny's report that Claude Opus 5.5 helped formally verify the Claude Agent SDK in Lean, producing 16 bug-fix PRs.

Sep 22

Sep 22Tue
  1. Alex AlbertAI score37

    Claude prompt recreates 1906 Market Street in Blender for video

    AIA prompt shared by Alex Albert asks Claude to recreate San Francisco's Market Street as it stood on April 17, 1906, before the earthquake, using Blender. It requires building a source file from Sanborn fire insurance maps, the Miles Brothers film, period photos, and USGS topography, with reusable Blender Python generators for facades, street lamps, and vehicles, ending in a 10-second video up the street.

  2. Mike KriegerAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  3. Boris ChernyAI score62

    Claude Opus 5.5 ports HAProxy to Rust faster and cheaper than Fable 5.1

    AIAnthropic introduced Claude Opus 5.5 as the first model in its Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 for most tasks at 40% lower run cost than Opus 5. Boris Cherny reports that Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of its tests, with Opus 5.5 finishing in 9.5 hours versus 12 hours and at 51% less cost.

  4. StepFunAI score27

    StepFun's Step Code tops Terminal-Bench 2.1 and Multi-Frame with fewer tokens

    AIStepFun's Step Code passed 72 of 89 tasks (80.9%) on Terminal-Bench 2.1, tying for the highest pass rate among evaluated harnesses while using fewer tokens than the other tied leaders. On Multi-Frame, it passed 110 of 150 tasks (73.3%) and averaged 5.09M tokens per task, the highest pass rate and lowest token use among six harnesses evaluated.

    Image from @StepFun_ai's post
  5. StepFunAI score52

    StepFun releases Step Code v0.1.0 as an open-source coding CLI

    AIStepFun has released Step Code v0.1.0, an open-source command-line tool under the MIT License that covers reading and editing code, running tests, and shipping from one CLI. The post reports 80.9% on Terminal-Bench 2.1 and 73.3% on Multi-Frame, a 150-task long-horizon benchmark from StepFun. It also includes one-command static site publishing with StepPage and links the GitHub repository.

    Image from @StepFun_ai's post
  6. WorkBuddyAI score18

    HKUST students build two AI workbenches with WorkBuddy, win Game Track

    AIHKUST's Anchor team used WorkBuddy to build two production-ready workbenches and won the Game Track championship. Kaiwu Producer creates a complete FPS game in 8 hours through full-pipeline 3D generation with an AI-driven narrative memory engine and zero human intervention. Anchor is a de-labeling narrative engine that automatically detects stereotypical dependencies.

    Video from @WorkBuddy_AI's post
  7. Tencent HyAI score44

    WebCraftBench Scores AI-Built Websites by Live Use and Human Preference

    AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.