Skip to contentSkip to stories

Updated

All AI news

Oct 2

Oct 2Fri
  1. PyTorch BlogAI score47

    Helion Linear Backend Boosts vLLM Hopper GPU Inference Throughput Over CUTLASS and DeepGEMM

    AIThe vLLM team integrated Helion, a PyTorch-native kernel DSL, into vLLM's linear backend, using per-shape autotuning to select among Standard GEMM, Split-K, and Swap-AB variants. On NVIDIA Hopper GPUs, the Helion backend outperformed the default CUTLASS and DeepGEMM backends across the evaluated models, with more than 10% throughput gains for some workloads. The work focuses on FP8 and INT8 quantized GEMM.

  2. Latent SpaceAI score43

    Airbnb CTO Ahmad Al-Dahle Details AI-Native Overhaul of Airbnb's Products and Workflows

    AIAirbnb CTO Ahmad Al-Dahle, who joined from Meta in January, says 60% of the company's code is now AI-authored and pull-request throughput per engineer is up about 1.6x. Roughly half of Airbnb's support tickets are now resolved purely by AI, which the company tested with synthetic data before production. Airbnb's internal context graph Everest helped speed up the grocery delivery and airport pickup services, which took eight to nine months and about six weeks to build, respectively.

  3. O'Reilly RadarAI score39

    Coding Agents Benefit From Architectural Decision Records, With Limits

    AIArchitectural Decision Records (ADRs) give coding agents durable project context, helping them distinguish intentional decisions from implementation details. Agents can over-apply accepted but obsolete ADRs, so the author recommends explicit AGENTS.md instructions treating accepted ADRs as binding, prompting agents to flag conflicts, and keeping each ADR current rather than recording amendment logs.

  4. KhazixAI score18

    Khazix rewrites desk pixel clock in Rust, tracks Claude and Codex agents

    AIUsing an AI agent, the author rewrote a desk hardware pixel clock in Rust and linked it to the working status of both Claude and Codex agents. The device also monitors quota resets in real time and shows the day's token consumption. The quoted post notes the project ties into Claude Code's session state with parallel-session support, and says Claude's visual design was far stronger than Codex's.

  5. Hugging Face BlogAI score62

    AutoSynthData generates targeted training data for enterprise agents from failures

    AIServiceNow CoreAI introduced AutoSynthData, which uses a target model's failures and a stronger teacher's successes to generate and validate new agent training tasks. In EnterpriseOps Gym experiments, the Hybrid domain produced 2,000 samples and raised Gemma-4-26B-A4B-it mean Pass@1 by 7.2 percentage points, while the ITSM domain produced 1,994 samples and raised it from 18.77% to 27.18%.

    Why it matters: The post shows how failure analysis, teacher demonstrations, and verifier checks combine into a repeatable pipeline for generating targeted agent training data.

  6. EveryAI score40

    How to Get Better at AI by Asking AI

    AIEvery's senior editor describes moving from single-thread chatbot prompting to delegating complex projects to teams of coordinating subagents, using skills, orchestrator threads, context packets, MCPs, and computer use. He says a subagent workflow verified employee equity costs across multiple grants, strike prices, and vesting schedules, and returned a draft Slack message for approval. The shift was prompted by a June tweet in which Codex placed a colleague at Level 5 of the "Eight Levels of AI Adoption" framework.

  7. Kling AI BlogAI score58

    Kling 4.0 Hands-On Test by Johnson Sheng Shows Stable Motion and Consistency

    AICreative director Johnson Sheng tested Kling 4.0 for commercial video production, focusing on stability during fast camera moves and dynamic action. He reports stable motion in whip pan and handheld push-in shots, a 30-second single-take fight scene, and consistent props and characters across scene changes. The post also covers performance and emotion control through prompt adjustments and multilingual generation. Kling states Kling 4.0 is in closed beta with an official launch planned for October, supporting up to 4K resolution and 10-bit HDR output.

Oct 1

Oct 1Thu
  1. Sophia YangAI score38

    Fireworks details numerical mismatch fixes for stable RL training

    AIFireworks reports that numerical mismatch between training and rollout engines can destabilize reinforcement learning, with a GLM 5.2 experiment showing collapsing reward without alignment and stable reward with it over 25 steps. The post notes that MoE models add further alignment challenges, as Qwen3.5-MoE differences in expert output combination caused disagreement even when one implementation used higher precision. Fireworks says it co-develops its trainer and rollout engine to keep frontier RL training aligned across numerics, kernels, and MoEs.

  2. Comfy BlogAI score47

    852話 Hakoniwa Creates YUI Short Film Using Comfy Agent

    AIArtist 852話 Hakoniwa made YUI, the first animated short film created with Comfy Agent, in three days of focused work for about 200,000 yen excluding labor. She used the Agent to regenerate shots, compare video models such as Wan, MiniMax H3 and Seedance, and check outputs against her character sheet. The main film was made mostly with Seedance 2.5, and the making-of video with MiniMax H3.

  3. Lewis TunstallAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

  4. Anthropic ResearchAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

  5. LangChain BlogAI score58

    LangChain shows how to build a model router in its Open SWE coding agent

    AILangChain built a model router inside its open source coding agent Open SWE that picks one of three models for each thread. In an A/B test against always using GPT-6 Astra, the median cost per thread fell 64% with no measurable change in merged PR rate. The router runs on the thread's first message, using a base prompt, per-tier criteria, and a classifier model, and the post lists next steps including subagent routing and mid-thread re-routing.

Sep 30

Sep 30Wed