A few of you asked what mods are, here's a quick walkthrough!
AIThey’re really just plugins with a few special functions that let your code run inside Claude Code, like middleware. (don't worry, you can just ask Claude to write them for you 😛)
Updated
Updated
AIThey’re really just plugins with a few special functions that let your code run inside Claude Code, like middleware. (don't worry, you can just ask Claude to write them for you 😛)
AIThe vLLM team integrated Helion, a PyTorch-native kernel DSL, into vLLM's linear backend, using per-shape autotuning to select among Standard GEMM, Split-K, and Swap-AB variants. On NVIDIA Hopper GPUs, the Helion backend outperformed the default CUTLASS and DeepGEMM backends across the evaluated models, with more than 10% throughput gains for some workloads. The work focuses on FP8 and INT8 quantized GEMM.
AIFine-tuning NVIDIA Nemotron 3.5 ASR cut word error rate on Najdi and Hijazi Saudi Arabic from 55% to 30%. Check out our new tutorial and learn how to adapt it for other dialects and languages:
AI…shared context repo any harness can read 2/ half a day of research in 5 min 3/ evals that test our product the way agents use it 🎧:
AIIn this DevEx program sprint, we tested agent governance on Gemini Enterprise from the ground up—using standard external accounts and zero internal shortcuts to find + fix developer friction fast ↓
AIA: Binary labels force clearer thinking and more consistent labeling, and involves significantly lower complexity to operationalize.
AIAirbnb CTO Ahmad Al-Dahle, who joined from Meta in January, says 60% of the company's code is now AI-authored and pull-request throughput per engineer is up about 1.6x. Roughly half of Airbnb's support tickets are now resolved purely by AI, which the company tested with synthetic data before production. Airbnb's internal context graph Everest helped speed up the grocery delivery and airport pickup services, which took eight to nine months and about six weeks to build, respectively.
AIingi_erlingsson wrote down how he reads a moodboard: the feeling first, then color, tone, and camera, then @armandgw made it a Krea Agent skill. learn how below👇
AIArchitectural Decision Records (ADRs) give coding agents durable project context, helping them distinguish intentional decisions from implementation details. Agents can over-apply accepted but obsolete ADRs, so the author recommends explicit AGENTS.md instructions treating accepted ADRs as binding, prompting agents to flag conflicts, and keeping each ADR current rather than recording amendment logs.
AIUsing an AI agent, the author rewrote a desk hardware pixel clock in Rust and linked it to the working status of both Claude and Codex agents. The device also monitors quota resets in real time and shows the day's token consumption. The quoted post notes the project ties into Claude Code's session state with parallel-session support, and says Claude's visual design was far stronger than Codex's.
AIFollowing @karpathy's advice, write ASD-STE100 into the system prompt.
AIServiceNow CoreAI introduced AutoSynthData, which uses a target model's failures and a stronger teacher's successes to generate and validate new agent training tasks. In EnterpriseOps Gym experiments, the Hybrid domain produced 2,000 samples and raised Gemma-4-26B-A4B-it mean Pass@1 by 7.2 percentage points, while the ITSM domain produced 1,994 samples and raised it from 18.77% to 27.18%.
Why it matters: The post shows how failure analysis, teacher demonstrations, and verifier checks combine into a repeatable pipeline for generating targeted agent training data.
AIEvery's senior editor describes moving from single-thread chatbot prompting to delegating complex projects to teams of coordinating subagents, using skills, orchestrator threads, context packets, MCPs, and computer use. He says a subagent workflow verified employee equity costs across multiple grants, strike prices, and vesting schedules, and returned a draft Slack message for approval. The shift was prompted by a June tweet in which Codex placed a colleague at Level 5 of the "Eight Levels of AI Adoption" framework.
AICreative director Johnson Sheng tested Kling 4.0 for commercial video production, focusing on stability during fast camera moves and dynamic action. He reports stable motion in whip pan and handheld push-in shots, a 30-second single-take fight scene, and consistent props and characters across scene changes. The post also covers performance and emotion control through prompt adjustments and multilingual generation. Kling states Kling 4.0 is in closed beta with an official launch planned for October, supporting up to 4K resolution and 10-bit HDR output.
AIClaude Code can now be modded to change how it behaves, customize its UI, and add custom features. Mods are written in a few lines of TypeScript or generated by Claude, and ship inside plugins installed with /plugin in the CLI or desktop app.
AIAndrej Karpathy says people will spend more time understanding language model outputs and suggests better formats than plain text. He recommends asking for ASD-STE100 controlled-language explanations, diagrams, interactive HTML pages, or custom explainer videos, and he is most bullish on the video format.
AIThe article compares LangChain/LangGraph and CrewAI workflow orchestration with OpenRouter's native model and provider routing. It says OpenRouter's models parameter provides an ordered, error-driven fallback list, while LangGraph and CrewAI handle state, memory, and delegation. Frameworks can also run on OpenRouter as the model layer underneath.
AIFireworks reports that numerical mismatch between training and rollout engines can destabilize reinforcement learning, with a GLM 5.2 experiment showing collapsing reward without alignment and stable reward with it over 25 steps. The post notes that MoE models add further alignment challenges, as Qwen3.5-MoE differences in expert output combination caused disagreement even when one implementation used higher precision. Fireworks says it co-develops its trainer and rollout engine to keep frontier RL training aligned across numerics, kernels, and MoEs.
AI@MicrosoftAI @MicrosoftAI STT plugin guide here:
AIAnalyze your funnel data in Replit and ask: “Where can we increase activation?” Then follow up: “Make it a dashboard.” You don’t need to know the final output before you start. One question can lead to an insight, then something useful.
AIAnthropic's Thariq describes Claude Code mods, which can read conversation scope such as turn count and token usage. Mods run in process, so they can spawn subagents, parse their results with structured output, and modify the UI, which hooks cannot do. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app.
AIArtist 852話 Hakoniwa made YUI, the first animated short film created with Comfy Agent, in three days of focused work for about 200,000 yen excluding labor. She used the Agent to regenerate shots, compare video models such as Wan, MiniMax H3 and Seedance, and check outputs against her character sheet. The main film was made mostly with Seedance 2.5, and the making-of video with MiniMax H3.
AIWe hope you find them as useful (and fun) as we do, and can't wait to see what you build!
AIMods run with the same access to your machine as Claude Code itself, so make sure to only install mods from sources you trust.
AIunderstand tasks 2. understand the models 3. build the router **in the harness** 4. track outcomes lower costs, with no performance hit
AIverifiers to build the environment 2. Hosted Training to run the RL loop 3. Prime Sandboxes to execute the model's code and return the execution score 4. Prime Inference to serve the LLM judge and frontier baselines
AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.
AIA: No. Different types of evaluators have different costs (code vs. LLM judge), which you need to weigh before building an eval.
AII used AI to further develop this pixel clock at home, directly linking it to all the status of Claude Code on my computer, with support for parallel sessions. To be honest, when it comes to aesthetics, Claude is still far too strong. The results Codex produced were genuinely ugly, I'm in tears...
AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.
Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.
AILangChain built a model router inside its open source coding agent Open SWE that picks one of three models for each thread. In an A/B test against always using GPT-6 Astra, the median cost per thread fell 64% with no measurable change in merged PR rate. The router runs on the thread's first message, using a base prompt, per-tier criteria, and a classifier model, and the post lists next steps including subagent routing and mid-thread re-routing.
AILatent.Space shares a podcast episode in which OpenAI's AriX and Nikunj Handa discuss OpenAI's new agent stack. The episode covers Computer Use, Dots cloud computers for agents, and the Decisions API, which the speakers say went from idea to product in weeks.
AIIn Ep 3 of Agent Clinic, Dani Zamora & @matt_feroz build an automated eval suite for a LangGraph agent in 60 mins. Here's the 4-step framework to go from vibes to benchmark ANY AI agent →