Skip to contentSkip to stories

Updated

#Tutorial/How-to

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 1

Oct 1Thu
  1. Comfy BlogAI score47

    Hakoniwa uses Comfy Agent to make the animated short YUI

    AIArtist 852 Hakoniwa made YUI, described as the first animated short created with Comfy Agent, which the ComfyUI team says took three days of focused work by one person at about 200,000 yen in total cost, excluding labor. The source says the film was made mostly with Seedance 2.5 and the making-of video with MiniMax H3, with Comfy Agent used to regenerate shots and compare video models.

  2. Prime IntellectAI score32

    Extropic uses Prime Intellect to post-train Qwen3.6-35B-A3B for thermodynamic ML

    AIExtropic post-trained Qwen3.6-35B-A3B with Prime Intellect for thermodynamic ML research, nearly tripling its held-out eval results in about 100 GRPO steps. The team built a custom RL environment with verifiers and trained on Hosted Training, Prime Sandboxes, and Prime Inference. This let Extropic avoid managing multi-node GPU infrastructure and focus on research.

  3. Lewis TunstallAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

  4. Cloudflare Blog · AIAI score58

    Cloudflare releases open-source Clef decision models and an RL fine-tuning service

    AICloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and launched a reinforcement learning fine-tuning service. In Cloudflare's tests, Clef classified a domain in 2.2s versus 4.7s for gpt-oss-120b, and the models are Jev-API compatible. The company is offering fine-tuning first through a forward-deployed engineering team, with a self-serve platform planned later.

  5. Anthropic ResearchAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

  6. Manus BlogAI score45

    Manus 2.0 Adds Video Editor for Creating and Editing Publishable Videos

    AIManus 2.0 introduces Video Editor, which lets users refine videos Manus generates, including changes to music, captions, and cut timing, without regenerating the entire video. The article describes Manus creating explainers, launch films, and animations from a single prompt, drawing on web search, video models such as Seedance 2.5, and code for motion graphics.

  7. LangChain BlogAI score58

    LangChain shows how to build a model router in its Open SWE coding agent

    AILangChain built a model router inside its open source coding agent Open SWE that picks one of three models for each thread. In an A/B test against always using GPT-6 Astra, the median cost per thread fell 64% with no measurable change in merged PR rate. The router runs on the thread's first message, using a base prompt, per-tier criteria, and a classifier model, and the post lists next steps including subagent routing and mid-thread re-routing.

Sep 30

Sep 30Wed
  1. Google FlowAI score38

    Google's Gemini Omni Flash guide offers prompting tips for Flow videos.

    AIGoogle Flow publishes a guide to creative prompting with Gemini Omni Flash, covering video generation for films, marketing, and visual assets. The guide recommends high-level constraints, first and last frame visual anchors, tagged image, video, and storyboard ingredients, and granular mid-scene pacing edits. It also suggests transferring style and motion from reference images and videos.

  2. O'Reilly RadarAI score45

    The Agentic Data Science Playbook: Delegating Analysis to AI Agents

    AIAgentic data science has AI agents explore datasets, choose modeling approaches, run analyses, and explain findings while data scientists frame questions and verify evidence. In an experiment, Claude Opus 5.0 given the vague prompt "Build me a model to detect fraudulent nodes" on a modified Elliptic Bitcoin dataset reported F1 0.87 and ROC AUC 0.99 using a random split that leaked a planted label proxy.

  3. Google Cloud · AI & Machine LearningAI score41

    Google Cloud Rolls Out Agent Substrate, GKE Agent Sandbox RL Tools in September

    AIGoogle Cloud introduced GKE Agent Substrate, an open-source execution runtime it says can run millions of sandboxes with 10x higher density than standard container runtimes. It also made GKE Agent Sandbox optimized for reinforcement learning generally available, alongside an orchestration SDK and native RL gym integrations. Google said GKE Pod snapshots can reduce AI inference start-up by as much as 89%, based on internal tests.

  4. Karl's AI WattsAI score38

    Can you keep your session after switching models in magpie?

    AIKarl's AI Watts asks whether a menu-bar tool can switch models while preserving the existing conversation, so users avoid re-explaining their project each time. The post frames this as the reason they want to keep the menu bar tool, which the quoted post describes as magpie, a menu-bar switcher for 20+ agents including Claude Code and Codex that also offers a local gateway.

  5. Hamel HusainAI score42

    Hamel Husain Tests Anthropic's Claude Eval Plugin on Leasing Assistant Traces

    AIHamel Husain reviewed Anthropic's new build_eval and hill-climb commands in the claude-api plugin for Claude Code, finding it useful for discovering issues like human handoff, formatting, and voice agent problems. He criticized it for pushing evaluator creation before data review, asking for label validation in Markdown files, and bundling four failure checks into one broad call-transfer evaluator. Husain says he would hold off on using it for now.

Sep 29

Sep 29Tue
  1. Google Developers BlogAI score47

    Google Details Sparse Attention Speedup for Video Diffusion on TPUs

    AIGoogle Developers Blog describes how Sparse VideoGen (SVG) routes video diffusion attention heads into spatial or temporal sparse masks and implements them as custom JAX and Pallas Splash Attention kernels on TPU v6e. In isolated single-chip tests with 75.6K tokens and 10 heads, the sparse variants retain about 38.87% of query-key pairs. The article argues that theoretical sparsity must be converted into hardware tile skipping to yield real speedups.