Skip to contentSkip to stories

Updated

#Tutorial/How-to

Showing low-relevance items too. Hide low-relevance items

Oct 1

Oct 1Thu
  1. Harrison ChaseXAI score22

    Harrison Chase outlines a four-step approach to model routing

    AIHarrison Chase says model routing is a provocative term that lacks a clear definition, but offers a practical approach. His four steps are to understand tasks, understand the models, build the router inside the harness, and track outcomes, aiming to lower costs without a performance hit.

    Image from @hwchase17's post
  2. Google WorkspaceOfficialAI score32

    Google Sheets canvas turns spreadsheets into interactive mini-apps via prompts

    AIGoogle Workspace says Sheets canvas can turn static spreadsheet data into interactive tools such as Kanban boards, dashboards, and visual workflows from a simple prompt. Derek Snyder, Director of Product Marketing for Google Workspace, demonstrates the feature in the latest AI Boost Bite video.

    Video from @GoogleWorkspace's post
  3. Prime IntellectOfficialAI score12

    Extropic and Prime Intellect Divide Roles in an RL Training Setup

    AIExtropic designed the tasks and reward, while Prime Intellect supplied the RL infrastructure, including verifiers for environment construction, Hosted Training for the RL loop, and Prime Sandboxes for executing model code. Prime Inference serves the LLM judge and frontier baselines in the same workflow.

    Image from @PrimeIntellect's post
  4. Prime IntellectOfficialAI score32

    Extropic uses Prime Intellect to post-train Qwen3.6-35B-A3B for thermodynamic ML

    AIExtropic post-trained Qwen3.6-35B-A3B with Prime Intellect for thermodynamic ML research, nearly tripling its held-out eval results in about 100 GRPO steps. The team built a custom RL environment with verifiers and trained on Hosted Training, Prime Sandboxes, and Prime Inference. This let Extropic avoid managing multi-node GPU infrastructure and focus on research.

    Image from @PrimeIntellect's post
  5. Lewis Tunstall @ COLM 🌉XAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

    Video from @_lewtun's post
  6. Cloudflare Blog · AIOfficialAI score58

    Cloudflare releases open-source Clef decision models and an RL fine-tuning service

    AICloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and launched a reinforcement learning fine-tuning service. In Cloudflare's tests, Clef classified a domain in 2.2s versus 4.7s for gpt-oss-120b, and the models are Jev-API compatible. The company is offering fine-tuning first through a forward-deployed engineering team, with a self-serve platform planned later.

  7. Hamel HusainXAI score25

    Hamel Husain says not every failure mode needs an automated evaluator

    AIHamel Husain advises against building automated evaluators for every failure mode discovered during AI development. He argues that teams should weigh the costs of different evaluator types, such as code-based checks versus LLM judges, before building an eval.

    Image from @HamelHusain's post
  8. merveXAI score4

    Hugging Face points to four YouTube video series for learning

    AIHugging Face's Merve Noyan says the company already offers four video series on YouTube for viewers who want to learn how to do it. She links to one playlist in the post, but the post does not specify which topic the series cover.

  9. Anthropic ResearchOfficialAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

  10. Manus BlogOfficialAI score45

    Manus 2.0 Adds Video Editor for Creating and Editing Publishable Videos

    AIManus 2.0 introduces Video Editor, which lets users refine videos Manus generates, including changes to music, captions, and cut timing, without regenerating the entire video. The article describes Manus creating explainers, launch films, and animations from a single prompt, drawing on web search, video models such as Seedance 2.5, and code for motion graphics.

  11. LangChain BlogOfficialAI score58

    LangChain shows how to build a model router in its Open SWE coding agent

    AILangChain built a model router inside its open source coding agent Open SWE that picks one of three models for each thread. In an A/B test against always using GPT-6 Astra, the median cost per thread fell 64% with no measurable change in merged PR rate. The router runs on the thread's first message, using a base prompt, per-tier criteria, and a classifier model, and the post lists next steps including subagent routing and mid-thread re-routing.

Sep 30

Sep 30Wed
  1. Hamel HusainXAI score4

    Hamel Husain thanks Lance Martin for updating Claude eval skill

    AIHamel Husain posted a short thank-you emoji reply to Lance Martin's update on an evaluation skill for Claude. Martin says the latest skill now instructs Claude to build a viewer for eval examples, but it does not walk users through the data first, which he agrees could help them prioritize which evals to write.

  2. Hamel HusainXAI score38

    Hamel Husain Reviews Claude's New Auto Eval Plugin for Evaluations

    AIHamel Husain has published a longer review of a new Claude Auto Eval plugin after many users asked about it. He invites readers to share their experiences using the plugin and how it went for them. The plugin is part of Claude's ability to help build evaluations and hillclimb on them, as described by @ClaudeDevs.

    Image from @HamelHusain's post
  3. Google FlowOfficialAI score38

    Google's Gemini Omni Flash guide offers prompting tips for Flow videos.

    AIGoogle Flow publishes a guide to creative prompting with Gemini Omni Flash, covering video generation for films, marketing, and visual assets. The guide recommends high-level constraints, first and last frame visual anchors, tagged image, video, and storyboard ingredients, and granular mid-scene pacing edits. It also suggests transferring style and motion from reference images and videos.

  4. Google Cloud TechOfficialAI score28

    Agent Clinic Ep 3 builds automated eval suite for LangGraph agent

    AITerminal test runs miss multi-turn agent regressions, so Agent Clinic Episode 3 builds an automated eval suite for a LangGraph agent in 60 minutes. The post presents a four-step framework for moving from informal checks to benchmarking AI agents, with a link to the full guide.

    Image from @GoogleCloudTech's post
  5. Google Cloud TechOfficialAI score15

    Google Cloud tips for capping GPU and replica settings to control costs

    AIGoogle Cloud recommends limiting accelerator count to a single GPU, setting replica count to 1-1, and avoiding capacity reservations to keep monthly bills predictable. These strict hardware limits apply to auto-scaling configurations for AI workloads.

  6. NVIDIA AIOfficialAI score40

    NVIDIA Shows Visual AI Agent Built in Under 30 Minutes

    AINVIDIA says a single prompt can build and deploy a visual AI agent for a manufacturing line in under 30 minutes, with alerts, video search, and incident reports. The method uses the new Build Vision AI skill in NVIDIA VSS Blueprint 3.3, and a tutorial is available for readers who want to build one.

    Video from @NVIDIAAI's post
  7. Ant LingOfficialAI score38

    Ling-3.1-flash ports C image library to Rust with 8.015× speedup

    AIAnt Ling reports that its Ling-3.1-flash model completed a roughly 20-hour Rust port of a C image library. After a performance regression caused by busy-waiting workers and a parallelism adjustment, the model recovered and reached an 8.015× speedup. All 30 correctness checks passed.

    Image from @AntLingAGI's post
  8. NVIDIA AIOfficialAI score27

    NVIDIA NeMo Relay Traces Hermes Agent Runs in Arize Phoenix

    AINVIDIA and Nous Research published a hands-on walkthrough of NVIDIA NeMo Relay for collecting traces from Hermes Agent. The guide runs two example scenarios and shows the agent's calls and retries in Arize Phoenix. It also covers how Nous used traces and task results to evaluate fixes across repeated runs.

    Video from @NVIDIAAI's post
  9. O'Reilly RadarBlogAI score45

    The Agentic Data Science Playbook: Delegating Analysis to AI Agents

    AIAgentic data science has AI agents explore datasets, choose modeling approaches, run analyses, and explain findings while data scientists frame questions and verify evidence. In an experiment, Claude Opus 5.0 given the vague prompt "Build me a model to detect fraudulent nodes" on a modified Elliptic Bitcoin dataset reported F1 0.87 and ROC AUC 0.99 using a random split that leaked a planted label proxy.

  10. Google Cloud · AI & Machine LearningOfficialAI score41

    Google Cloud Rolls Out Agent Substrate, GKE Agent Sandbox RL Tools in September

    AIGoogle Cloud introduced GKE Agent Substrate, an open-source execution runtime it says can run millions of sandboxes with 10x higher density than standard container runtimes. It also made GKE Agent Sandbox optimized for reinforcement learning generally available, alongside an orchestration SDK and native RL gym integrations. Google said GKE Pod snapshots can reduce AI inference start-up by as much as 89%, based on internal tests.