Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Oct 2

Oct 2Fri
  1. PyTorch BlogAI score47

    Helion Linear Backend Boosts vLLM Hopper GPU Inference Throughput Over CUTLASS and DeepGEMM

    AIThe vLLM team integrated Helion, a PyTorch-native kernel DSL, into vLLM's linear backend, using per-shape autotuning to select among Standard GEMM, Split-K, and Swap-AB variants. On NVIDIA Hopper GPUs, the Helion backend outperformed the default CUTLASS and DeepGEMM backends across the evaluated models, with more than 10% throughput gains for some workloads. The work focuses on FP8 and INT8 quantized GEMM.

  2. Epoch AI · The Epoch BriefAI score62

    Epoch AI estimates 2026 compute could run hundreds of millions of AI agents

    AIEpoch AI estimates that compute built from projected 2025 to 2027 high-bandwidth memory shipments could support tens to hundreds of millions of frontier AI agents, or billions of cheaper ones. Running nonstop, the top-tier agents would match the working hours of 140 million to 700 million full-time employees, and the central DeepSeek V4 Pro estimate of about 1.9 billion agents would match 8 billion workers.

    Why it matters: The estimate converts memory shipments into agent capacity and revenue ranges, showing how hardware supply could translate into labor and sales if demand keeps up.

  3. Nathan LambertAI score35

    Nathan Lambert launches Trillium Labs, a nonprofit for open frontier AI science

    AINathan Lambert and Tom Zick have unveiled Trillium Labs, a new non-profit focused on the open science of frontier AI. The lab plans to build open post-training recipes and expand into open infrastructure to study topics such as RSI, reward hacking, and multi-agent systems. It is hiring, fundraising, and seeking compute, with support from Halcyon Futures and Schmidt Sciences.

    Image from @natolambert's post
  4. Google AIAI score62

    Google launches Project Suncatcher prototype satellite to test TPUs in orbit

    AIGoogle AI announced that its Project Suncatcher prototype satellite, built with Planet, has launched into orbit on SpaceX's Transporter-18 rideshare mission. The initial mission will gather data on how Google TPUs handle the physical stress and extremes of spaceflight. The post says low Earth orbit systems could generate up to 8x more solar power than on Earth, and that future work may link multiple satellite constellations for scaled machine learning.

    Video from @GoogleAI's post
  5. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    AIAi2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    Why it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  6. Google ResearchAI score60

    Google's TEE-based federated learning system adds verifiable privacy guarantees

    AIGoogle announces a next-generation federated learning system that uses Trusted Execution Environments to provide verifiable, auditable data anonymization. The system publishes access policies to a public transparency log and is deployed in Gboard, which has launched English and Japanese next-word prediction models with stronger privacy guarantees and improved accuracy. Training time has also sped up significantly because computation moved to the server and is parallelized across many machines.

    Why it matters: The post shows how Trusted Execution Environments make federated learning's privacy claims externally verifiable, rather than relying on trust in the server operator.

  7. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

    Image from @huggingface's post
  8. Ai2 (Allen Institute for AI)AI score67

    Ai2 open-sources AstaBrief 8B, a fast open-weights scientific report model

    AIAi2 released AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. In Asta's Generate a report feature, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The model is built on Qwen3-8B with supervised fine-tuning and DPO, and institutions can run its open weights on their own infrastructure.

    Why it matters: The post explains the data filtering and one-pass generation choices behind a fast open-weights report model, showing what worked and what did not.

  9. TinkerAI score33

    Tinker praises Fulcrum's cheap, effective style-customization training approach

    AITinker says Fulcrum trains its Echo writing model by building on a base model that already writes well, tailoring both SFT and RL to separate the default LLM voice from authors' voices. The post calls this customization approach both cheap and effective. Fulcrum says Echo beats frontier models at writing tasks such as fiction and technical explanations, at a training cost under $5K.

  10. Hugging Face BlogAI score62

    AutoSynthData generates targeted training data for enterprise agents from failures

    AIServiceNow CoreAI introduced AutoSynthData, which uses a target model's failures and a stronger teacher's successes to generate and validate new agent training tasks. In EnterpriseOps Gym experiments, the Hybrid domain produced 2,000 samples and raised Gemma-4-26B-A4B-it mean Pass@1 by 7.2 percentage points, while the ITSM domain produced 1,994 samples and raised it from 18.77% to 27.18%.

    Why it matters: The post shows how failure analysis, teacher demonstrations, and verifier checks combine into a repeatable pipeline for generating targeted agent training data.

Oct 1

Oct 1Thu
  1. NVIDIA AIAI score44

    CoreWeave RL rollouts reload model weights 15× faster with Dynamo

    AICoreWeave's new RL rollouts service uses ModelExpress and Router in NVIDIA Dynamo to speed up model weight reloads during RL post-training with minimal downtime. Working with NVIDIA and You.com, CoreWeave achieved 15× faster model reloads than its baseline while post-training Nemotron 3.5 Lightning. The speedup addresses GPUs sitting idle while inference workers wait to load updated weights between training iterations.

  2. Apple Machine Learning ResearchAI score28

    Language Discrimination Narrows Multilingual Speech Model Gap, Study Finds

    AIResearchers Maureen de Seyssel, Jie Chi, and Zakaria Aldeneh found that strengthening language discrimination during pretraining reduces the performance gap between multilingual and monolingual HuBERT speech models. In a controlled English/French setting, phone-ABX error fell from 11.6% to 10.4%, close to the monolingual 10.8%, while lexical sWUGGY scores rose from 52.1% to 56.7%. The gains were largest when language discrimination was introduced in the first training iteration.

  3. Epoch AIAI score62

    Epoch AI estimates how many concurrent AI agents 2025–27 memory shipments could run

    AIEpoch AI estimates that high-bandwidth memory shipped in 2025–27 could eventually support about 30–170 million concurrent frontier-model agents once fully deployed and allocated. Using DeepSeek V4 Pro serving benchmarks, the estimate rises to about 1.9 billion concurrent agents. The authors compare the implied API-equivalent spending of $2.6–5.3 trillion per year with projected developer revenue of roughly $1 trillion by end-2027, suggesting demand may lag supply.

    Why it matters: The analysis converts HBM shipment data into concurrent agent capacity and compares it with projected API revenue, showing where compute buildout may outpace demand.

  4. Apple Machine Learning ResearchAI score34

    Limits of Confidence-Based Sampling in Discrete Diffusion Models

    AIApple Machine Learning Research reports that discrete diffusion steps match the training distribution only when simultaneously written token positions are conditionally independent given already-fixed tokens. The authors show that per-position distributions cannot determine such dependence, and on the synthetic ScanAndAdd task, confidence-ranked groups of two or more positions were dependent and produced a generated distribution 29 times the sampling-noise floor in total variation.

  5. Google · Innovation & AIAI score56

    Google's Project Suncatcher prototype satellite launches into orbit with Planet

    AIGoogle's prototype satellite for Project Suncatcher, built with Planet, launched into orbit on the Transporter-18 rideshare mission with SpaceX. The team confirmed contact and says the satellite is operating as expected. Over the coming weeks, it will gather in-orbit data on how Google's TPUs handle spaceflight stress, radiation, and thermal extremes, and a peer-reviewed paper detailing the research is available in Joule.

  6. Sophia YangAI score38

    Fireworks details numerical mismatch fixes for stable RL training

    AIFireworks reports that numerical mismatch between training and rollout engines can destabilize reinforcement learning, with a GLM 5.2 experiment showing collapsing reward without alignment and stable reward with it over 25 steps. The post notes that MoE models add further alignment challenges, as Qwen3.5-MoE differences in expert output combination caused disagreement even when one implementation used higher precision. Fireworks says it co-develops its trainer and rollout engine to keep frontier RL training aligned across numerics, kernels, and MoEs.

  7. Goodfire ResearchAI score60

    Goodfire proposes protein embedding monitors for biosecurity risks in AI agents

    AIGoodfire Research developed sequence-aware monitors using protein language model embeddings to flag concerning biological sequences in dual-use AI agent tasks. On a custom benchmark, the monitors outperformed frontier model safeguards with fewer refusals on benign requests, and they held up better against paraphrasing and fragmentation attacks. The paraphrase results rely on in-silico estimates and do not establish whether the redesigned proteins keep biological activity, and the monitors run in milliseconds per sequence.

    Why it matters: The post gives a concrete benchmark setup and fragmentation results, showing how sequence embeddings can separate dual-use biology requests that task-based safeguards handle poorly.

  8. Prime IntellectAI score34

    Qwen3.6 reward rises 2.8x via GRPO on Hosted Training

    AIPrime Intellect reports that after about 100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain. Qwen3.5, trained the same way, reached 0.356, suggesting the method works across model families. Both post-trained models finished well ahead of other open models and narrowed the gap to Claude Opus 4.8, with Qwen3.6 activating only 3B parameters per token.

    Image from @PrimeIntellect's post
  9. Yellowbrick InvestingAI score18

    Yellowbrick 2.0 launches with leaderboards, API access, and custom feeds

    AIYellowbrick 2.0 is live, tracking 35,000+ stock pitches from 4,000+ authors and adding 300+ new pitches weekly. The rebuilt platform adds author leaderboards, custom feeds and alerts, API access, and paid research partner discounts. Premium subscribers get a 30% discount on Koyfin, which the company says covers the cost of Yellowbrick Premium.

  10. Jerry LiuAI score42

    LlamaIndex launches Extract v2.5 document extraction agents with improved accuracy

    AILlamaIndex introduced Extract v2.5, a series of agents tuned for document extraction across cost-effective, agentic, and agentic plus tiers. The company reports the agents outperform Opus 5.5 and GPT-6 Sol while costing 30% to 4x less, with accuracy gains on long lists (86.1% to 95.5%), multi-page records (85.5% to 96.5%), and scanned forms (90.9% to 95.7%) on its agentic tier. The release adds advanced citations with bounding boxes and structural reasoning, and the agents are available on LlamaParse.

    Video from @jerryjliu0's post
  11. merveAI score46

    Hugging Face clarifies ml-intern options, one trained model for $6

    AIHugging Face says ml-intern is an open-source ML engineering and research harness usable free on local setups, and it is also hosted on Hugging Chat with no-code access. A second hosted option runs on Hugging Face infrastructure, where ml-intern selects the cheapest GPU for a task so models can be trained for a few dollars. MaziyarPanahi reportedly trained a model by prompting alone for $6.60 on an NVIDIA A100 in 16 minutes.