Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Oct 6

Oct 6Tue
  1. Epoch AIAI score60

    Epoch AI finds frontier models fall short of an end-to-end AI research task

    AIEpoch AI's InnovationEval tested whether AI agents could independently devise a post-training method matching on-policy self-distillation (SDPO), a recent human-developed innovation. GPT-5.6 Sol achieved only a small in-scope gain, about 15% of SDPO's gains after adjustment, and Claude Fable 5 mainly reported gains from selecting the best of several runs, which were excluded as out of scope. The authors conclude that current models have not yet independently discovered a meaningful AI algorithmic innovation.

    Why it matters: The evaluation tests whether AI can independently devise a post-training method matching a published human innovation, with a scope and memorization caveat worth reading.

  2. Epoch AIAI score36

    US Adults' Cyber Incident Rates Unchanged Since Claude Fable 5 Launch, Epoch AI Finds

    AIEpoch AI reports that the share of US adults reporting at least one cyber incident in the past 12 months was 45% in September, essentially unchanged from 46% in June. The poll found no detectable change among frequent AI users, who moved from 53% to 51%. Epoch notes that its polling measures ordinary Americans' experiences, separate from its documented rise in serious vulnerability disclosures and frontier-model offensive capabilities.

  3. PyTorch BlogAI score46

    PyTorch Introduces FBTriton Kernels to Speed Table Batched Embedding Operations

    AIPyTorch's blog describes a Triton-based implementation of Table Batched Embedding (TBE) forward and backward kernels for recommendation-system embedding lookups, which the post says outperforms legacy CUDA kernels on these workloads. On B200, an updated CUDA bounds-check step reaches up to 1.24x speedup on that component, and an optional forward-side preprocessing path cuts combined latency from 79.537 ms to 66.183 ms (−16.8%) on a large configuration.

  4. PrismaXAI score49

    Hand makers and Boston Dynamics steal the spotlight at IROS 2026

    AIAt IROS 2026 in Pittsburgh, at least 17 dexterous hand companies exhibited, 11 of them Chinese, with WUJI reportedly shipping 800 to 900 units a month. Boston Dynamics skipped a booth but released a video on the final day of a new four-finger, 13-degree-of-freedom Atlas hand, down from 7 DOF on its previous gripper. Hand makers are also selling capture gloves and data services, since labs need far more demonstrations than the hardware alone provides.

  5. GoogleAI score52

    Google Earth AI uses agents and satellite data to predict disease spread

    AIGoogle Earth AI combines environmental signals and other data sources with AlphaEarth Foundations, a Population Dynamics Foundation Model (PDFM), and a prototype Geospatial Reasoning agent. Researchers ask questions such as where a disease is likely to spread next, and the system automatically gathers relevant models and datasets to build a prediction model. By combining satellite views with population patterns, the tool aims to reveal hidden risk factors and identify issues earlier.

    Image from @Google's post
  6. GoogleAI score30

    Google Earth AI helps forecast disease outbreak spread faster

    AIGoogle Earth AI, according to new research, can help communities respond to public health crises more quickly and proactively. The post says it combines behavioral trends, geospatial AI models, and other insights beyond simple statistics to help public health teams understand complex issues and bridge reporting gaps. The aim is to shift emergency response from reactive management toward proactive prevention.

    Image from @Google's post
  7. Azure BlogAI score22

    Microsoft Named a Leader in 2026 Gartner Magic Quadrant for Industrial AIoT Platforms

    AIMicrosoft has been named a Leader in the 2026 Gartner Magic Quadrant for Global Industrial AIoT Platforms. The company says its Azure platform, including Azure IoT, Azure Arc, Microsoft Fabric, and Microsoft Foundry, connects cloud and edge operations to apply AI-powered reasoning and close the loop between insight and action.

  8. NVIDIA Technical BlogAI score37

    Scale Bitwise-Deterministic Pretraining with NVIDIA Megatron Core

    AINVIDIA's technical blog describes bitwise determinism for large-scale pretraining with Megatron Core, which makes training runs easier to debug, validate, and resume reproducibly. The source says these benefits matter most for models with trillions of parameters trained across thousands of GPUs, where multiple parallelism dimensions, low-precision computation, and distributed checkpointing complicate failure reproduction and fix validation.

  9. elvisAI score41

    Parsewave audit fixes 206 verifier bugs in AutomationBench

    AIParsewave audited all 600 public tasks in Zapier's AutomationBench and human review confirmed 206 real verifier bugs, all of which were fixed in AutomationBench Verified. Replaying 1,235 Kimi K3 runs on the old and fixed verifiers changed 27.9% of grades, with pass rate rising from 18.8% to 43.8% where verifiers were too strict and falling from 60.2% to 49.7% where they were too lenient.

  10. Ai2AI score4

    Ai2 Agents post-training team invites COLM 2026 attendees to connect

    AIAi2 applied scientist Shashank Gupta says he will attend COLM 2026 from Tuesday through Friday and invites people to talk with the Ai2 Agents post-training team. Topics include post-training for coding and long-horizon agents, such as agentic RL, OPD, data and infrastructure, and multi-agent training, plus opportunities at Ai2. He lists an Ai2 booth session Tuesday 1:30–3pm and an Ai2 mixer Tuesday 6–9pm.

  11. Microsoft ResearchAI score36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    AIMicrosoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

  12. Google ResearchAI score51

    Google's PDFM location embeddings improve five global public health tasks

    AIGoogle Research reports that Population Dynamics Foundation Model (PDFM) embeddings, built from search trends, mobility, built environment, and weather signals, were tested by partners across five public health tasks. The embeddings improved results in cross-border MMR vaccination coverage, dengue forecasting, postpartum depression screening, and cholera outbreak prediction, and matched census inputs for cardiovascular mortality nowcasting.

  13. SemiAnalysisAI score18

    ClusterMAX rates FarmGPU underperform on Slurm and Kubernetes testing

    AISemiAnalysis rated FarmGPU as ClusterMAX Underperform after its Slurm layer failed to advertise GPU resources and Kubernetes exposed no RDMA devices for scale-out networking. The post credits FarmGPU's Grafana monitoring, provisioning notes, and trustworthy technical team, while noting the team may be stretched thin across small clusters.

    Image from @SemiAnalysis_'s post
  14. The Next PlatformAI score20

    When One Datacenter Is No Longer Enough: Cisco on Scale-Across AI Networking

    AICisco SVP Rakesh Chopra discusses the "Scale-Across" approach to networking AI training workloads spread across multiple data centers. He describes how Silicon One architecture and Intelligent Collective Networking aim to manage synchronized GPU traffic over long-distance fiber links. The interview covers power efficiency and hardware-accelerated MACsec and IPsec security.

  15. Sophia YangAI score26

    Reinforcement learning infrastructure scales to tens of thousands of parallel rollouts

    AIThe post describes a reinforcement learning system that autoscales an actor fleet to run tens of thousands of rollouts in parallel with asynchronous training, designed for trajectories of millions of tokens with multiple compactions and low staleness. New methods at both stages reduce off-policy drift, and the setup runs on 3k GPUs producing about 33B tokens per day, with roughly 16B trainable after filtering and masking. Rewards rise across representative environments as the policy learns harder tasks.

    Image from @sophiamyang's post
  16. Guillaume Lample @ NeurIPS 2024AI score26

    Mistral's ML4 trained on 3,800 NVIDIA Grace Blackwell GPUs in Europe

    AIMistral says its ML4 model was trained on 3,800 NVIDIA Grace Blackwell GPUs in its European datacenters, including its Bruyères-le-Châtel cluster built with Series B funding. The company is investing further, with Series C and D clusters coming online soon to support longer training, more ambitious post-training, and faster iteration. Mistral expects large and rapid improvements in the weeks and months ahead.

    Image from @GuillaumeLample's post
  17. NVIDIA BlogAI score32

    Telecom Operators Build AI Strategies on Open Models, Citing Control and Customization

    AITelecom operators are building AI strategies on open models for reasons beyond cost, including control, customization, and trust across workloads from autonomous networks to customer care. NVIDIA's State of AI in Telecommunications report found 89% of respondents say open source models and software are important to their company's AI strategy. The NVIDIA Nemotron family offers open weights, training data, and recipes, and the 30-billion-parameter Nemotron 3 Large Telco Model was fine-tuned by AdaptKey on open telecom datasets.

  18. The Next PlatformAI score38

    Dell Adds Data Context, Prep, and Storage Features to Its AI Data Platform

    AIDell is adding agentic AI capabilities to its AI Data Platform, including a Unified Semantic Layer with a searchable glossary and an Enterprise Knowledge Graph built with Nvidia's Auto-Ontology open source library. The features are designed to give agents shared context, reducing repeated token generation and compute costs. The platform's layers include the Data Orchestration Engine, Data Engines, and Storage Engines such as PowerScale, ObjectScale, and the Lightning File System.

  19. ElevenLabs BlogAI score21

    What Conversation Intelligence Is and How Businesses Can Use It

    AIConversation intelligence records and transcribes sales and support calls, then uses AI to tag sentiment, objections, and action items for team-wide review. The guide explains how the pipeline works, from data capture and transcription to analysis and CRM sync. It also outlines benefits such as faster coaching and less manual data entry.

  20. ChinaTalkAI score33

    Bharat Patel on why data, not models, is the hard part of military AI

    AIAccenture defense AI lead Bharat Patel argues that data quality depends on the use case and that "AI-ready data" is a myth. He cites Project Maven, which began in 2017, where early imagery lacked relevant targets and models underperformed until teams continuously collected targeted data. The conversation also covers why fully autonomous tanks remain distant and the risks of data poisoning.

  21. IThome · AIAI score41

    Strata engine runs 125B Qwen3.8 model on 12GB GPU at 94 tokens/s

    AIDeveloper Niko1221 has open-sourced Strata, an engine that runs a quantized 125B-parameter Qwen3.8-Flash-Next model on consumer GPUs with at least 12GB of VRAM. Strata loads the MoE model into RAM and keeps only frequently used experts in VRAM, and uses a lightweight model for speculative decoding. On an NVIDIA RTX 5070 with 12GB VRAM, the Q2_0 quantization reaches 94 tokens per second for output.

Oct 5

Oct 5Mon
  1. Epoch AIAI score43

    Epoch AI finds China more exposed than US to chip supply shocks

    AIChina is more exposed than the US to semiconductor supply disruptions, with semiconductor producers earning $15.2 per $1,000 of Chinese final demand in 2022 versus $5.7 for US spending. In a combined Taiwan disruption and China–West decoupling scenario, Chinese advanced processor prices rise 17-fold and real gross national expenditure falls 3%, compared with about a 20% price rise and 0.6% fall for the US. The authors report the gap persists across robustness checks, though the exact size carries significant uncertainty.

  2. Apple Machine Learning ResearchAI score23

    RISED uses rubrics to guide multi-environment LLM agent training and data selection

    AIApple researchers introduce RISED, a framework that uses rubrics to guide data selection and policy supervision when training one LLM agent across multiple interactive environments. An LLM judge tags rollouts with a shared rubric vocabulary, positive rubrics provide privileged context for an on-policy self-distillation teacher, and negative rubrics steer generation away from recurring failures. The authors report that RISED achieves the highest mean pass rate across environments and ranks first or second in each environment, across model backbones.

  3. Mike KnoopAI score62

    Dust pretrains transformers with zeroth-order optimization, approaching backprop results

    AIDust is a zeroth-order method that pretrains transformers and sometimes matches or exceeds backprop given large compute. The authors report it is about 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. The post also cites the gradient-alignment result up to 1B tokens and the virtual population idea for scaling.

  4. PyTorch BlogAI score40

    PyTorch Consolidates Media Decoding and Encoding Into TorchCodec, Narrows TorchVision and TorchAudio

    AIPyTorch has consolidated all media decoding and encoding for images, video, and audio into TorchCodec, which now runs on CPU and CUDA. TorchVision and TorchAudio are narrowed to focus on their transforms, with models, datasets, and pipelines no longer under active development. All three libraries are now ABI stable and no longer need rebuilding for each PyTorch release.

  5. Harrison ChaseAI score50

    Cognition's Devin adds "Dreaming" offline memory cleanup, open-sourced as a standard

    AIHarrison Chase praises Cognition's "Dreaming" feature, which lets Devin clean stale memory records and surface latent information offline. He argues agent memory needs an offline cleanup loop rather than only better retrieval, and questions how inferred memories get validated before use. He also welcomes Cognition's plan to release Agent Memory Repo as an open standard.

  6. SemiAnalysisAI score10

    Classifiers map inputs to fixed labels via encoders and softmax or sigmoid

    AIA classifier assigns an input to a fixed label set, covering binary, multiclass, and multilabel variants, such as spam versus not spam or movie genres. It encodes the input into a vector using hand-built features like logistic regression or a learned encoder such as a CNN or BERT. A linear layer then projects that vector into K logit scores, which softmax or sigmoid turns into probabilities.

    Image from @SemiAnalysis_'s post
  7. ReflectionAI score23

    Reflection AI's Beam model pretrained in four weeks on 24T tokens

    AIReflection AI says its Beam model was pretrained in 4 weeks on 24T high-quality tokens, giving it innate coding capabilities. The company credits MoE stability improvements and large-scale data curation and deduplication for a base model it claims outperforms open-source base models of the same class. It presents this strong reasoning foundation as what makes sustained reinforcement learning gains possible.

    Image from @reflection_ai's post
  8. ReflectionAI score44

    Reflection scales Beam on 10.5k GB300s in record RL run

    AIReflection says it ran Beam, its reinforcement learning system, on 10.5k GB300 GPUs for four weeks, which it describes as the largest publicly documented RL run it knows of. The company credits algorithmic advances combined with distributed infrastructure for making the system scale. Across its eval suite, capabilities kept improving as RL increased, with no sign of a plateau.

    Image from @reflection_ai's post
  9. Google AIAI score46

    Gemma 4 and BOTANIC-1 pinpoint crop-yield DNA mutations in minutes

    AILiving Models paired Google's Gemma 4 with BOTANIC-1, a plant-DNA model trained on 320 species, to identify causal genetic variants. In a melon yield test, the pipeline ranked the target mutation first out of 2,494 possibilities in under four minutes. The approach aims to speed up breeding of climate-resilient crops that would otherwise take years of field trials.