Skip to contentSkip to stories

Updated

#Data/Training

Sep 8

Sep 8Tue
  1. John SchulmanAI score40

    Schulman distinguishes risks of training AI on user data

    AIJohn Schulman argues that training on user data carries very different privacy and IP risks depending on method. Pretraining on user tokens poses high regurgitation risk, while distillation from prompts and RL from user traces carry lower regurgitation risk but can still leak customer IP. He notes de-identification is weak because long traces can still identify users, and AI companies rarely disclose what they do.

  2. Dwarkesh PodcastAI score62

    Data improvements drove more pretraining efficiency gains than model changes from 2019 to 2025

    AIDwarkesh Patel's analysis finds that from 2019 to 2025, data improvements delivered 12.0x compute efficiency gains versus 3.7x for model improvements at the 1e19 FLOPs budget. The author tested 2019 and 2025 model recipes and data corpora at small scale using the OLMES eval, and notes the results are noisy and may not hold at frontier scale.

  3. Google DeepMindAI score74

    Google DeepMind launches AlphaGenome Atlas to predict 9 billion DNA variant effects

    AIGoogle DeepMind has introduced AlphaGenome Atlas, a platform with predicted molecular effects for 9 billion single-nucleotide variants in the human genome. It is free for academic research through a web portal, and the AlphaGenome Variant Impact score condenses predictions from AlphaGenome and AlphaMissense into one number for ranking variants. The source says collaborators used it to identify variants in unsolved rare disease cases and to find rare non-coding variants linked to traits.

    Why it matters: The source details how precomputed variant predictions, a single impact score, and linked feature attributions make genome-wide mutation effects searchable for researchers without coding skills.

  4. Google DeepMind · The KeywordAI score72

    Google DeepMind launches AlphaGenome Atlas, a database of DNA variant effect predictions

    AIGoogle DeepMind has released AlphaGenome Atlas, a web portal that predicts the regulatory effects of all 9 billion possible single-letter genetic changes in the human genome. The Atlas provides an AlphaGenome Variant Impact (AVI) score that combines coding and non-coding predictions to help researchers prioritize variants. The source says the portal requires no coding skills and is available to researchers and biologists worldwide.

    Why it matters: The source details how the Atlas's AVI score is used in real rare disease and UK Biobank analyses, showing a practical route for prioritizing non-coding variants.

  5. Google DeepMind · YouTubeAI score78

    DeepMind releases AlphaGenome Atlas, a predictive map of every possible DNA letter change

    AIGoogle DeepMind has used AlphaGenome to predict the molecular impact of every possible single-letter change in the human genome, around nine billion variants. The resulting AlphaGenome Atlas is a 1PB dataset that assigns each variant an AlphaGenome Variant Impact (AVI) score, covering both coding and non-coding variations, and is available to researchers worldwide. The video notes that AlphaGenome has not been validated or approved for any clinical use.

    Why it matters: The release supplies a precomputed impact score for every possible single-letter genome change, which lets researchers look up variants without running the model themselves.

  6. NVIDIA · new models on Hugging FaceAI score46

    NVIDIA Releases NV-Reason-CT, a 3D Vision-Language Model for Chest and Abdominal CT

    AINVIDIA's NV-Reason-CT is a 3D vision-language model for CT image analysis that combines a native 3D vision encoder with a language model. It is designed for radiology report generation, question answering, and multi-step reasoning across chest and abdominal CT volumes. The model converts a 384×384×384-mm input into 13,824 visual tokens without spatial downsampling and is available on Hugging Face under the OpenMDW-1.1 License.

Sep 7

Sep 7Mon
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score45

    openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning

    AIOpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning. The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline. The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.

Sep 6

Sep 6Sun
  1. Sebastian RaschkaAI score22

    Raschka's Reasoning From Scratch video covers LLM text generation and KV caching

    AISebastian Raschka released a video in his Reasoning From Scratch series covering text generation in LLMs and KV caching. The walkthrough uses a pretrained Qwen3 model from the Reasoning From Scratch package, covering tokenization, greedy decoding, end-of-sequence handling, and a benchmarked KV caching speedup. It prepares the base model for reasoning techniques in later episodes.

  2. OpenBMB (MiniCPM) · new models on Hugging FaceAI score35

    UltraData-Code-L2-Classifier scores files for algorithmic code selection

    AIOpenBMB released UltraData-Code-L2-Classifier, a suite of language-specific file-level scorers for 11 programming languages in UltraData-Code-L1. The L2 corpus selected with these scorers contains approximately 400B tokens and retains about 12.23% of L1 files, and a 10B-token test on a 1B model raised EvalPlus pass@1 by 7.80 points over L1 training.

Sep 5

Sep 5Sat
  1. AI at MetaAI score43

    AIRA₃ coordinates long-running agents through a shared forum and filesystem

    AIMeta's AIRA₃ replaces a central controller with many long-running agents, each pairing a model with a coding harness in its own isolated environment. The agents coordinate asynchronously through a shared forum for hypotheses and findings and a shared filesystem for solution artifacts. According to the post, performance gains compound over time as agents build on each other's discoveries.

Sep 4

Sep 4Fri
  1. Lewis TunstallAI score46

    Meta paper uses research preference models to guide AI agents' experiments

    AILewis Tunstall praises a new Meta paper on research preference models (RPMs), which instill "research taste" in agents by treating experiments as tree nodes. An RPM acts as an LLM judge that selects the most promising candidate experiment before it is run, reducing wasted compute. Tunstall notes the resulting trajectories could train domain-specific RPMs, which would be valuable in hard fields such as the natural sciences.

  2. Tencent · new models on Hugging FaceAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

Sep 3

Sep 3Thu
  1. Understanding AI (Timothy B. Lee)AI score43

    Robot startups are trying everything they can think of to get more data

    AIRobot startups are racing to collect training data, from companies paying cleaners to wear cameras to firms recording VR-controlled humanoid robots. The article says the largest openly available robot task dataset, ABC-130K, contains only 3,500 hours of demonstrations. Skild CEO Deepak Pathak argues companies must gather high-quality data before robots can do enough useful work to generate it through deployment.

  2. Google DeepMind · The KeywordAI score72

    Google DeepMind releases WeatherNext 3, a global weather model with hourly satellite-based forecasts

    AIGoogle DeepMind and Google Research introduced WeatherNext 3, which generates hourly global forecasts at up to 5-kilometer resolution using live geostationary satellite data. The company reports that precipitation forecasts improved by up to 60% against IMERG in medium-range evaluations, and that longer-range precipitation forecasts are up to 50% more accurate. The model is now available across Search, Gemini, Google Maps, Google Maps Platform Weather API, Google Earth Engine, BigQuery, and Google Cloud Storage.

    Why it matters: The post explains how training on live satellite data and station observations changes resolution and update frequency, with precipitation accuracy gains reported against named baselines.

  3. Google DeepMind · YouTubeAI score72

    Google DeepMind's WeatherNext 3 offers hourly, 5km-resolution weather forecasts

    AIGoogle DeepMind introduced WeatherNext 3, a weather forecasting model that learns directly from satellite feeds and ground-level weather station data. It produces a fresh forecast every hour, compared with the six-hour refresh typical of traditional models, with native 5km resolution for temperature and humidity. It is available through Google Search, Gemini, Google Maps and more.

    Why it matters: The source shows a shift from six-hourly to hourly refresh and 5km local resolution, which matters for energy planning and local forecasting.

  4. Prime Intellect BlogAI score59

    Prime Intellect rebuilds GLM-5.2 RL weight transfer on NIXL, cutting sync to 3.9 seconds

    AIPrime Intellect reports that rebuilding RL weight transfer for GLM-5.2 on NIXL and ModelExpress cut sync time from 86.1 seconds with NCCL to 3.9 seconds in its fastest setting. The method traces vLLM's loader to find each tensor's runtime layout, then reads only the needed source bytes over RDMA and replays the rest locally. Most remaining latency comes from vLLM's pause consensus, which the team reduced by syncing every wave instead of every 32.

Sep 2

Sep 2Wed
  1. The Register · AIAI score39

    AI Models Misidentify Mushrooms in Test, Sometimes Calling Deadly Species Edible

    AIPiotr Migdał tested 16 AI models on 1,040 mushroom photos covering 55 species, and the best, Gemini-3.8-flash, was correct on its first guess only 65 percent of the time. Dangerous mistakes were common, with the death cap called edible 16 percent of the time, and Qwen3.8-27b wrongly labeled poisonous mushrooms edible 36 percent of the time. Migdał warns users not to eat any mushroom because an AI says it is safe.

  2. Engineering at MetaAI score55

    Meta details an AI agent that learns from expert corrections without retraining

    AIMeta Engineering describes an AI agent for a compliance domain that stores expert knowledge in structured, auditable files and separates it from reasoning procedures called recipes. Expert feedback is diagnosed, compiled into verified text edits, tested against regression suites, and reviewed by humans, all without retraining the underlying model. Meta reports that domain experts rated outputs useful almost all the time and that assessment time fell from days to minutes.

Sep 1

Sep 1Tue
  1. Ai2 (Allen Institute for AI)AI score38

    Ai2 Panel Identifies Five Hard Challenges for AI-Assisted Science

    AIAt an August 27 Ai2 event on expanding its work with the Paul G. Allen Research Center at Providence Swedish Cancer Institute, panelists identified five persistent challenges for scientific AI. The main ones are keeping AI steerable as research evolves, deciding which tasks to delegate, and avoiding the amplification of weak study design or bad data.

  2. Ai2 (Allen Institute for AI)AI score56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    AIAi2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

  3. OpenBMB (MiniCPM) · new models on Hugging FaceAI score49

    MiniCPM5-2B-Midtrain: OpenBMB releases mid-training checkpoint of 2B-class model

    AIOpenBMB released MiniCPM5-2B-Midtrain, a BF16 mid-training checkpoint taken before SFT in the MiniCPM5-2B series, on Hugging Face and ModelScope. The series is a 2B dense Transformer with 2,516,756,480 total parameters and a 131,072-token context length, and the final MiniCPM5-2B reports an average score of 53.9 against 51.1 for the best larger comparison model. The release also includes GGUF, MLX, and GPTQ variants, along with the UltraData datasets.

Aug 31

Aug 31Mon

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP Releases Core-Embed 8B for Compositional Multimodal Retrieval

    AIAlibaba NLP has released core-emb-8b, an MLLM-based multimodal embedding model that distills a reranker's compositional judgments to distinguish attribute-object bindings such as "a white plate and a black chair" versus "a black plate and a white chair." The 8B dense embedding model, built on the Qwen3-VL-based VL-Emb backbone, scores 0.666 total average on compositional benchmarks, 5.7 points above its backbone. It is part of a family that also includes 2B embedding and reranker models.

  2. Jazzyear · InsightsAI score40

    Helical Fusion's Stellarator Design Uses AI to Cut Parameters to Three

    AIHelical Fusion CTO Wei Xishuo told the NFEC2026 AI-for-fusion forum that the company uses an autoencoder to compress hundreds of stellarator shape parameters into three. The company says this lets it predict zonal flow residuals and turbulent transport from more than 15,000 global simulations and generate new configurations with up to 100x better confinement in simulation.

  3. Fireworks AI BlogAI score57

    Fireworks AI makes its Training API generally available for custom model training

    AIFireworks AI announced general availability of its Training API, which connects a customer's Python training loop to managed distributed training and rollout infrastructure. Serverless training bills per token for LoRA adapters, while Dedicated training provides per-GPU-hour capacity for full-parameter runs and larger models. The post cites customer results, including Heidi moving a clinical scribe from proof of concept to production in four weeks with 3.5x lower latency.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.