Updated
#Data/Training
Updated
Sep 15
Elad GilAI score38 Sundar PichaiAI score42 Google outlines AI for science, weather, languages, and economic research
AIGoogle says it is focusing AI efforts on health, disaster and weather resilience, learning, and economic opportunity. Recent examples include AlphaGenome Atlas, which maps all 9B possible single-letter genetic changes across the human genome and is openly available to researchers, and WeatherNext 3, described as its most accurate and capable global weather AI model to date. The post also cites AI & Economy ATLAS, an open-access look at global AI usage, and says its translation services now cover nearly 300 languages spoken by 7B people.
Jason WeiAI score40 Jason Wei says wet-lab data lets a specialized model beat GPT-6 Astra
AIJason Wei argues that specialized, often private wet-lab data can let a task-specific model outperform a general frontier model on scientific tasks. He cites Neon, an open-source model that Liam Fedus says was mid-trained and RL-tuned on experimental data using 1,300 H200s to surpass GPT-6 Astra on an analysis benchmark. The post frames this data as a potential moat as work moves toward the frontier of science.
Dwarkesh PatelAI score38 Visited the lab - was struck both by how wide the search space is for materials synthesis experiments, and also how amenable it is to depth…
AI…first search, where the design and informativeness of your next experiment improves as you pile up more data from previous runs.
OdysseyAI score38 Odyssey-3 can also generate environments that AIs can inhabit, taking actions and learning from their consequences.
AIThese worlds support a recursive learning system, with an intelligence operating inside another intelligence, each pushing the other to become more capable.
Google · Innovation & AIAI score44 Google says its language tools now support over 300 languages used by 7 billion people
AIGoogle says its technologies now support more than 300 languages spoken by 7 billion people, representing 86% of the global population. The company also released its AI & Economy ATLAS, which it describes as a look at how people are using AI globally. The post highlights recent AI science work, including AlphaGenome Atlas, WeatherNext 3, and a Planetary Prediction Engine.
Baseten BlogAI score40 LangChain uses Baseten Loops to train custom models for LangSmith Engine
AILangChain is using Baseten Loops, a managed fine-tuning service, to train custom models for LangSmith Engine, its in-platform agent that debugs and improves AI agents. The article says LangChain fine-tunes large open-weight models on agent traces and trains smaller open-weight models such as Qwen for tasks like failure-mode categorization. Baseten Loops supports supervised fine-tuning, reinforcement learning, and long-context workloads, and lets checkpoints be evaluated and deployed directly to inference.
Tencent HunyuanAI score38 EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle
AITencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.
Sep 14
vLLM BlogPickAI score62 How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72
AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.
Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.
LlamaIndexAI score29 LlamaIndex proposes two-pass just-in-time OCR for agent document pipelines
AILlamaIndex proposes a two-pass just-in-time OCR pattern for agents working through document collections, avoiding parsing every page upfront. LiteParse, an open-source Rust tool supporting 50+ formats, performs a fast layout-aware first pass with bounding boxes, headings, tables, and a per-page complexity flag, processing a full data room in 32 seconds. LlamaParse then parses only the pages needing deeper analysis, returning cell-level tables, bounding boxes, and confidence scores.
Intern Large ModelsAI score62 Intern-S2-397B released in BF16 and FP8 under Apache 2.0
AIShanghai AI Laboratory's Intern Large Models announced Intern-S2-397B, available in BF16 and FP8 under Apache 2.0. The post reports 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both, and says it was jointly trained across 20+ scientific domains with long-horizon agent RL.
Tencent · new models on Hugging FaceAI score44 Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face
AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.
Sep 13
Ian JohnsonAI score34 what if visualizing your data was actually a video game?
AIflying through 30 million jina-v5-nano embeddings of 12 billion tokens of multilingual FineWeb + starcoder + pile + RedPajama looking at your data is supposed to be like eating your veggies, what if we made tasty veggies?
Sep 11
Dwarkesh PatelAI score18 Was really interesting to hear John, Beren, and Charlie speculate about why Sonnet 5 and Opus 5 feel like worse models than GLM 5.3…
AI…despite the fact that Anthropic can do raw logit distillation from Fable, and can also train Sonnet/Opus on the environments from which Fable was trained). Led to some interesting thoughts about value of distillation, what it takes to do distillation effectively, and what kinds of model behaviors are hard to extract from distillation.
Dwarkesh PatelAI score42 Dwarkesh Patel releases podcast with AI researchers on frontier progress
AIDwarkesh Patel announced a new episode featuring John Schulman, Chris O'Neill, and Beren Millidge, three AI researchers from openish companies. The discussion covers the case against recursive self-improvement, drivers of Chinese labs' progress, training of automated AI researchers, long-horizon RL, the sim-to-real gap, and the role of data and RL in recent progress.
Dwarkesh PodcastAI score62 AI researchers debate how close we are to recursive self-improvement
AIJohn Schulman, Beren Millidge, and Charlie O'Neill discuss whether current training methods could produce recursive self-improvement. They argue that progress depends on whether models can learn their own objectives and on sample efficiency, and that distillation keeps frontier capabilities from centralizing quickly.
Sep 10
Together AI BlogAI score52 Together AI expands Fine-Tuning with live metrics, expert LoRA, and early stopping
AITogether AI expanded its Fine-Tuning service with support for newer open-weight models, live metrics tracking, and finer training controls. Expert LoRA adapters can be applied to Mixture-of-Experts expert layers, and early stopping keeps the checkpoint with the best validation loss. Dataset previews, sample weights, pre-flight validation, and lower prices on selected models are also included.
Google Developers BlogAI score55 Google details autonomous LLM post-training loops using Tunix on TPUs
AIGoogle Developers Blog describes autofinetune, a project applying autonomous agent loops to LLM post-training with Tunix, Gemma, and Cloud TPUs. In an SFT case study on FunctionGemma, an agent ran 20 automated experiments on a Cloud TPU v5e-1 to adjust LoRA settings, optimizers, and learning rates. In a GRPO case study on Gemma 3 1B for GSM8K math reasoning, the agent ran 40 experiments on a Cloud TPU v6e-1 and improved total reward by about 10%.
Amazon ScienceAI score40 Amazon research explains why ML research agents don't overfit benchmarks
AIAmazon Science researchers propose that machine learning research agents avoid overfitting benchmarks despite years of iteration against the same tests. They attribute this to generalizable strategies being expressed compactly, leaving no room for memorization, while overfitting strategies fail to survive a compression bottleneck.
Sebastian RaschkaAI score19 Nice showcase that interesting LLM work can be done on single GPU!
Sherwin WuAI score62 OpenAI launches ChatGPT for Financial Services with GPT-6 Astra reasoning
AIOpenAI has made ChatGPT for Financial Services available, a tailored ChatGPT Work experience that combines built-in financial data with GPT-6 Astra's reasoning. Teams can use it to develop research, build financial models, and create customized client materials. The author says it integrates financial data sources including Daloopa, PitchBook, and LSEG.
Leandro von WerraAI score34 https://x.com/lvwerra/status/2098110576207544637
Google LabsAI score38 Google Labs' Dreambeans personalized daily story app now available to all U.S. accounts
AIGoogle Labs has made Dreambeans, its experimental app that creates personalized daily story collections, available to all U.S. accounts aged 18 and over on Android and iOS. Each daily collection combines personalized topics with information distilled from connected Google apps, including Calendar, Gmail, Photos, Search, YouTube, and Gemini. Users can dive deeper into stories, bookmark them, share them, and give feedback to improve future collections.
John SchulmanAI score40 Schulman says user data gains in math are unlikely; disclosure norms needed
AIJohn Schulman argues that training on user data contributes little to frontier math gains, which come mainly from scaling pretraining and RLVR. He says user data is more likely used to find failure modes that hired annotators struggle to recreate. He calls for stronger norms on disclosing how companies train on user data, including the methods and capabilities targeted.
Amazon ScienceAI score55 Research agents avoid overfitting when their winning strategies compress into few tokens
AIAmazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.
Aidan GomezAI score7 Encoders are back baby
Replit BlogAI score36 Replit and Databricks Integration Becomes Generally Available with Lakebase Support
AIReplit's integration with Databricks is now generally available, adding native Databricks Lakebase support that lets Replit Agent automatically provision a Lakebase database when an app is ready to deploy. Apps built with Replit can read live Databricks warehouse data while storing new app data in Lakebase, inheriting existing Unity Catalog security and governance controls. The update also adds automated preview deploys that keep test data isolated from live business data.
Mistral AIAI score36 Cloudera and Mistral Partner to Deliver Sovereign AI on Enterprise Data
AIMistral AI and Cloudera announced a partnership that integrates Mistral's models with Cloudera's hybrid data platform for enterprise AI. Customers can run inference across private and public cloud, on-prem, and fully air-gapped environments, and train custom models on proprietary data while retaining ownership. Cloudera cited 30 exabytes of customer-managed data on its platform.
Tencent HunyuanAI score60 Tencent Hunyuan releases open-source AuK audio model for speech generation and editing
AITencent Hunyuan has released AuK, an open-source foundation model for unified speech generation and editing that takes natural-language instructions and reference audio. It supports tasks including zero-shot TTS, timbre, style and emotion editing, denoising, and music separation. A companion AuK-Flash variant runs 4-step inference and is about 4.5 times faster under matched conditions, with code, weights, and a demo now available.
Sep 9
BAAI · new models on Hugging FaceAI score24 BAAI open-sources EPT, UniPath, and MiSI AIDD molecular and crystal modeling resources
AIBAAI released open-source resources for three AIDD projects on Hugging Face: EPT, an equivariant pretrained transformer for unified 3D molecular representation learning, and UniPath, a learnable-time flow matching method for crystal structure and energy prediction. The repository mirrors their GitHub source code and READMEs, with setup, preprocessing, training, and evaluation documentation. The MiSI benchmark is released separately on Hugging Face.
Fireworks AI BlogAI score58 Fireworks AI outlines a staged path from closed APIs to owned specialized models
AIFireworks AI describes a four-stage path for teams moving from renting closed frontier models to training their own, starting with API use and prompt, context, and harness engineering. The post uses the UIPad computer-use dataset to show that Kimi K3 ties GPT 5.6 Sol overall at 87.7 but wins three of four categories while costing about half as much, suggesting routing. After roughly three hours of training on the training split, the tuned Kimi K3 outperforms GPT 5.6 Sol on the held-out test set.
Fireworks AI BlogPickAI score60 Genspark's Gen-1 Slides matches Opus 5 decks at about one-tenth the cost per deck
AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.
Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.
LlamaIndexAI score23 LlamaParse now available as a ChatGPT connector for document parsing
AILlamaIndex has made LlamaParse available in the ChatGPT plugin directory, following its earlier Claude integration. The connector parses scanned, table-heavy, and chart-filled documents into Markdown, JSON, or HTML, extracts fields into a user-defined schema, searches document collections, and classifies and splits files into sections.
Mistral AIAI score54 Mistral details how AI agents migrated 40,000 lines of Fortran to C++
AIMistral AI helped a European energy operator migrate 40,000 lines of Fortran 77 to C++ for a reservoir simulator with no test suite. The post explains a parity harness that checks numerical agreement between the two codebases, and a workflow where agents coder, tester, and reviewer migrate modules under human review. Its authors note the approach covered the self-contained first sprint of 40,000 of 300,000 lines and that dependent systems would bring additional challenges.
Ai2 (Allen Institute for AI)AI score39 Goodfire Traces Olmo Safety Regression to Preference Training Data
AIGoodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.
Sep 8
Ian JohnsonAI score22 I've come up with a new way of exploring large datasets with UMAP called Latent Craft.
AIFly through this amazing dataset and mine for images you want to collect. all 1 million explorable in the browser!