Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Jun 19

Jun 19Fri
  1. AI Futures ProjectAI score60

    Forecast puts China's commercial EUV lithography in late 2030s

    AIThe post argues that China's commercial-scale EUV machines should be forecast for the late 2030s and immersion DUV for the mid-2030s, using ASML's development timeline as a reference. It also weighs factors that could push these estimates earlier or later, including state funding, espionage, talent flows, and the use of AI in R&D. The authors note that forecasts placing either milestone in the 2020s would need strong justification.

Jun 18

Jun 18Thu
  1. OpenAI Alignment Research BlogAI score62

    OpenAI study finds beneficial-trait RL improves alignment across untrained domains

    AIOpenAI reports that reinforcement learning on realistic conversations targeting traits such as honesty, epistemic humility, and corrigibility improved a model across 44 out-of-distribution alignment evaluations. Gains included reward hacking, deception, and health benchmarks, and training only on health conversations still improved non-health alignment scores. The trained model was also harder to steer toward harmful behavior with adversarial persona prompts or harmful fine-tuning.

    Why it matters: The post tests whether reinforcement learning on beneficial traits in one domain transfers to unrelated alignment benchmarks and holds up under adversarial steering.

Jun 17

Jun 17Wed
  1. John SchulmanAI score40

    PPO's LLM-era revival and the unexpected reasons behind it

    AIJohn Schulman says PPO gained a second wave in the LLM era for reasons not anticipated in the original paper. He points to the importance-ratio objective, which corrects biases from numeric error, asynchronous training, and forward-pass noise, and to the clipping objective, whose effect on entropy was unknown at publication, citing DAPO's arXiv paper.

Jun 16

Jun 16Tue
  1. OpenAI Alignment Research BlogAI score60

    WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x

    AIOpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude. The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.

    Why it matters: The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.

Jun 10

Jun 10Wed
  1. ByteDance · new models on Hugging FaceAI score52

    ByteDance open-sources Bernini-Diffusers for semantic video generation and editing

    AIByteDance open-sourced inference code and model weights for Bernini-Diffusers, a full video generation and editing pipeline with an MLLM-based semantic planner and a DiT-based renderer. The release bundles a Qwen2.5-VL planner and Wan2.2 diffusion components in one self-contained directory, and the source recommends it over the renderer-only Bernini-R for complex instruction following.

Jun 8

Jun 8Mon
  1. Xiaomi MiMoAI score62

    Xiaomi MiMo open-sources a 1T model running over 1,000 tps on 8 GPUs

    AIXiaomi MiMo and the TileRT team say a 1T model exceeds 1,000 tps on a single standard 8-GPU node using general-purpose GPUs. The speedup comes from FP4 quantization and DFlash, a block-masked parallel speculative decoding method that accepts more tokens per verification, with TileRT tailoring its compiler and kernels to these techniques. Open weights for the FP4 + DFlash checkpoint are available on Hugging Face.

Jun 6

Jun 6Sat
  1. Ahead of AI (Sebastian Raschka)AI score32

    Raschka Lists 2026 LLM Research Papers from January Through May, Heavy on Reasoning and Efficiency

    AISebastian Raschka has published a curated list of LLM research papers he bookmarked from January through May 2026, not a complete survey of the field. The list is weighted toward reasoning models, reinforcement learning, and efficient inference, with added interest in agent harnesses, long context, and diffusion language models. He highlights Nvidia's Nemotron 3 Super, a 120B-A12B hybrid model alternating attention and Mamba-2 layers, as a must-read, and notes a 4B Nano variant and the 550B-A55B Nemotron 3 Ultra released two days earlier.

May 29

May 29Fri
  1. Fei-Fei LiAI score38

    Fei-Fei Li Highlights GPIC, a Permissive Image Corpus for Visual Generation

    AIFei-Fei Li praised GPIC, a new benchmark dataset for visual generation built for modern large-scale generative models. The corpus includes 100M VLM-captioned image-text pairs for training and 1M pairs for benchmarking, totaling about 28 trillion pixels. It is centrally hosted and fully permissive for research and commercial use.

May 26

May 26Tue
  1. One Useful Thing (Ethan Mollick)AI score40

    Mollick Warns AI Writing Defaults Erode Learning and Human Thinking

    AIEthan Mollick argues that using AI as a default for writing, without thinking, risks undermining the human effort that builds skill and style. He cites two Wharton-linked studies: a Turkish high school experiment where ChatGPT access hurt test performance, and a Taipei Python course where a personalized AI tutor raised exam scores by 0.15 standard deviations. Mollick calls the difference how AI is used, not whether, and notes that the tools for tutor-style learning are not intuitive to access.

May 25

May 25Mon
  1. MiniMax BlogAI score67

    MiniMax explains why its LLM failed to generate the name Ma Jiaqi

    AIMiniMax says its M2 series could not output the name Ma Jiaqi, a failure it traced to post-training data that rarely included the token. Its tests found the input embedding stayed stable while the lm_head weights for low-frequency tokens drifted during SFT. A synthetic full-vocabulary repetition dataset restored generation for affected tokens and reduced Japanese-to-Russian confusion from 47% to 1%.

    Why it matters: The post traces a community-noticed token failure through tokenizer, embedding, and lm_head tests, showing how post-training data coverage can cause low-frequency token drift.

May 21

May 21Thu
  1. Tri DaoAI score44

    Transformers reduce to GEMM-plus-epilogue, enabling LLM-written fast kernels

    AITri Dao says that after a mathematical rewrite, all transformer operations can be expressed as a series of GEMMs with epilogues. Given a few optimized primitives, LLMs and novice humans can write near speed-of-light kernels for transformer ops. The related CODA work fuses memory-bound surrounding ops into the matmul epilogue, and LLMs can also write CODA kernels approaching speed-of-light.

May 20

May 20Wed
  1. Stability AIAI score62

    Stability AI releases Stable Audio 3.0 model family with open-weight music models

    AIStability AI released Stable Audio 3.0, a family of four audio models trained on fully licensed data. Three of them, Small SFX, Small and Medium, have open weights on Hugging Face, while Large is available through the Stability AI API and enterprise self-hosting. Outputs can be distributed and commercialized under the Stability AI Community License, and organizations with more than $1M in annual revenue can use the Enterprise License.

    Why it matters: The source specifies which models are open-weight, their licensing terms, and clip-length limits, which matters for anyone deciding whether to build on them.

May 16

May 16Sat
  1. Ahead of AI (Sebastian Raschka)AI score62

    Recent LLM architecture changes that cut long-context KV cache and attention cost

    AISebastian Raschka reviews recent open-weight LLM architecture changes aimed at reducing long-context memory and compute costs. He covers KV sharing and per-layer embeddings in Gemma 4, per-layer query-head budgeting in Laguna XS.2, Compressed Convolutional Attention in ZAYA1-8B, and mHC with CSA/HCA compressed attention in DeepSeek V4. The article reports that DeepSeek V4-Pro uses 27% of single-token inference FLOPs and 10% of the KV cache size of DeepSeek V3.2 at a 1M-token context.

Apr 29

Apr 29Wed

Apr 23

Apr 23Thu
  1. Apple · new models on Hugging FaceAI score40

    Apple releases CADD-Base-7B, a masked diffusion model for code generation

    AIApple has released CADD-Base-7B on Hugging Face, a 7B masked diffusion language model for code generation that uses Continuously Augmented Discrete Diffusion (CADD) to guide discrete denoising with a continuous flow-matching signal. The model loads through Transformers with trust_remote_code, and its diffusion_generate method supports CADD sampling modes "weighted" and "argmax" with alg options such as "entropy" and "maskgit_plus". The release builds on DiffuCoder and reuses Dream's modeling architecture and generation utilities.

Apr 22

Apr 22Wed

Apr 21

Apr 21Tue

Apr 20

Apr 20Mon
  1. Berkeley AI ResearchAI score44

    GRASP: A Gradient-Based Planner for Long-Horizon World Model Planning

    AIBerkeley AI Research introduces GRASP, a gradient-based planner for learned world models that aims to make long-horizon planning more robust. GRASP lifts trajectories into virtual states for parallel optimization across time, adds stochasticity to state iterates for exploration, and reshapes gradients to avoid brittle state-input gradients through high-dimensional vision models. The post identifies ill-conditioned gradients and non-greedy loss landscapes as core failure modes of standard rollout-based planning.

Apr 17

Apr 17Fri
  1. OpenAI · new models on Hugging FaceAI score41

    OpenAI Releases Privacy Filter, an Open-Weight PII Detection Model on Hugging Face

    AIOpenAI released Privacy Filter, a bidirectional token-classification model that detects and masks personally identifiable information in text under the Apache 2.0 license. The model has 1.5B total parameters with 50M active, supports a 128,000-token context window, and can run in a web browser or on a laptop. Users can fine-tune it and adjust precision/recall tradeoffs through preset operating points.

Apr 16

Apr 16Thu

Apr 14

Apr 14Tue

Apr 13

Apr 13Mon
  1. ARC PrizeAI score58

    ARC Prize Releases Human Performance Dataset for ARC-AGI-3 Benchmark

    AIARC Prize Foundation released an open-source human dataset for ARC-AGI-3, covering 342 step-by-step replays across 25 public environments from a study of 458 participants. The source reports that every environment was solved by at least two humans, and it updates scoring by moving the per-level baseline to the median human player and raising the per-level cap from 100% to 115%.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    AICognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    Why it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

Apr 7

Apr 7Tue
  1. Werner VogelsAI score62

    Amazon S3 Files lets users mount any S3 bucket as a filesystem

    AIWerner Vogels announced S3 Files, which lets users mount any S3 bucket as a filesystem without making copies, running sync scripts, or choosing between file and object storage. He linked to a detailed post by Andy Warfield on the feature and its design history, including the filerectories concept that did not make the final release.

Apr 6

Apr 6Mon

Apr 4

Apr 4Sat

Apr 1

Apr 1Wed

Mar 26

Mar 26Thu
  1. Hamel HusainAI score38

    Data Scientists Face New Pressures as LLM APIs Let Teams Ship AI Without Them

    AIHamel Husain argues data scientists remain essential as foundation-model APIs let teams ship AI without them, because much of the work lies in evaluation, debugging, and metric design. He says teams often rely on generic off-the-shelf metrics and unverified LLM judges instead of examining their own data. He lists five eval pitfalls, starting with generic metrics, and recommends looking at traces and doing error analysis.

Mar 23

Mar 23Mon
  1. Jim FanAI score40

    Jim Fan says robot learning from human video replaces teleoperation in 2026

    AIJim Fan argues that behavior cloning directly from humans, following EgoScale and its dexterity scaling law, has become the way to move past teleoperation. He says 2026 will focus on scaling robot learning without robots. The post is cited alongside EgoVerse, an ecosystem for egocentric human data with 1300+ hours across 240 scenes and 2000+ tasks.

Mar 22

Mar 22Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score32

    PrismAudio Adds Reinforcement Learning to Video-to-Audio Generation with Chain-of-Thought Planning

    AIPrismAudio is a framework that integrates reinforcement learning into video-to-audio generation, using a Chain-of-Thought planning mechanism. It builds on ThinkSound by splitting single-step reasoning into four CoT modules for semantic, temporal, aesthetic, and spatial dimensions, each with targeted reward functions. Code, model weights, and datasets are released for research and educational use under the MIT License, and commercial use requires explicit author authorization.

Mar 19

Mar 19Thu
  1. Tri DaoAI score52

    Tri Dao Says Nonlinear RNNs Differ From Attention and Linear SSMs

    AITri Dao says nonlinear RNNs seem to do something genuinely different from attention and linear RNNs or SSMs. He reports they already perform well with the right parametrization, and adding just one nonlinear RNN layer substantially improves a transformer-Mamba/DeltaNet hybrid. The post quotes the M²RNN paper, which introduces non-linear RNNs with matrix-valued states for language modeling, with links to the paper, code, and models.

Mar 17

Mar 17Tue
  1. Apple · new models on Hugging FaceAI score43

    Apple releases SimpleSD-4B-thinking, a self-distilled Qwen model for code generation

    AIApple has published SimpleSD-4B-thinking on Hugging Face, a research checkpoint built on Qwen that improves code generation through Simple Self-Distillation without rewards, verifiers, teacher models, or reinforcement learning. On LiveCodeBench, it lifts Qwen3-4B-Thinking-2507 from 54.5% to 57.8% pass@1 on LCBv6 and from 59.6% to 63.1% pass@1 on LCBv5. The model is released as a reproducibility checkpoint under the Apple Machine Learning Research Model License, not as an optimized Qwen release.

  2. Apple · new models on Hugging FaceAI score46

    Apple releases SimpleSD-4B-instruct, a self-distilled Qwen code model

    AIApple has released SimpleSD-4B-instruct on Hugging Face, a research checkpoint fine-tuned from Qwen3-4B-Instruct-2507 on its own sampled outputs to improve code generation. On LiveCodeBench, the model scores 41.5% pass@1 on LCBv6, up from the base model's 34.0%, and 45.7% pass@1 on LCBv5, up from 34.3%. The model is released under the Apple Machine Learning Research Model License and is intended for reproducibility rather than as an optimized Qwen release.