Skip to contentSkip to stories

Updated

#Data/Training

Jun 6

Jun 6Sat
  1. Ahead of AI (Sebastian Raschka)AI score32

    Raschka Lists 2026 LLM Research Papers from January Through May, Heavy on Reasoning and Efficiency

    AISebastian Raschka has published a curated list of LLM research papers he bookmarked from January through May 2026, not a complete survey of the field. The list is weighted toward reasoning models, reinforcement learning, and efficient inference, with added interest in agent harnesses, long context, and diffusion language models. He highlights Nvidia's Nemotron 3 Super, a 120B-A12B hybrid model alternating attention and Mamba-2 layers, as a must-read, and notes a 4B Nano variant and the 550B-A55B Nemotron 3 Ultra released two days earlier.

May 29

May 29Fri

May 26

May 26Tue
  1. One Useful Thing (Ethan Mollick)AI score40

    Mollick Warns AI Writing Defaults Erode Learning and Human Thinking

    AIEthan Mollick argues that using AI as a default for writing, without thinking, risks undermining the human effort that builds skill and style. He cites two Wharton-linked studies: a Turkish high school experiment where ChatGPT access hurt test performance, and a Taipei Python course where a personalized AI tutor raised exam scores by 0.15 standard deviations. Mollick calls the difference how AI is used, not whether, and notes that the tools for tutor-style learning are not intuitive to access.

May 25

May 25Mon
  1. MiniMax BlogAI score67

    MiniMax Explains Why Its Model Failed to Output Certain Rare Chinese Tokens

    AIMiniMax says the M2 series could not generate the rare token "嘉祺" in names like Ma Jiaqi, and its investigation traced the cause to post-training. The company found the token was learned in pretraining, but low coverage of rare tokens in post-training data caused lm_head vectors to drift. Adding synthetic full-vocabulary repetition data restored generation for these tokens and reduced Japanese-to-Russian mixing from 47% to 1%.

    Why it matters: The post traces a specific token failure through tokenizer, embedding, and lm_head checks, showing a reusable way to diagnose post-training generation problems.

May 21

May 21Thu
  1. Tri DaoAI score44

    Transformers reduce to GEMM-plus-epilogue, enabling LLM-written fast kernels

    AITri Dao says that after a mathematical rewrite, all transformer operations can be expressed as a series of GEMMs with epilogues. Given a few optimized primitives, LLMs and novice humans can write near speed-of-light kernels for transformer ops. The related CODA work fuses memory-bound surrounding ops into the matmul epilogue, and LLMs can also write CODA kernels approaching speed-of-light.

May 20

May 20Wed
  1. Stability AIAI score62

    Stability AI releases Stable Audio 3.0 model family with open-weight music models

    AIStability AI released Stable Audio 3.0, a family of four audio models trained on fully licensed data. Three of them, Small SFX, Small and Medium, have open weights on Hugging Face, while Large is available through the Stability AI API and enterprise self-hosting. Outputs can be distributed and commercialized under the Stability AI Community License, and organizations with more than $1M in annual revenue can use the Enterprise License.

    Why it matters: The source specifies which models are open-weight, their licensing terms, and clip-length limits, which matters for anyone deciding whether to build on them.

May 16

May 16Sat
  1. Ahead of AI (Sebastian Raschka)AI score62

    Recent LLM architecture changes that cut long-context KV cache and attention cost

    AISebastian Raschka reviews recent open-weight LLM architecture changes aimed at reducing long-context memory and compute costs. He covers KV sharing and per-layer embeddings in Gemma 4, per-layer query-head budgeting in Laguna XS.2, Compressed Convolutional Attention in ZAYA1-8B, and mHC with CSA/HCA compressed attention in DeepSeek V4. The article reports that DeepSeek V4-Pro uses 27% of single-token inference FLOPs and 10% of the KV cache size of DeepSeek V3.2 at a 1M-token context.

Apr 29

Apr 29Wed

Apr 23

Apr 23Thu
  1. Apple · new models on Hugging FaceAI score40

    Apple releases CADD-Base-7B, a masked diffusion model for code generation

    AIApple has released CADD-Base-7B on Hugging Face, a 7B masked diffusion language model for code generation that uses Continuously Augmented Discrete Diffusion (CADD) to guide discrete denoising with a continuous flow-matching signal. The model loads through Transformers with trust_remote_code, and its diffusion_generate method supports CADD sampling modes "weighted" and "argmax" with alg options such as "entropy" and "maskgit_plus". The release builds on DiffuCoder and reuses Dream's modeling architecture and generation utilities.

Apr 22

Apr 22Wed

Apr 21

Apr 21Tue

Apr 20

Apr 20Mon
  1. Berkeley AI ResearchAI score44

    GRASP: A Gradient-Based Planner for Long-Horizon World Model Planning

    AIBerkeley AI Research introduces GRASP, a gradient-based planner for learned world models that aims to make long-horizon planning more robust. GRASP lifts trajectories into virtual states for parallel optimization across time, adds stochasticity to state iterates for exploration, and reshapes gradients to avoid brittle state-input gradients through high-dimensional vision models. The post identifies ill-conditioned gradients and non-greedy loss landscapes as core failure modes of standard rollout-based planning.

Apr 17

Apr 17Fri
  1. OpenAI · new models on Hugging FaceAI score41

    OpenAI Releases Privacy Filter, an Open-Weight PII Detection Model on Hugging Face

    AIOpenAI released Privacy Filter, a bidirectional token-classification model that detects and masks personally identifiable information in text under the Apache 2.0 license. The model has 1.5B total parameters with 50M active, supports a 128,000-token context window, and can run in a web browser or on a laptop. Users can fine-tune it and adjust precision/recall tradeoffs through preset operating points.

Apr 16

Apr 16Thu

Apr 14

Apr 14Tue

Apr 13

Apr 13Mon
  1. ARC PrizeAI score58

    ARC Prize Releases Human Performance Dataset for ARC-AGI-3 Benchmark

    AIARC Prize Foundation released an open-source human dataset for ARC-AGI-3, covering 342 step-by-step replays across 25 public environments from a study of 458 participants. The source reports that every environment was solved by at least two humans, and it updates scoring by moving the per-level baseline to the median human player and raising the per-level cap from 100% to 115%.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    AICognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    Why it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

Apr 7

Apr 7Tue

Apr 6

Apr 6Mon

Apr 4

Apr 4Sat

Apr 1

Apr 1Wed

Mar 26

Mar 26Thu
  1. Hamel HusainAI score38

    Data Scientists Face New Pressures as LLM APIs Let Teams Ship AI Without Them

    AIHamel Husain argues data scientists remain essential as foundation-model APIs let teams ship AI without them, because much of the work lies in evaluation, debugging, and metric design. He says teams often rely on generic off-the-shelf metrics and unverified LLM judges instead of examining their own data. He lists five eval pitfalls, starting with generic metrics, and recommends looking at traces and doing error analysis.

Mar 23

Mar 23Mon
  1. Jim FanAI score40

    Jim Fan says robot learning from human video replaces teleoperation in 2026

    AIJim Fan argues that behavior cloning directly from humans, following EgoScale and its dexterity scaling law, has become the way to move past teleoperation. He says 2026 will focus on scaling robot learning without robots. The post is cited alongside EgoVerse, an ecosystem for egocentric human data with 1300+ hours across 240 scenes and 2000+ tasks.

Mar 22

Mar 22Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score32

    PrismAudio Adds Reinforcement Learning to Video-to-Audio Generation with Chain-of-Thought Planning

    AIPrismAudio is a framework that integrates reinforcement learning into video-to-audio generation, using a Chain-of-Thought planning mechanism. It builds on ThinkSound by splitting single-step reasoning into four CoT modules for semantic, temporal, aesthetic, and spatial dimensions, each with targeted reward functions. Code, model weights, and datasets are released for research and educational use under the MIT License, and commercial use requires explicit author authorization.

Mar 19

Mar 19Thu
  1. Tri DaoAI score52

    Tri Dao Says Nonlinear RNNs Differ From Attention and Linear SSMs

    AITri Dao says nonlinear RNNs seem to do something genuinely different from attention and linear RNNs or SSMs. He reports they already perform well with the right parametrization, and adding just one nonlinear RNN layer substantially improves a transformer-Mamba/DeltaNet hybrid. The post quotes the M²RNN paper, which introduces non-linear RNNs with matrix-valued states for language modeling, with links to the paper, code, and models.

Mar 17

Mar 17Tue
  1. Apple · new models on Hugging FaceAI score43

    Apple releases SimpleSD-4B-thinking, a self-distilled Qwen model for code generation

    AIApple has published SimpleSD-4B-thinking on Hugging Face, a research checkpoint built on Qwen that improves code generation through Simple Self-Distillation without rewards, verifiers, teacher models, or reinforcement learning. On LiveCodeBench, it lifts Qwen3-4B-Thinking-2507 from 54.5% to 57.8% pass@1 on LCBv6 and from 59.6% to 63.1% pass@1 on LCBv5. The model is released as a reproducibility checkpoint under the Apple Machine Learning Research Model License, not as an optimized Qwen release.

  2. Apple · new models on Hugging FaceAI score46

    Apple releases SimpleSD-4B-instruct, a self-distilled Qwen code model

    AIApple has released SimpleSD-4B-instruct on Hugging Face, a research checkpoint fine-tuned from Qwen3-4B-Instruct-2507 on its own sampled outputs to improve code generation. On LiveCodeBench, the model scores 41.5% pass@1 on LCBv6, up from the base model's 34.0%, and 45.7% pass@1 on LCBv5, up from 34.3%. The model is released under the Apple Machine Learning Research Model License and is intended for reproducibility rather than as an optimized Qwen release.

  3. BAAIAI score46

    BAAI unveils RoboBrain-Dex, dexterous manipulation trained on human egocentric data

    AIBAAI has released RoboBrain-Dex, a dexterous manipulation model for embodied intelligence trained on large-scale, diverse human egocentric data rather than massive robot teleoperation datasets. The company says this approach yields strong generalization, marking a shift from small-data, weakly generalizing methods toward big-data robotic manipulation. The code has been open-sourced on GitHub.

Mar 13

Mar 13Fri
  1. Berkeley AI ResearchAI score34

    SPEX and ProxySPEX Identify Influential LLM Interactions at Scale with Fewer Ablations

    AIBerkeley AI Research introduces SPEX, a signal-processing framework that identifies influential interactions in LLMs using far fewer ablations than exhaustive analysis. A hierarchy-based extension, ProxySPEX, matches SPEX performance with around 10x fewer ablations. The methods apply to feature, data, and model component attribution.

  2. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score44

    Fun-CineForge Releases Open-Source Dubbing Pipeline, Model, and CineDub-CN Dataset

    AIFun-CineForge, from FunAudioLLM, is an open-source toolkit with an end-to-end dataset pipeline and an MLLM-based model for zero-shot movie dubbing across diverse cinematic scenes. The team built CineDub-CN, described as the first large-scale Chinese television dubbing dataset, and reports that its model outperforms state-of-the-art methods on audio quality, lip-sync, timbre transition, and instruction following. Inference code and checkpoints were released on March 16, 2026, and the model runs on a consumer-grade GPU.

Feb 25

Feb 25Wed
  1. Jim FanAI score75

    EgoScale trains a 22-DoF humanoid mostly on 20,000 hours of human video

    AIResearchers trained a humanoid with 22-DoF dexterous hands mainly on over 20,000 hours of egocentric human video, with no robot in the loop, to perform tasks such as assembling model cars and folding shirts. They report a log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and state that this loss predicts real-robot success rate. The recipe, called EgoScale, pre-trains GR00T N1.5 on the video, adds only 4 hours of robot play data, and reports a 54% gain over training from scratch across five dexterous tasks.

Feb 24

Feb 24Tue
  1. Jim FanAI score62

    NVIDIA's SONIC trains a 42M transformer to control a humanoid robot

    AINVIDIA researchers trained SONIC, a 42M-parameter transformer, to control a humanoid robot's whole body using motion tracking on over 100M mocap frames. After three days of training in simulation, the policy transferred zero-shot to the real G1 robot and reported a 100% success rate across 50 real-world motion sequences. One policy supports VR teleoperation, webcam human video, text prompts, music, and GR00T N1.5 VLA integration with 95% success on mobile tasks, and the code and checkpoints are open-sourced.

Feb 20

Feb 20Fri
  1. Jim FanAI score75

    DreamDojo: Open-source world model trained on 44K hours of human video

    AIJim Fan announced DreamDojo, an open-source interactive world model that takes robot motor controls and generates future frames in pixels. It is pre-trained on 44K hours of human egocentric video using latent actions, then post-trained onto specific robot hardware, and a real-time version runs at 10 FPS for live teleoperation, policy evaluation, and model-based planning. The author reports a +17% real-world success gain on a fruit packing task, and weights, code, datasets, and the whitepaper are released.

Feb 13

Feb 13Fri
  1. MiniMax BlogAI score62

    MiniMax details Forge, a scalable agent RL framework behind M2.5

    AIMiniMax describes Forge, its internal reinforcement learning framework for training real-world agents, which was used during the development of MiniMax M2.5. The post explains a Windowed FIFO scheduler, prefix tree merging that the post says yields a 40x training speedup, and CISPO-based training across more than one hundred thousand agent scaffolds and environments.

    Why it matters: The post details how the Forge framework balances throughput, stability, and agent flexibility, with concrete scheduling and prefix-merging methods for training agent RL at scale.

Feb 4

Feb 4Wed
  1. Anthropic EngineeringAI score72

    Anthropic finds container resource limits can shift agentic coding eval scores

    AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

    Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.