Skip to contentSkip to stories

Updated

#Data/Training

Jul 27

Jul 27Mon

Jul 26

Jul 26Sun
  1. Philipp SchmidAI score62

    EvoCode-Bench Tests Coding Agents Across Multi-Turn Iterative Specification Changes

    AIEvoCode-Bench is a multi-turn coding benchmark with 26 tasks spanning 227 sequential rounds, where agents keep a persistent workspace and must pass cumulative tests after each evolving instruction. The results show that agents perform much worse when building on their own prior work than when starting from a clean, human-completed codebase. Regressions, not failure to implement new features, are the main bottleneck, and agents that maintained a persistent requirements document more than doubled their success rates.

  2. Berkeley AI ResearchAI score44

    Berkeley AI Research Trains LLMs to Update Beliefs for Long Tasks

    AIBerkeley AI Research introduces ABBEL, a framework that replaces full interaction histories with natural-language belief states that models update as new observations arrive. On CollabBench collaborative coding, belief grading closes about half the performance gap to full-context models while using fewer peak tokens and training in 50 steps instead of 100.

Jul 25

Jul 25Sat
  1. Ali GhodsiAI score26

    Longer-running AI agents often perform worse than faster ones, says Ghodsi

    AIAli Ghodsi argues that AI agents which take longer to work through a task are often worse, while Genie reaches results faster. He adds that ontology will be key to giving agents the context they need to answer correctly and quickly. The related post reports that Genie Code outperformed three general-purpose coding agents on more than 400 real user data tasks.

Jul 23

Jul 23Thu
  1. Sequoia CapitalAI score58

    Western AI Builders Depend on Chinese Open Models Through Distillation

    AIThe essay argues that Western companies increasingly rely on Chinese open-weight models like Qwen and Kimi for post-training, while Western labs cannot lawfully distill from American frontier models. It says Qwen's share of new open-model fine-tunes rose from 1% in January 2024 to 69% by February 2026, citing ATOM's Report. The authors propose controlled teacher access and tighter enforcement against foreign distillation as a domestic alternative.

  2. Matei ZahariaAI score36

    Berkeley STAR Lab packages AI research optimizers into one GEPA API

    AIBerkeley's STAR Lab packaged multiple LLM-based "autoresearch" algorithms into a single API within the GEPA package, letting users mix and match them. The optimizers can be applied to tasks including prompt writing, agent design, and code optimization. The quoted thread adds that GEPA, AutoResearch, and Meta-Harness each win on different tasks, and that the new optimize_anything omni meta-optimizer beats every standalone optimizer at a matched budget.

  3. Ahmad Al-DahleAI score62

    Ahmad Al-Dahle outlines five myths about AI model distillation

    AIAl-Dahle argues that distillation is a standard training method used inside labs, under licenses, or without authorization, so it does not by itself show theft. He says a few million conversations are small against trillion-token runs, yet can matter in late-stage training, reinforcement learning bootstrapping, or training a grader. He also argues that model outputs are hard to trace after paraphrasing or mixing, and that transferred capability is difficult to measure.

Jul 22

Jul 22Wed

Jul 21

Jul 21Tue
  1. Meta AI BlogAI score44

    Meta's SAM 3 and DINOv3 Power SYNAPS-I's Genesis Mission Imaging Pipeline

    AISYNAPS-I, a multi-lab Genesis Mission project led by Lawrence Berkeley National Laboratory, uses Meta's open-source SAM 3 and DINOv3 models to segment X-ray and micro-CT scientific imagery. The fine-tuned pipeline, run on 300 A100 GPUs, reduced a grapevine xylem analysis from a month of expert annotation per time step to about 15 minutes. The team can deploy the open models inside secure national lab infrastructure, where research data must remain.

Jul 15

Jul 15Wed
  1. Fei-Fei LiAI score60

    RoboTTT scales robot policy context to 8,000 timesteps using test-time training

    AIStanford SVL and NVIDIA Robotics introduced RoboTTT, which uses test-time training to give robot policies up to 8,000 timesteps of context at constant inference cost. The source reports that 8K-context pretraining beats 1K by 62%, and that performance keeps improving from 128 to 8K timesteps with no sign of saturation. The authors also describe one-shot imitation from human video and in-episode error recovery.

  2. Liquid AI NewsletterAI score38

    Liquid AI Releases Antidoom and IFStruct to Fix Reasoning Loops and Schema Errors

    AILiquid AI released Antidoom, an open-source method that retrains a single overtrained token to eliminate "doom loops" in small reasoning models. On LFM2.5-2.6B and Qwen3.5-4B, loop rates fell from 10.2% to 1.4% and from 22.9% to 1%, respectively. The company also released IFStruct, an open-source benchmark measuring whether model outputs satisfy a schema, where LFM2.5-350M rose from 21.10% to 44.90% after training.

  3. Jim FanAI score62

    RoboTTT scales robot policy context to 8,000 timesteps with constant inference cost

    AIJim Fan introduced RoboTTT, a robot model that uses test-time training to compress history into a tiny inner model updated at each sensor reading. The post reports closed-loop performance rising steadily from 128 to 8K timesteps, and 8K-context pretraining beating 1K by 62%. It also claims one-shot in-context learning from human video and mid-episode error recovery, with learning continuing after deployment.

Jul 12

Jul 12Sun
  1. ByteDance · new models on Hugging FaceAI score41

    ByteDance releases UniVR-34B-Planning for visual-space reasoning and planning

    AIByteDance's UniVR-34B-Planning, built on Emu3.5 at 34B parameters, learns visual reasoning, physical dynamics, and long-term planning from visual demonstrations using a next-token objective and two-stage training on the VR-X dataset with VR-GRPO reinforcement learning. On the VR-X benchmark it scores 58.2 overall, up 18.4 points from the Emu3.5 34B baseline of 39.8. The Planning checkpoint is available on Hugging Face under CC BY 4.0, alongside a General checkpoint.

Jul 9

Jul 9Thu
  1. Thinking Machines LabAI score44

    Thinking Machines Argues the Future Worth Building Keeps Humans Central to AI Decisions

    AIThinking Machines Lab says AI should extend human will and judgment, with people shaping its goals through continuous feedback rather than relying on models trained once and frozen. The company outlines three technical directions: training strong models, building tools for customization including training model weights, and developing interfaces that let personal judgment influence AI work. It also says it will publish research for the scientific community.

  2. Benedict EvansAI score60

    Benedict Evans argues AI token prices face unstable, commodity-leaning equilibrium

    AIBenedict Evans argues that token prices are unstable amid a supply crunch, and that foundation models may end up as low-margin commodity infrastructure rather than holding lasting pricing power. He cites inference gross margins of 40-50% that exclude training costs, which currently exceed revenue, and compares the outlook with mobile data and semiconductor manufacturing. He concludes that the outcome remains uncertain and that value capture above the model layer would require changes not yet visible.

Jul 8

Jul 8Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition releases SWE-1.7, a coding model trained with long-horizon RL

    AICognition launched SWE-1.7, which it says reaches frontier-level coding performance at lower cost, trained from a Kimi K2.7 base. The post describes RL methods including top-p sampling replay to preserve entropy, compressed weight deltas across multi-cluster training, and self-compaction for rollouts up to six hours. SWE-1.7 is available in Devin via Cerebras at 1000 TPS.

    Why it matters: The post details entropy preservation, multi-cluster weight sync, and self-compaction, offering concrete RL training techniques for long-horizon coding agents to compare against one's own pipeline.

Jul 7

Jul 7Tue
  1. Berkeley AI ResearchAI score62

    Berkeley researchers outline how data systems must change as agents take over knowledge work

    AIBerkeley AI Research authors argue that near-free inference will make agents the dominant workload for data systems, requiring redesign for agentic speculation, agent-run state and coordination, and agent-synthesized systems. The post cites inference prices falling 9x to 900x per year with a median near 50x, and reports that about 80-90% of sub-queries in a text-to-SQL benchmark were duplicates. It frames the three directions as data systems for, of, and by agents.

    Why it matters: The piece maps three concrete data-system challenges posed by near-free inference, useful for anyone designing infrastructure for agent workloads and memory.

Jul 3

Jul 3Fri
  1. Arthur MenschAI score34

    Mistral urges enterprises to adopt open-source models and own their AI data

    AIMistral CEO Arthur Mensch argues enterprises should use open-source models and store their own data to avoid dependence on closed providers that retain customer data. He says companies should build continuous training loops from employee and user interactions, and shrink models to cut deployment costs. Mistral positions its Studio control plane and Forge training platform, deployed on customer infrastructure or via zero-data-retention hosting, as tools for this shift.

Jul 1

Jul 1Wed
  1. Jim FanAI score51

    Jim Fan introduces ASPIRE, a self-evolving robot skills library for continual learning

    AIJim Fan announces ASPIRE, a system where coding agents use multimodal sensory traces from simulation and real robots to run evolutionary search over control programs and add the results to a growing skills library. The post claims up to a roughly 10x reduction in transfer learning tokens for sim-to-real and single-arm to bimanual transfer, and says the full stack will be open-sourced.

Jun 30

Jun 30Tue
  1. John SchulmanAI score38

    Bridgewater fine-tuning with expert data beats prompting-only approaches

    AIJohn Schulman argues that fine-tuning with the right data, such as expert judgments, can substantially outperform prompting-only approaches even as general-purpose models improve. He cites Bridgewater's work, where an expert-labeled dataset and on-policy distillation were used to fine-tune a model to triage financial documents reliably and cheaply.

  2. Jim FanAI score60

    ASPIRE lets robots build an evolving skills library that transfers across tasks

    AIJim Fan introduces ASPIRE, a system in which coding agents observe multimodal sensory traces and run evolutionary search over control programs to distill skills into a growing library. The post says ASPIRE shares know-how rather than pixels or weights across the sim-to-real gap, reducing transfer learning tokens by up to about 10x. The author also says the full stack will be open-sourced and provides a gallery of 150+ tasks and 90+ skills.

Jun 29

Jun 29Mon
  1. Hamel HusainAI score54

    Why Hard-to-Eval AI Products Need Designs That Support Verification

    AIHamel Husain argues that an AI product whose output is hard to verify is a product design problem, not just an evaluation problem. He shows before-and-after sketches for an AI data agent, a PE lesson planner, and a workers' compensation report tool, each adding provenance, scoped edits, and checkable evidence. He notes that designing for verification also makes evals easier to build and grade.

  2. Meta AI BlogAI score68

    Meta's Brain2Qwerty v2 decodes sentences from non-invasive brain recordings

    AIMeta released Brain2Qwerty v2, an end-to-end deep learning pipeline that decodes sentences in real time from non-invasive brain recordings. The model reached 61% word accuracy across participants, compared with 8% for other non-invasive methods, and 78% for the best participant. Meta also released the v1 and v2 training code, and partner BCBL released the v1 dataset.

    Why it matters: The source reports word accuracy and data-scaling results for non-invasive decoding, offering a benchmark against surgical brain-computer interfaces and prior non-invasive methods.

Jun 28

Jun 28Sun
  1. PaddlePaddleAI score46

    PaddlePaddle announces Unlimited-OCR now runs in vLLM

    AIUnlimited-OCR, Baidu's long-context OCR model, now runs in vLLM, with a recipe provided for developers to try it. The background post says it parses entire books in one pass using Reference Sliding Window Attention (R-SWA), which keeps the KV cache fixed during decoding, and claims 35% faster throughput than DeepSeek-OCR at 6K output tokens.

Jun 27

Jun 27Sat
  1. PaddlePaddleAI score36

    PaddleFormers 1.2 adds DeepSeek-V4 training with 128K+ context support

    AIPaddleFormers 1.2 is released with support for training DeepSeek-V4 and 128K+ long-context training. The update adds Context Parallel, Packing, Document Mask Attention, and the Muon optimizer, plus ultra-fused mHC, CSA, and HCA operators, DeepEP/HybridEP communication, and lossless FP8 training with AutoSubbatch memory balancing. The project is presented as fully open-source and is available on GitHub.

Jun 26

Jun 26Fri
  1. Qwen · new models on Hugging FaceAI score44

    Qwen3-ForcedAligner-0.6B-hf Adds Timestamp Alignment for Speech Transcripts

    AIQwen released Qwen3-ForcedAligner-0.6B-hf, a Transformers-format forced aligner that predicts timestamps for arbitrary units within up to 5 minutes of speech in 11 languages. The model accepts transcripts from any ASR system, and the documentation shows it paired with Qwen3-ASR-0.6B and NVIDIA Parakeet CTC. Until it ships in an official Transformers release, users must install Transformers from source.

Jun 25

Jun 25Thu
  1. Lilian WengAI score40

    Lilian Weng's Overview of Scaling Laws and Compute-Optimal Allocation

    AILilian Weng published a long blog post on scaling laws, which help estimate the best split of compute between data and model size before a large training run. The post covers what scaling laws predict, how compute-optimal allocation works, and why Kaplan et al. and Chinchilla reach different conclusions. It also addresses how data limits and fitting details make extrapolation difficult.

  2. PaddlePaddleAI score38

    PP-OCRv6 recognition uses CTC and NRTR heads to curb hallucination

    AIPP-OCRv6's recognition module uses a CTC plus NRTR dual-head design so text is decoded from visual features rather than language priors, reducing hallucination. In hallucination tests, PP-OCRv6_medium reaches 93.2%, versus 85.0% for the best VLM, and recognition accuracy across 15 scenarios is 83.2%, above PP-OCRv5_server's 78.1%. NRTR is used only during training, adding language regularization at no inference cost, and it contributes +1.16% accuracy.

Jun 24

Jun 24Wed
  1. PaddlePaddleAI score30

    PP-OCRv6 Detection Module Outperforms VLMs on Text Localization Benchmarks

    AIPaddlePaddle says its PP-OCRv6_medium text detector reached an 86.2% detection Hmean in benchmarks, versus 46.8% for Gemini-3.1-Pro and 38.3% for GPT-5.5. The detector's design uses RepLKFPN with 7×7 kernels to cut FPN neck parameters from 172K to 118K, auxiliary deep supervision heads on P2–P4, and Focal Loss paired with Dice Loss, which adds +1.15% Hmean in ablation.

Jun 23

Jun 23Tue
  1. Lil'Log (Lilian Weng)AI score40

    Scaling Laws, Carefully: Early Empirical Power-Law Studies of Loss, Data and Model Size

    AILil'Log examines early empirical work showing that deep learning generalization error follows power-law curves as training data and model size grow. Hestness et al. (2017) found the exponent reflects the problem domain rather than the architecture, while Rosenfeld et al. (2020) modeled loss jointly as a function of model size N and data size D, fitting parametric forms on small configurations to extrapolate to larger ones.

Jun 19

Jun 19Fri
  1. AI Futures ProjectAI score60

    Forecast puts China's commercial EUV lithography in late 2030s

    AIThe post argues that China's commercial-scale EUV machines should be forecast for the late 2030s and immersion DUV for the mid-2030s, using ASML's development timeline as a reference. It also weighs factors that could push these estimates earlier or later, including state funding, espionage, talent flows, and the use of AI in R&D. The authors note that forecasts placing either milestone in the 2020s would need strong justification.

Jun 18

Jun 18Thu
  1. OpenAI Alignment Research BlogAI score62

    OpenAI study finds beneficial-trait RL improves alignment across untrained domains

    AIOpenAI reports that reinforcement learning on realistic conversations targeting traits such as honesty, epistemic humility, and corrigibility improved a model across 44 out-of-distribution alignment evaluations. Gains included reward hacking, deception, and health benchmarks, and training only on health conversations still improved non-health alignment scores. The trained model was also harder to steer toward harmful behavior with adversarial persona prompts or harmful fine-tuning.

    Why it matters: The post tests whether reinforcement learning on beneficial traits in one domain transfers to unrelated alignment benchmarks and holds up under adversarial steering.

Jun 17

Jun 17Wed

Jun 16

Jun 16Tue
  1. OpenAI Alignment Research BlogAI score60

    WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x

    AIOpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude. The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.

    Why it matters: The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.

Jun 10

Jun 10Wed
  1. ByteDance · new models on Hugging FaceAI score52

    ByteDance open-sources Bernini-Diffusers for semantic video generation and editing

    AIByteDance open-sourced inference code and model weights for Bernini-Diffusers, a full video generation and editing pipeline with an MLLM-based semantic planner and a DiT-based renderer. The release bundles a Qwen2.5-VL planner and Wan2.2 diffusion components in one self-contained directory, and the source recommends it over the renderer-only Bernini-R for complex instruction following.