Skip to contentSkip to stories

Updated

#Paper/Research

Oct 8

Oct 8Thu
  1. MIT News · AIAI score34

    MIT's Sasha Rakhlin outlines how universities should respond to AI in research and training

    AIMIT Statistics and Data Science Center director Sasha Rakhlin argues that AI progress is fastest where results can be verified quickly, citing a model reaching gold-medal level at the International Mathematical Olympiad a year before models produced new research results. He says departments should reconsider how they reward work, emphasizing question-asking, replication, and disclosure of AI's role in a researcher's contributions. He also urges universities to build shared lab infrastructure that captures failed experiments and tacit expertise.

  2. Lewis TunstallAI score62

    Lewis Tunstall Shares a Physics Paper Proof Developed with OpenAI's Astra Model

    AILewis Tunstall quotes Kyle Cranmer's post about a paper by Nate Gunnarsson on a non-perturbative approach to chiral fermions in the Standard Model, extending Lüscher's abelian result. The paper's acknowledgments state that OpenAI's GPT-6 Astra model was essential, proposing refinement strategies, writing rewrites of the proof, and carrying out Lean verification.

  3. SantiagoAI score40

    Seedance 2.5 tops evaluation of world models for physical consistency

    AISantiago says physical consistency is the most important and hardest feature of a world model, and that many generated videos show objects defying gravity. He reports that Seedance 2.5 is currently the best among the evaluated world models. The post links to a physics evaluation benchmark in which eight video world models reached a top score of 57.76/100.

Oct 7

Oct 7Wed
  1. Exponential ViewAI score72

    OpenAI's 722 machine-generated math results may split mathematics into two layers

    AIOpenAI released 722 mathematical manuscripts in 372 families, produced by an unreleased frontier model, with the average result taking the equivalent of three hours of ChatGPT Pro thinking. The author notes many results are verified in Lean but not all, and suggests mathematics could divide into vast machine-verified work and a compressed human 'effective theory' that people can actually understand.

Oct 6

Oct 6Tue

Oct 5

Oct 5Mon
  1. Mike KnoopAI score62

    Dust pretrains transformers with zeroth-order optimization, approaching backprop results

    AIDust is a zeroth-order method that pretrains transformers and sometimes matches or exceeds backprop given large compute. The authors report it is about 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. The post also cites the gradient-alignment result up to 1B tokens and the virtual population idea for scaling.

Oct 3

Oct 3Sat

Oct 1

Oct 1Thu
  1. Latent.SpaceAI score60

    Recursive Language Models explained by MIT's Alex Zhang on coding agents

    AIA Latent.Space podcast episode features MIT researcher Alex Zhang explaining recursive language models (RLMs). He discusses why Claude Code, Codex, and Pi are basically the same, and how RLMs use code, context offloading, and recursive subagents to generalize across tasks. The episode also covers OpenAI's 10,000-agent, 130B-output-token experiment and academia's freedom to pursue ambitious research bets.

  2. AnthropicAI score38

    Harvard physicist builds toolkit to match Claude with science calculations

    AIHarvard physicist Matthew Schwartz argues that LLMs are poorly matched to science when used as human-style collaborators, so he built a toolkit for exact quantitative calculations. Working with Claude, the approach surfaced connections to ecology, population genetics, and a dozen other fields, with domain experts steering it toward interesting questions.

Sep 4

Sep 4Fri
  1. Lewis TunstallAI score46

    Meta paper uses research preference models to guide AI agents' experiments

    AILewis Tunstall praises a new Meta paper on research preference models (RPMs), which instill "research taste" in agents by treating experiments as tree nodes. An RPM acts as an LLM judge that selects the most promising candidate experiment before it is run, reducing wasted compute. Tunstall notes the resulting trajectories could train domain-specific RPMs, which would be valuable in hard fields such as the natural sciences.

Aug 13

Aug 13Thu
  1. Air Street PressAI score52

    Air Street Press argues logged research decisions could teach AI scientific taste

    AIThe article argues that scientific papers omit the failed experiments and rejected branches that could train AI systems to develop scientific judgment. It describes Alasdair Russell's Cambridge group logging discovery paths as graphs of ideas, and proposes recording six fields per decision, including candidates and outcomes, to test whether this taste transfers to unfamiliar projects.

Jun 17

Jun 17Wed

May 7

May 7Thu

Apr 16

Apr 16Thu

Apr 14

Apr 14Tue