Skip to contentSkip to stories

Updated

#Paper/Research

Items with an AI score under 20 are hidden. Show low-relevance items

Feb 25

Feb 25Wed
  1. Jim FanAI score75

    EgoScale trains a 22-DoF humanoid mostly on 20,000 hours of human video

    AIResearchers trained a humanoid with 22-DoF dexterous hands mainly on over 20,000 hours of egocentric human video, with no robot in the loop, to perform tasks such as assembling model cars and folding shirts. They report a log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and state that this loss predicts real-robot success rate. The recipe, called EgoScale, pre-trains GR00T N1.5 on the video, adds only 4 hours of robot play data, and reports a 54% gain over training from scratch across five dexterous tasks.

  2. Quoc LeAI score65

    Aletheia Agent Solves 6 of 10 FirstProof Math Problems Autonomously

    AIGoogle researchers used the Aletheia agent, powered by Gemini 3 Deep Think, to attempt 10 FirstProof challenge problems without modification. The agent operated fully autonomously and solved 6 of the 10 problems, according to the post, with methodology and expert evaluations described in the linked arXiv paper.

    Why it matters: The post gives the autonomous setup and expert-evaluated results for an AI agent on FirstProof math problems, useful for judging how far such systems go on research-level math.

Feb 19

Feb 19Thu
  1. Guillaume LampleAI score40

    Mistral releases Voxtral Realtime paper, Apache 2.0 speech model

    AIMistral has published the technical report for Voxtral Realtime, a speech transcription model released under the Apache 2.0 license. The model reportedly achieves state-of-the-art transcription performance at sub-500ms latency. Mistral also launched a Realtime playground in Mistral Studio and made the model available in Hugging Face Transformers.

Feb 14

Feb 14Sat

Feb 13

Feb 13Fri
  1. MiniMax BlogAI score62

    MiniMax details Forge, a scalable agent RL framework behind M2.5

    AIMiniMax describes Forge, its internal reinforcement learning framework for training real-world agents, which was used during the development of MiniMax M2.5. The post explains a Windowed FIFO scheduler, prefix tree merging that the post says yields a 40x training speedup, and CISPO-based training across more than one hundred thousand agent scaffolds and environments.

    Why it matters: The post details how the Forge framework balances throughput, stability, and agent flexibility, with concrete scheduling and prefix-merging methods for training agent RL at scale.

Feb 11

Feb 11Wed
  1. Yi TayAI score67

    Aletheia math research agent produces two papers and solves open Erdős problems

    AIYi Tay introduces Aletheia, a math research agent powered by an advanced version of Gemini Deep Think. The post says it produced two publishable papers, one fully automatic and one human-AI collaboration, and solved multiple open Erdős problems. The attached image shows a Google DeepMind paper titled "Towards Autonomous Mathematics Research" with a generator, verifier, and reviser loop.

Feb 4

Feb 4Wed
  1. Anthropic EngineeringAI score72

    Anthropic finds container resource limits can shift agentic coding eval scores

    AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

    Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.

Feb 2

Feb 2Mon
  1. BAAIAI score43

    BAAI's Emu3 published in Nature, first Chinese-led large model paper there

    AIThe Beijing Academy of Artificial Intelligence (BAAI) published its Emu3 multimodal large model research in Nature, which the post describes as the first large-model achievement led by a Chinese research institution in that journal. Emu3 learns from text, image, and video at scale using next-token prediction alone, reaching generation and perception performance comparable to task-specific methods. The authors frame this as a step toward scalable, unified multimodal intelligence systems.

Jan 27

Jan 27Tue
  1. Tim DettmersAI score72

    Tim Dettmers Details How SERA Built an Open Coding Agent on 32 GPUs

    AIAi2's Open Coding Agents family, with SERA as its first release, was built by Tim Dettmers and collaborators on 32 GPUs. The method generates synthetic bug trajectories with soft verification, comparing patches by line overlap instead of running tests. The post reports that a 32B model fine-tuned on about 7,000 trajectories for one private repository matched its GLM 4.5-Air teacher, and that the baseline costs $500 to run.

Jan 10

Jan 10Sat
  1. Berkeley AI ResearchAI score36

    Information-Driven Design Framework Evaluates Imaging Systems by Mutual Information

    AIBerkeley AI Research proposes an information-based framework that evaluates and optimizes imaging systems using mutual information estimated directly from noisy measurements. The team reports that the metric predicts decoder performance across color photography, radio astronomy, lensless imaging, and microscopy, and that optimized designs match end-to-end methods while requiring less memory and compute.

Jan 9

Jan 9Fri
  1. BAAIAI score47

    DrugCLIP screens 10 trillion protein-molecule pairs per day for drug discovery

    AITsinghua AIR and BAAI's DrugCLIP screened 10,000 proteins against 500 million molecules, identifying over 2 million drug candidates. The post claims a 1-million-fold speedup, reaching 10 trillion protein-molecule pairs per day, and positions DrugCLIP as bridging AlphaFold structures to drug candidates. The work is published in Science, with a platform available at drugclip.com.

Dec 16, 2025

Dec 16, 2025Tue
  1. MiniMax · new models on Hugging FaceAI score38

    MiniMax Releases VTP-Large-f16d64 Visual Tokenizer With Technical Report and Pretrained Weights

    AIMiniMax released the technical report and pretrained weights for VTP-Large-f16d64, a visual tokenizer that jointly optimizes contrastive, self-supervised, and reconstruction losses. The model scores 78.2 zero-shot accuracy, 85.7 linear probing, and 0.36 rFID, and its generation performance scales with pretraining compute, parameters, and data. Checkpoint weights were listed as "released very soon" in the source.

Dec 12, 2025

Dec 12, 2025Fri
  1. Apple · new models on Hugging FaceAI score46

    Apple's SHARP Turns a Single Photo into a 3D Scene in Under a Second

    AIApple has released SHARP, a model that generates a 3D Gaussian representation of a scene from a single photograph in less than a second on a standard GPU. The output renders in real time as high-resolution photorealistic views of nearby camera positions, with metric absolute scale, and the paper reports reductions of 25–34% in LPIPS and 21–43% in DISTS versus the best prior model.

Dec 11, 2025

Dec 11, 2025Thu
  1. OpenAI · new models on Hugging FaceAI score42

    OpenAI Releases circuit-sparsity Sparse Model Weights on Hugging Face

    AIOpenAI has published weights for a sparse model from Gao et al. 2025, used for qualitative results on bracket counting and variable binding, on Hugging Face under the openai/circuit-sparsity repository. The release includes a standalone Hugging Face implementation that loads the converted model and tokenizer with trust_remote_code and runs sample generation. The project is licensed under Apache License 2.0.

Oct 27, 2025

Oct 27, 2025Mon

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    AIThinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    Why it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.

May 5, 2025

May 5, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score39

    Kevin-32B Uses Multi-Turn Reinforcement Learning to Write Faster CUDA Kernels

    AIStanford and Cognition AI researchers introduced Kevin-32B, a 32B-parameter model trained with multi-turn reinforcement learning to write CUDA kernels. On KernelBench, it solves 89% of tasks at best@16 and achieves 65% average correctness over eight refinement steps, versus 53% for o4-mini and 51% for o3. Its best@16 speedup is 1.41x, and multi-turn training outperforms single-turn training as refinement steps increase.

Nov 30, 2024

Nov 30, 2024Sat
  1. Liquid AI BlogAI score60

    Liquid AI's STAR uses evolutionary search to synthesize tailored model architectures

    AILiquid AI reports STAR, an evolutionary algorithm that synthesizes tailored neural network architectures from a numerical genome representation. The authors say it produced hundreds of designs that outperform Transformer and hybrid architectures in quality, with smaller caches and parameter counts, and can optimize for latency on target hardware. The full method is described in the arXiv technical report 2411.17800.

    Why it matters: The post explains how evolutionary search over a new architecture design space produced designs beating Transformers and hybrids, giving a concrete method for quality versus latency and memory trade-offs.

Mar 14, 2024

Mar 14, 2024Thu
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition reports Devin resolves 13.86% of SWE-bench issues end to end

    AICognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.

    Why it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.

Aug 11, 2021

Aug 11, 2021Wed