Skip to contentSkip to stories

Updated

#Paper/Research

Showing low-relevance items too. Hide low-relevance items

Apr 23

Apr 23Thu
  1. OpenAI Alignment Research BlogAI score44

    OpenAI Open-Sources Chain-of-Thought Monitorability Evaluation Datasets and Code

    AIOpenAI is releasing a subset of datasets, reference code, and the g-mean 2 metric for evaluating chain-of-thought monitorability. The release includes most datasets from its monitorability suite, while some evaluations relying on private or restricted data were omitted. The company says it will keep reporting monitorability results in future frontier reasoning model system cards.

Apr 21

Apr 21Tue

Apr 20

Apr 20Mon
  1. Berkeley AI ResearchAI score44

    GRASP: A Gradient-Based Planner for Long-Horizon World Model Planning

    AIBerkeley AI Research introduces GRASP, a gradient-based planner for learned world models that aims to make long-horizon planning more robust. GRASP lifts trajectories into virtual states for parallel optimization across time, adds stochasticity to state iterates for exploration, and reshapes gradients to avoid brittle state-input gradients through high-dimensional vision models. The post identifies ill-conditioned gradients and non-greedy loss landscapes as core failure modes of standard rollout-based planning.

Apr 16

Apr 16Thu

Apr 14

Apr 14Tue

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    AICognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    Why it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

Mar 31

Mar 31Tue
  1. Intern Large ModelsAI score52

    Intern Large Models unveils Kernel-Smith for generating GPU kernels and operators

    AIIntern Large Models introduced Kernel-Smith, a framework for generating high-performance GPU kernels and operators using an evolutionary agent and post-training recipe. The post says it outperforms Gemini-3.0-pro and Claude-4.6-opus on Kernel-Bench, and that optimized kernels have been merged into SGLang and LMDeploy. The accompanying figure compares best program score trajectories across evolution steps, with Kernel-Smith-235B-RL reaching the highest peak.

    Image from @intern_lm's post

Mar 26

Mar 26Thu
  1. Guillaume Lample @ NeurIPS 2024AI score62

    Mistral releases Voxtral TTS, its first open-weight speech model

    AIMistral's Voxtral TTS is its first speech model, presented as an open-weight text-to-speech model that reportedly delivers SOTA performance at significantly lower cost with very low latency. It combines autoregressive generation of semantic speech tokens with flow-matching for acoustic tokens, and a technical report on its training methodology is being released.

    Image from @GuillaumeLample's post

Mar 25

Mar 25Wed

Mar 24

Mar 24Tue
  1. ARC PrizeAI score70

    ARC Prize announces ARC-AGI-3, an interactive benchmark for frontier agents

    AIARC Prize has released ARC-AGI-3, a set of hundreds of interactive, turn-based environments with thousands of game-style levels, with no instructions or stated goals. Humans score 100% while frontier AI scores 0.51%. ARC Prize 2026 offers over $2 million in prizes for open-source solutions to ARC-AGI-2 and ARC-AGI-3.

    Why it matters: The benchmark's human versus frontier AI gap and its interactive design show how agent evaluation is shifting from instruction-following toward exploration and adaptation.

Mar 19

Mar 19Thu
  1. Tri DaoAI score52

    Tri Dao Says Nonlinear RNNs Differ From Attention and Linear SSMs

    AITri Dao says nonlinear RNNs seem to do something genuinely different from attention and linear RNNs or SSMs. He reports they already perform well with the right parametrization, and adding just one nonlinear RNN layer substantially improves a transformer-Mamba/DeltaNet hybrid. The post quotes the M²RNN paper, which introduces non-linear RNNs with matrix-valued states for language modeling, with links to the paper, code, and models.

Mar 17

Mar 17Tue
  1. BAAIAI score46

    BAAI unveils RoboBrain-Dex, dexterous manipulation trained on human egocentric data

    AIBAAI has released RoboBrain-Dex, a dexterous manipulation model for embodied intelligence trained on large-scale, diverse human egocentric data rather than massive robot teleoperation datasets. BAAI says this shifts robotic dexterous manipulation research from small data with weak generalization to big data with strong generalization. The code is open-sourced on GitHub.

Mar 13

Mar 13Fri
  1. Berkeley AI ResearchAI score34

    SPEX and ProxySPEX Identify Influential LLM Interactions at Scale with Fewer Ablations

    AIBerkeley AI Research introduces SPEX, a signal-processing framework that identifies influential interactions in LLMs using far fewer ablations than exhaustive analysis. A hierarchy-based extension, ProxySPEX, matches SPEX performance with around 10x fewer ablations. The methods apply to feature, data, and model component attribution.

Mar 5

Mar 5Thu
  1. Tri DaoAI score62

    FlashAttention-4 paper: attention on Blackwell GPUs nears matmul speed

    AIThe FlashAttention-4 paper is out, reporting that attention on Blackwell GPUs now runs at roughly matmul speed, reaching about 1600 TFLOPs. The forward pass is bottlenecked by exponential computation and the backward pass by shared memory bandwidth, and the redesign uses polynomial exponential emulation, a new online softmax that avoids 90% of softmax rescaling, and 2CTA MMA instructions that let two thread blocks share operands to cut shared memory traffic.

Mar 4

Mar 4Wed
  1. Tri DaoAI score62

    Tri Dao Shares Speculative Speculative Decoding, a Claimed Up-to-2x LLM Inference Speedup

    AITri Dao reposts a quoted post from @tanishqkumar07 introducing Speculative Speculative Decoding (SSD), an LLM inference algorithm claimed to be up to 2x faster than leading inference engines. The quoted post credits collaborators @tri_dao and @avnermay and links to a thread with details. Tri Dao's own text says the approach applies an asynchronous-machines principle seen in GPU kernels to speculative decoding.

Feb 25

Feb 25Wed
  1. Jim FanAI score75

    EgoScale trains a 22-DoF humanoid mostly on 20,000 hours of human video

    AIResearchers trained a humanoid with 22-DoF dexterous hands mainly on over 20,000 hours of egocentric human video, with no robot in the loop, to perform tasks such as assembling model cars and folding shirts. They report a log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and state that this loss predicts real-robot success rate. The recipe, called EgoScale, pre-trains GR00T N1.5 on the video, adds only 4 hours of robot play data, and reports a 54% gain over training from scratch across five dexterous tasks.

    Video from @DrJimFan's post
  2. Quoc LeAI score65

    Aletheia Agent Solves 6 of 10 FirstProof Math Problems Autonomously

    AIGoogle researchers used the Aletheia agent, powered by Gemini 3 Deep Think, to attempt 10 FirstProof challenge problems without modification. The agent operated fully autonomously and solved 6 of the 10 problems, according to the post, with methodology and expert evaluations described in the linked arXiv paper.

    Why it matters: The post gives the autonomous setup and expert-evaluated results for an AI agent on FirstProof math problems, useful for judging how far such systems go on research-level math.

    Image from @quocleix's post

Feb 19

Feb 19Thu
  1. Guillaume Lample @ NeurIPS 2024AI score40

    Mistral releases Voxtral Realtime paper, Apache 2.0 speech model

    AIMistral has published the technical report for Voxtral Realtime, a speech transcription model released under the Apache 2.0 license. The model reportedly achieves state-of-the-art transcription performance at sub-500ms latency. Mistral also launched a Realtime playground in Mistral Studio and made the model available in Hugging Face Transformers.

    Image from @GuillaumeLample's post

Feb 14

Feb 14Sat

Feb 13

Feb 13Fri
  1. MiniMax BlogAI score62

    MiniMax details Forge, a scalable agent RL framework behind M2.5

    AIMiniMax describes Forge, its internal reinforcement learning framework for training real-world agents, which was used during the development of MiniMax M2.5. The post explains a Windowed FIFO scheduler, prefix tree merging that the post says yields a 40x training speedup, and CISPO-based training across more than one hundred thousand agent scaffolds and environments.

    Why it matters: The post details how the Forge framework balances throughput, stability, and agent flexibility, with concrete scheduling and prefix-merging methods for training agent RL at scale.

Feb 11

Feb 11Wed
  1. Yi TayAI score67

    Aletheia math research agent produces two papers and solves open Erdős problems

    AIYi Tay introduces Aletheia, a math research agent powered by an advanced version of Gemini Deep Think. The post says it produced two publishable papers, one fully automatic and one human-AI collaboration, and solved multiple open Erdős problems. The attached image shows a Google DeepMind paper titled "Towards Autonomous Mathematics Research" with a generator, verifier, and reviser loop.

    Image from @YiTayML's post

Feb 4

Feb 4Wed
  1. Anthropic EngineeringAI score72

    Anthropic finds container resource limits can shift agentic coding eval scores

    AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

    Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.

Feb 2

Feb 2Mon
  1. BAAIAI score43

    BAAI's Emu3 published in Nature, first Chinese-led large model paper there

    AIThe Beijing Academy of Artificial Intelligence (BAAI) published its Emu3 multimodal large model research in Nature, which the post describes as the first large-model achievement led by a Chinese research institution in that journal. Emu3 learns from text, image, and video at scale using next-token prediction alone, reaching generation and perception performance comparable to task-specific methods. The authors frame this as a step toward scalable, unified multimodal intelligence systems.

    Video from @BAAIBeijing's post

Jan 27

Jan 27Tue
  1. Tim DettmersAI score72

    Tim Dettmers Details How SERA Built an Open Coding Agent on 32 GPUs

    AIAi2's Open Coding Agents family, with SERA as its first release, was built by Tim Dettmers and collaborators on 32 GPUs. The method generates synthetic bug trajectories with soft verification, comparing patches by line overlap instead of running tests. The post reports that a 32B model fine-tuned on about 7,000 trajectories for one private repository matched its GLM 4.5-Air teacher, and that the baseline costs $500 to run.

Jan 10

Jan 10Sat
  1. Berkeley AI ResearchAI score36

    Information-Driven Design Framework Evaluates Imaging Systems by Mutual Information

    AIBerkeley AI Research proposes an information-based framework that evaluates and optimizes imaging systems using mutual information estimated directly from noisy measurements. The team reports that the metric predicts decoder performance across color photography, radio astronomy, lensless imaging, and microscopy, and that optimized designs match end-to-end methods while requiring less memory and compute.

Jan 9

Jan 9Fri
  1. BAAIAI score47

    DrugCLIP screens 10 trillion protein-molecule pairs per day for drug discovery

    AITsinghua AIR and BAAI's DrugCLIP screened 10,000 proteins against 500 million molecules, identifying over 2 million drug candidates. The post claims a 1-million-fold speedup, reaching 10 trillion protein-molecule pairs per day, and positions DrugCLIP as bridging AlphaFold structures to drug candidates. The work is published in Science, with a platform available at drugclip.com.

    Image from @BAAIBeijing's post

Dec 16, 2025

Dec 16, 2025Tue
  1. MiniMax · new models on Hugging FaceAI score38

    MiniMax Releases VTP-Large-f16d64 Visual Tokenizer With Technical Report and Pretrained Weights

    AIMiniMax released the technical report and pretrained weights for VTP-Large-f16d64, a visual tokenizer that jointly optimizes contrastive, self-supervised, and reconstruction losses. The model scores 78.2 zero-shot accuracy, 85.7 linear probing, and 0.36 rFID, and its generation performance scales with pretraining compute, parameters, and data. Checkpoint weights were listed as "released very soon" in the source.

Dec 12, 2025

Dec 12, 2025Fri
  1. Apple · new models on Hugging FaceAI score46

    Apple's SHARP Turns a Single Photo into a 3D Scene in Under a Second

    AIApple has released SHARP, a model that generates a 3D Gaussian representation of a scene from a single photograph in less than a second on a standard GPU. The output renders in real time as high-resolution photorealistic views of nearby camera positions, with metric absolute scale, and the paper reports reductions of 25–34% in LPIPS and 21–43% in DISTS versus the best prior model.

Dec 11, 2025

Dec 11, 2025Thu
  1. OpenAI · new models on Hugging FaceAI score42

    OpenAI Releases circuit-sparsity Sparse Model Weights on Hugging Face

    AIOpenAI has published weights for a sparse model from Gao et al. 2025, used for qualitative results on bracket counting and variable binding, on Hugging Face under the openai/circuit-sparsity repository. The release includes a standalone Hugging Face implementation that loads the converted model and tokenizer with trust_remote_code and runs sample generation. The project is licensed under Apache License 2.0.

Oct 27, 2025

Oct 27, 2025Mon

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    AIThinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    Why it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.