Skip to contentSkip to stories

Updated

#Paper/Research

Showing low-relevance items too. Hide low-relevance items

Sep 15

Sep 15Tue
  1. Stability AIOfficialAI score5

    Stability AI shares a link to an ECCV 2026 poster presentation

    AIStability AI posted a closing message about a productive week of conversations and linked to a poster from the ECCV 2026 virtual conference. The post does not describe the research's content, results, or specific model names.

    Image from @StabilityAI's post
  2. Stability AIOfficialAI score25

    Stability AI's team presents color consistency research at ECCV

    AIStability AI's interactive research team presented new work at the 19th European Conference on Computer Vision (ECCV) aimed at keeping colors consistent across shots and AI-generated reference photos. The post frames color consistency as a persistent friction point in production.

    Image from @StabilityAI's post
  3. TinkerOfficialAI score34

    Trained-on human stories shape how AI assistants behave in chat

    AIA Truthful AI paper trained models only on synthetic stories about humans, with no AI characters, and found the Assistant adopted quirky behaviors from those stories in ordinary chat. Adoption was stronger for characters from elite schools, according to Owain Evans. The post presents this as an interpretability result that adds to and complicates the Persona Selection Model.

  4. Lewis Tunstall @ COLM 🌉XAI score30

    Periodic Labs advances toward cracking condensed matter physics superconductor problem

    AIPeriodic Labs, the team behind high-throughput materials labs in Menlo Park, reports progress on one of condensed matter physics' hardest problems. Its open-source model Neon, trained with mid-training and RL on 1,300 H200s plus months of lab data, surpasses GPT-6 Astra on the company's analysis benchmark. The work targets materials science challenges including superconductors, magnets, and semiconductors.

  5. Tencent HyOfficialAI score38

    EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle

    AITencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.

    Image from @TencentHunyuan's post

Sep 14

Sep 14Mon
  1. TinkerOfficialAI score44

    RLVR trains models to design power transformers with physics-based verifiers

    AITinker says engineers are using physics-based verifiers and Tinker to train models that design power transformers meeting specifications at low cost. Background from @gentrajectory says an RL-trained Kimi base model met 93% of unseen transformer specs, compressing multi-week engineering work into minutes of inference.

Sep 12

Sep 12Sat
  1. Epoch AI · The Epoch BriefOfficialAI score60

    Epoch Brief covers Huawei chips, Nvidia's GDP effect, and GPT-6 Astra benchmarks

    AIEpoch AI's newsletter reports that Huawei is far behind Nvidia and is unlikely to catch up this decade due to export controls. It also finds official US GDP statistics understate growth by about 0.3 percentage points over the past year, and that GPT-6 Astra set new records on Epoch's evaluations, including the Epoch Capabilities Index.

    Why it matters: The newsletter bundles several analyses of AI chips, GDP measurement, and benchmarks, so it helps readers scan the research agenda behind each finding.

Sep 11

Sep 11Fri
  1. Redwood Research BlogBlogAI score62

    Prompt tuning lifts CoT controllability scores on open models

    AIRedwood Research reports that better prompt templates raise chain-of-thought controllability scores on the CoTControl eval for open-source reasoning models by roughly 2-3x or more. For example, GPT-OSS-120B rose from 5.5% to 15% in the zero-shot setting. The author concludes that current CoT controllability numbers may underestimate what models can do, though the finding does not significantly undermine the view that current models probably cannot consistently evade CoT monitoring.

Sep 10

Sep 10Thu
  1. Ai2 · new models on Hugging FaceOfficialAI score34

    AstaBrief-8B-SFT: Ai2's 8B model for cited scientific research reports

    AIAi2 released AstaBrief-8B-SFT, an 8B intermediate supervised fine-tuning checkpoint built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. On the ScholarQA-CS2 test set of 100 computer science questions, it scored an average of 83.7 versus 77.3 for base Qwen3-8B, with citation recall at 71.3 versus 64.6. The model is licensed under Apache 2.0 for research and educational use.

  2. Amazon ScienceOfficialAI score40

    Amazon research explains why ML research agents don't overfit benchmarks

    AIAmazon Science researchers propose that machine learning research agents avoid overfitting benchmarks despite years of iteration against the same tests. They attribute this to generalizable strategies being expressed compactly, leaving no room for memorization, while overfitting strategies fail to survive a compression bottleneck.

  3. Redwood Research BlogBlogAI score62

    Redwood Research proposes NLS depth to measure opaque serial reasoning in AI models

    AIRedwood Research defines NLS depth, a measure of how much unverbalized serial computation an AI system can perform, building on Brown-Cohen et al.'s opaque serial depth. The proposal counts only nodes that output natural language initialized from a pre-training prior as interpretable bottlenecks. Standard transformers scale with their layer count, while latent reasoning architectures would raise NLS depth sharply.

  4. Amazon ScienceOfficialAI score55

    Research agents avoid overfitting when their winning strategies compress into few tokens

    AIAmazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.

Sep 9

Sep 9Wed
  1. TinkerOfficialAI score28

    Tinker and OpenResearch automate auditing of self-distillation methods

    AITinker says it and OpenResearch let agents test dozens of competing published post-training methods automatically, with compute cost forecast to within a dollar. The main post cites a grant-supported effort, while the quoted alphaXiv post says agents reproduced SDFT's continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score82

    Cognition's Devin factors RSA-260 using a GPU lattice siever

    AICognition's Devin agent, directed by Eric Lu, factored the 260-digit RSA-260 number using a new GPU implementation of the general number field sieve built on CADO-NFS. The author estimates the run cost about 13.5 GPU-years, roughly $400k at market prices, and projects RSA-1024 factoring at around $30M, while RSA-2048 is not meaningfully affected.

    Why it matters: The source gives a full cost breakdown and scaling estimates for RSA factoring on GPUs, showing how far the cost of breaking RSA-1024 has fallen.

  3. Sebastian RaschkaXAI score14

    Raschka's mega write-up on GPT-6 Astra and looped transformers

    AISebastian Raschka published a long write-up covering how looped transformers and recurrent depth work, along with their cost tradeoffs. The post also examines whether these architectures hide reasoning traces and surveys recent looped transformer research, with many figures included.

    Image from @rasbt's post
  4. Ahead of AI (Sebastian Raschka)BlogAI score46

    GPT-6 Astra Leads Coding and Math Benchmarks, Shows Strong Computer Use

    AIOpenAI's GPT-6 Astra scores 99.9% on ARC-AGI-3, versus 7.8% for GPT-5.6 Sol, and leads Raschka's coding and math tests. Its strongest showing is in graphics and computer-use tasks, such as redrawing an image in a browser-based Paint app. The author notes that Artificial Analysis shows Astra at the frontier but not pulling far ahead on its Coding Agent Index.

  5. Ai2 (Allen Institute for AI)OfficialAI score39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    AIGoodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.

Sep 8

Sep 8Tue
  1. BAAIOfficialAI score34

    Robot models excel at single moves but fail chained tasks

    AIBAAI reports that robot models trained on individual skills such as grasping, placing, pulling, and opening performed poorly when asked to chain them into full tasks without extra practice. The best score was 16.7%, and some models scored zero. The post's example notes a robot may open a drawer yet get stuck on the handle.

    Image from @BAAIBeijing's post
  2. Google DeepMind · YouTubeOfficialAI score78

    DeepMind releases AlphaGenome Atlas, a predictive map of every possible DNA letter change

    AIGoogle DeepMind has used AlphaGenome to predict the molecular impact of every possible single-letter change in the human genome, around nine billion variants. The resulting AlphaGenome Atlas is a 1PB dataset that assigns each variant an AlphaGenome Variant Impact (AVI) score, covering both coding and non-coding variations, and is available to researchers worldwide. The video notes that AlphaGenome has not been validated or approved for any clinical use.

    Why it matters: The release supplies a precomputed impact score for every possible single-letter genome change, which lets researchers look up variants without running the model themselves.

Sep 5

Sep 5Sat
  1. AI at MetaOfficialAI score46

    AIRA₃ cuts GPU kernel latency 27% and reaches Kaggle gold level

    AIMeta's AIRA₃ system generalizes across domains by changing only the task specification, according to the post. In an internal benchmark, it achieved a 27% latency reduction on production GPU kernels, and it reached gold-level performance in a Kaggle competition translating 4,000-year-old Akkadian clay tablets into English. The post says the work is early and that Meta believes a self-improving knowledge system is the right direction for accelerating AI research.

  2. AI at MetaOfficialAI score43

    AIRA₃ coordinates long-running agents through a shared forum and filesystem

    AIMeta's AIRA₃ replaces a central controller with many long-running agents, each pairing a model with a coding harness in its own isolated environment. The agents coordinate asynchronously through a shared forum for hypotheses and findings and a shared filesystem for solution artifacts. According to the post, performance gains compound over time as agents build on each other's discoveries.

    Image from @AIatMeta's post

Sep 4

Sep 4Fri
  1. John SchulmanXAI score34

    Schulman praises metric and dataset for training models to explain behavior

    AIJohn Schulman says a metric for explanation quality, centered on counterfactual simulatability, enables hillclimbing, and praises Adam et al. for a more diverse and realistic dataset and pipeline. He notes that models can be trained to write better post-hoc explanations of their own behavior, as described in a linked thread by @a_karvonen. That thread reports training on thousands of self-explanations of in-the-wild behaviors, with generalization to held-out evals.

  2. BAAI · new models on Hugging FaceOfficialAI score26

    ConsiSpace: BAAI and Peking University release geometry-consistent video spatial reasoning model

    AIBAAI and Peking University researchers released official weights for ConsiSpace, a geometry-consistent multimodal framework for spatial reasoning in long-form visual observations. The model is described in the paper "ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning" (arXiv:2607.17599).

  3. Lewis Tunstall @ COLM 🌉XAI score60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

  4. Lewis Tunstall @ COLM 🌉XAI score22

    Research Preference Models Rank AI Research Ideas to Save Compute

    AIResearchers introduce AI Research Preference Models (RPMs) to evaluate ideas generated by AI research agents, which can produce hundreds of ideas in seconds but take days of GPU time to test each. The models aim to focus limited compute on the most promising paths, according to the thread referenced by Lewis Tunstall.

  5. Lewis Tunstall @ COLM 🌉XAI score46

    Meta paper uses research preference models to guide AI agents' experiments

    AILewis Tunstall praises a new Meta paper on research preference models (RPMs), which instill "research taste" in agents by treating experiments as tree nodes. An RPM acts as an LLM judge that selects the most promising candidate experiment before it is run, reducing wasted compute. Tunstall notes the resulting trajectories could train domain-specific RPMs, which would be valuable in hard fields such as the natural sciences.

    Image from @_lewtun's post
  6. Tencent · new models on Hugging FaceOfficialAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

  7. Tencent · new models on Hugging FaceOfficialAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

  8. Matei ZahariaXAI score36

    Lakebase VLDB paper details Neon's elastic database built on S3 storage

    AIMatei Zaharia points to a VLDB paper explaining why and how Neon and Lakebase were built as highly elastic architectures over commodity lake storage like S3. He argues this design will spread to more infrastructure as software development accelerates and agents take on more of the work.

Sep 3

Sep 3Thu
  1. TinkerOfficialAI score23

    Tinker used to test counterfactual simulatability for LLM interpretability

    AITinker, the platform from @tinkerapi, supported two recent papers testing counterfactual simulatability as a way to interpret LLM behavior. The core idea is that understanding a model means predicting how its output changes when the prompt changes, with causes ranging from specific words to abstract properties such as a user's angry tone.

  2. TinkerOfficialAI score51

    Bespoke Labs post-trains Inkling on one code repo and reports broader coding gains

    AIBespoke Labs post-trained the Inkling base model on a single GitHub repository using supervised fine-tuning and GRPO reinforcement learning. The post reports a 57-point improvement on the held-out fontTools evaluation over the base model, along with gains on Terminal-Bench 2.1 and SWE-bench Lite. It also says the post-trained model uses about 40% fewer tokens.

    Image from @tinkerapi's post