Skip to contentSkip to stories

Updated

#Paper/Research

Showing low-relevance items too. Hide low-relevance items

Sep 24

Sep 24Thu
  1. Goodfire ResearchOfficialAI score58

    Llama 3.1 8B Uses a Shared Circular Addition Module for Calendar Arithmetic

    AIGoodfire researchers found that Llama 3.1 8B solves month and day arithmetic, such as six months after August, through a single addition module in layer 18. The module represents numbers as Fourier-feature circles and computes modular sums in parallel, and steering those circles changed the model's predicted month.

  2. Goodfire ResearchOfficialAI score48

    Steering Along Manifolds Beats Linear Steering for Controlling Llama's Days-of-Week Behavior

    AIGoodfire Research shows that steering Llama-3.1 8B along the curved representation manifold of weekdays produces output probabilities that follow the model's natural cyclic behavior, shifting probability mass smoothly from Monday to Tuesday to Friday. Linear steering along a straight vector, by contrast, cuts across the behavior manifold and yields noisy off-target tokens, some not days of the week at all. The authors argue that representation geometry and behavior geometry are linked bidirectionally.

  3. Goodfire ResearchOfficialAI score54

    LLM activations trace emotional story arcs through neural geometry over time

    AIGoodfire Research examines how language models track emotional dynamics across a story, sentence by sentence. The authors prompt Llama 3.1 8B to rate six emotions after each sentence, then harvest activations from each sentence's last token and fit a manifold to show stories tracing trajectories through it.

  4. Goodfire ResearchOfficialAI score57

    Goodfire finds sparse autoencoder features capture curved neural geometry in three ways

    AIGoodfire Research examines how sparse autoencoder directions relate to curved manifolds in neural representations, identifying shattering, compact capture, and dilution as three ways lines can represent them. The team trained an autoencoder on synthetic data containing shapes such as donuts, spheres, and Möbius strips, and reports that real features in Llama 3.1 8B show dilution. It also describes an unsupervised pipeline that clusters features by firing patterns to surface manifolds in that model.

  5. AI at MetaOfficialAI score34

    Muse Realtime Avatar beats two commercial avatar systems in live-call tests

    AIMeta's Muse Realtime Avatar was rated ahead of two leading commercial avatar systems in live-call products, based on 2–3 minute conversations with matched avatar identities. Raters compared visual quality, sync, character consistency, and mannerisms, and Muse Realtime Avatar came out ahead on overall preference.

    Image from @AIatMeta's post
  6. AI at MetaOfficialAI score42

    Meta unveils Muse Realtime Voice and Avatar with shared speech-token streaming

    AIMeta's Muse Realtime Voice generates speech tokens encoding both content and prosody, and Muse Realtime Avatar consumes that shared stream to produce streaming video. Using a fixed-length history as motion context keeps computation bounded regardless of conversation length while synchronizing voice, lip motion, and expressions.

    Image from @AIatMeta's post
  7. AI at MetaOfficialAI score22

    Meta distills 40-step video diffusion into a 2-step live streaming model

    AIMeta distilled a 40-step diffusion teacher using 3-way CFG, requiring 120 evaluations per video chunk, into an unguided 2-step causal student with a fixed-length KV cache. The student uses self-forcing to resist drift and keep near-teacher quality while needing 60x fewer evaluations, enabling instant responses in live video streaming.

    Image from @AIatMeta's post
  8. Anthropic ResearchOfficialAI score60

    Anthropic study finds Claude agent trading limited by preference understanding

    AIAnthropic ran a controlled book-swapping market with 201 employees and Claude-powered agents, which reached 0.55 efficiency against a 0.89 optimum. Agents matched participants' own rankings on 61% of book pairs, and about 85% of the shortfall came from imprecise preference representation rather than the trading floor design. Stronger models produced more efficient markets than weaker ones, while instructions mattered less.

    Why it matters: The study separates agent misunderstanding of user preferences from negotiation failure, showing which failure mode limits outcomes in agent-run markets.

Sep 23

Sep 23Wed
  1. Tencent HyOfficialAI score38

    Tencent Hunyuan studies batch-size scaling for LLM reinforcement learning efficiency

    AITencent Hunyuan extends classical critical-batch-size theory to online LLM reinforcement learning, where models generate their own training data. Across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded range of batch sizes. On fixed hardware, larger batches raise PPO generation-stage throughput by up to 2.29×, and the best measured GRPO setup reaches the same validation target in 29% less time.

  2. Felix RiesebergXAI score38

    Anthropic's biology team discovers new enzyme system called ART

    AIAnthropic's experimental biology team reportedly discovered array-associated reverse transcriptases (ART), an enzyme system no human scientists had previously reported. The post, from Anthropic employee Felix Rieseberg, also praises Claude's contributions during the research.

    Image from @felixrieseberg's post
  3. Google Developers BlogOfficialAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    AIGoogle Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    Why it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  4. Dario AmodeiXAI score76

    Claude Helps Discover a Possible New Gene Editing Enzyme System

    AIAnthropic announced that Claude, working mostly on its own, identified a previously unknown enzyme system in bacteriophage DNA that may represent a new gene editing mechanism. Claude read literature and genome data, proposed experiments, and Anthropic's team carried them out. The function and biotechnological utility of the system remain unclear.

    Why it matters: The post pairs a Claude-led discovery with the lab workflow used to verify it, showing how AI and humans split the research work in biology.

  5. AnthropicOfficialAI score62

    Claude finds a previously unknown enzyme system in bacteriophage DNA

    AIClaude has identified a previously unknown enzyme system in bacteriophage DNA, located beside a long array of repeating DNA that somewhat resembles CRISPR. Anthropic says its function is not yet understood, but only a handful of known systems share its features, all of which can cut, copy, and paste DNA. The source notes that programmable systems like CRISPR have been important to medicine, but more work is needed to learn what this system does and whether it can be used similarly.

  6. Anthropic · YouTubeOfficialAI score65

    Anthropic launches a molecular biology lab where Claude hunts for unusual proteins

    AIAnthropic is introducing a molecular biology research group and lab to test whether Claude can help scientists find unusual proteins. Claude combs through large DNA datasets, flags uncharacterized proteins, and passes its most promising ideas to scientists, who test them at the bench. In one early program, Claude discovered a novel enzyme system with CRISPR-like repeats.

    Why it matters: The source shows Claude being used in a wet-lab workflow, from scanning DNA datasets to flagging proteins for scientists to test at the bench.

  7. ModelScopeOfficialAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    AIShanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

    Image from @ModelScope2022's post
  8. Anthropic NewsroomOfficialAI score73

    Claude agents discover a novel CRISPR-like enzyme system in bacteriophages

    AIAnthropic's new life sciences group reports that Claude autonomously identified a previously uncharacterized enzyme system, called array-associated reverse transcriptase (ART), in bacteriophages. Claude agents searched over 200,000 reverse transcriptases, narrowed 3,500 candidates to 20, and one agent flagged a CRISPR-like repeat array after about 21 hours. Human scientists then validated the finding in the lab, and the function of ART remains unknown.

    Why it matters: The post shows how Claude agents surveyed DNA sequence data, flagged a candidate, and then led to lab validation, which is a concrete workflow for AI-assisted biology research.

Sep 22

Sep 22Tue
  1. Redwood Research BlogBlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.

  2. whXAI score34

    MiMo-V2.6 paper details data and RL results for open model

    AIThe MiMo-V2.6 paper thread reports on the newest open model, which also streams its RL run, focusing on data and RL experimental results rather than architecture. The Pro model reportedly rose from 58.41 to 72.57 on DeepSWE after RL, with the top published DeepSWE score cited at 74.

    Image from @nrehiew_'s post
  3. Tencent HyOfficialAI score44

    WebCraftBench Scores AI-Built Websites by Live Use and Human Preference

    AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.

Sep 21

Sep 21Mon
  1. Xiaomi MiMoOfficialAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

    Image from @XiaomiMiMo's post
  2. Amazon ScienceOfficialAI score47

    Amazon Bio Discovery's three AI methods accelerate antibody drug design

    AIAmazon Bio Discovery developed three AI approaches for antibody drug design: MochiBind for sequence-based affinity ranking, CA-MAP for developability prediction with batch effect correction, and an agent-guided design system. The agent-guided system produced 46 lab-validated hits against a novel cancer target.

  3. Amazon ScienceOfficialAI score60

    Amazon Science reports AI models for designing and characterizing antibodies

    AIAmazon Science describes three papers on AI for antibody discovery: MochiBind ranks antibody binding strength from sequence alone, CA-MAP predicts developability properties using batch-aware context, and an agent-guided pipeline designed nanobody binders against a novel cancer target. In the pipeline, 116 candidates survived lab screening, and 46 were identified as strong binders, which are being used to train the next design cycle.

    Why it matters: The source reports the method, benchmark setup, and experimental validation in a single design workflow, showing how predictors, agents, and lab screening connect in antibody discovery.

  4. Microsoft ResearchOfficialAI score34

    RetroChimera model aims to speed up custom molecule synthesis

    AIMicrosoft Research published a Nature paper on RetroChimera, a predictive model designed to accelerate chemical synthesis. The post says custom-made molecules advance medicine, materials, and agriculture, but producing them remains slow and expensive. The model is meant to help researchers explore a wider range of molecules.

    Video from @MSFTResearch's post
  5. Microsoft ResearchOfficialAI score50

    Microsoft Research open-sources RetroChimera, a retrosynthesis model published in Nature

    AIMicrosoft Research published RetroChimera, a retrosynthesis framework that combines the R-SMILES 2 Transformer model and the NeuralLoc graph neural network through learned ensembling to propose synthesis routes for small molecules. In blind tests, PhD-level chemists preferred its individual reaction predictions over those from preceding models and recorded literature reactions. The implementation and weights are open-sourced for researchers developing new medicinal molecules and materials.

Sep 19

Sep 19Sat
  1. Sebastian RaschkaXAI score42

    Muon reduces memorization compared with AdamW in nanoGPT training experiments

    AIMuon appears to outperform AdamW because it suppresses memorization, according to WeightWatcher experiments on a single-head nanoGPT model across five seeds. At 10,000 steps, teacher-forced recall of planted sequences was about 62% for AdamW versus under 1% for Muon. The author notes that some Muon layers also show α < 2, so α alone does not explain memorization and individual layers and their ESDs should be examined.

Sep 18

Sep 18Fri
  1. Google ResearchOfficialAI score22

    Google Research releases MilleMiglia, a public middle-mile logistics benchmark

    AIGoogle Research has introduced MilleMiglia, a standardized benchmark for optimizing middle-mile logistics, the segment that moves goods across hundreds of miles overnight. The benchmark uses spatial clustering and gravity models to simulate realistic middle-mile delivery scenarios. It addresses the difficulty of optimizing these networks without public data.

    Image from @GoogleResearch's post
  2. SemiAnalysisBlogAI score52

    Engram offloading to DRAM beats SSD for DeepSeek-V4.1-Flash serving on B200

    AISemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.

Sep 17

Sep 17Thu
  1. Google ResearchOfficialAI score52

    Google Research enables teachers to create generative UI learning interactives

    AIGoogle Research is sharing an experiment that lets educators generate custom, guided STEM simulations tailored to their curriculum using generative UI. It is releasing a sample library of over 30 English interactives for physics, chemistry, biology, and math, all AI-generated and reviewed by teachers. Schools using Google Workspace for Education can sign up through the Google for Education Pilot Program to give feedback.

  2. SenseTimeOfficialAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

    Image from @SenseTime_AI's post
  3. Google DeepMindOfficialAI score47

    AlphaGenome Atlas boosts rare genetic signal detection in UK Biobank study

    AIResearchers at the University of Exeter, working with Google DeepMind, used AlphaGenome Atlas on data from more than 54,000 UK Biobank participants. The approach raised detection of rare genetic signals by over 22% and uncovered new DNA variants that influence levels of PLA2G7, a protein linked to metabolic health.

    Image from @GoogleDeepMind's post

Sep 16

Sep 16Wed
  1. Matei ZahariaXAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

Sep 15

Sep 15Tue
  1. Stability AIOfficialAI score5

    Stability AI shares a link to an ECCV 2026 poster presentation

    AIStability AI posted a closing message about a productive week of conversations and linked to a poster from the ECCV 2026 virtual conference. The post does not describe the research's content, results, or specific model names.

    Image from @StabilityAI's post
  2. Stability AIOfficialAI score25

    Stability AI's team presents color consistency research at ECCV

    AIStability AI's interactive research team presented new work at the 19th European Conference on Computer Vision (ECCV) aimed at keeping colors consistent across shots and AI-generated reference photos. The post frames color consistency as a persistent friction point in production.

    Image from @StabilityAI's post
  3. TinkerOfficialAI score34

    Trained-on human stories shape how AI assistants behave in chat

    AIA Truthful AI paper trained models only on synthetic stories about humans, with no AI characters, and found the Assistant adopted quirky behaviors from those stories in ordinary chat. Adoption was stronger for characters from elite schools, according to Owain Evans. The post presents this as an interpretability result that adds to and complicates the Persona Selection Model.

  4. Lewis Tunstall @ COLM 🌉XAI score30

    Periodic Labs advances toward cracking condensed matter physics superconductor problem

    AIPeriodic Labs, the team behind high-throughput materials labs in Menlo Park, reports progress on one of condensed matter physics' hardest problems. Its open-source model Neon, trained with mid-training and RL on 1,300 H200s plus months of lab data, surpasses GPT-6 Astra on the company's analysis benchmark. The work targets materials science challenges including superconductors, magnets, and semiconductors.