Skip to content

#Paper/Research

Oct 8

TodayOct 8Thu61 items
  1. SantiagoAI score40

    Physical consistency is the most important feature of a world model, and the hardest to get right. That's why you see videos with objects defying gravity and people posing in impossible ways. Here is a complete evaluation of existing world models. Seedance 2.5 is the best right now.

    Physical consistency is the most important feature of a world model, and the hardest to get right. That's why you see videos with objects defying gravity and people posing in impossible ways. Here is a complete evaluation of existing world models. Seedance 2.5 is the best right now.

  2. GoogleAI score40

    In our prospective clinical study, our models were able to pinpoint the gestational age within 4 days of accuracy which is a tight enough zone that can have a really meaningful clinical impact. If we're able to expand these tools to low resource settings, we can start to move the needle on maternal deaths and bridge that gap in care that we see all over the world.

    In our prospective clinical study, our models were able to pinpoint the gestational age within 4 days of accuracy which is a tight enough zone that can have a really meaningful clinical impact. If we're able to expand these tools to low resource settings, we can start to move the needle on maternal deaths and bridge that gap in care that we see all over the world.

  3. ZyphraAI score34

    Our approach is lossless. The same tokens visit the same experts. The architecture, routing decisions and training objective stay unchanged. The model computes the same function. By organizing where experts and tokens live we reduce the communication needed to do the same work.

    Our approach is lossless. The same tokens visit the same experts. The architecture, routing decisions and training objective stay unchanged. The model computes the same function. By organizing where experts and tokens live we reduce the communication needed to do the same work.

  4. ZyphraAI score38

    The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster

    The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster

  5. ZyphraAI score22

    Crucially, we find that these patterns emerge early enough to be useful during training and are cheap to measure. Using just 4,096 tokens, our placement and routing predictions come within roughly one percentage point of results using 524,000 tokens.

    Crucially, we find that these patterns emerge early enough to be useful during training and are cheap to measure. Using just 4,096 tokens, our placement and routing predictions come within roughly one percentage point of results using 524,000 tokens.

  6. ZyphraAI score18

    Routing is also predictable across layers: the experts a token used in one layer tell us which it will likely need next. Token shuffling moves each token to the GPU holding those experts, inside a transfer that already runs after attention, so it adds no network traffic.

    Routing is also predictable across layers: the experts a token used in one layer tell us which it will likely need next. Token shuffling moves each token to the GPU holding those experts, inside a transfer that already runs after attention, so it adds no network traffic.

  7. ZyphraAI score37

    So we put experts that are picked together on the same GPU, and send each token to a GPU once no matter how many of its experts live there. We call this correlated expert placement. On 8 GPUs it removes up to 58% of the token copies sent.

    So we put experts that are picked together on the same GPU, and send each token to a GPU once no matter how many of its experts live there. We call this correlated expert placement. On 8 GPUs it removes up to 58% of the token copies sent.

  8. ZyphraAI score32

    MoE models route each token sparsely to just a few experts. As these models grow, the experts are spread across GPUs and nodes, and tokens have to be sent to their experts and back. In our runs that exchange took 13-24% of step time on one node and 45-60% across four nodes.

    MoE models route each token sparsely to just a few experts. As these models grow, the experts are spread across GPUs and nodes, and tokens have to be sent to their experts and back. In our runs that exchange took 13-24% of step time on one node and 45-60% across four nodes.

  9. ZyphraAI score23

    Those routing decisions have patterns. For instance, in one layer, just 64 of 8,128 possible expert pairs accounted for 42% of tokens. If experts were picked independently, those pairs would carry 1.6%. We show how to exploit this structure to improve where experts are placed.

    Those routing decisions have patterns. For instance, in one layer, just 64 of 8,128 possible expert pairs accounted for 42% of tokens. If experts were picked independently, those pairs would carry 1.6%. We show how to exploit this structure to improve where experts are placed.

  10. ZyphraAI score22

    In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime. At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.

    In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime. At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.

  11. OpenBMBAI score36

    ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%

    ReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.

  12. QbitAI (量子位)AI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    UniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  13. MarkTechPostAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    NVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

  14. Hacker News · AI (150+ points)AI score40

    OpenAI Withdraws Three Math Papers Over Sign Error in Weil Classes Proof

    OpenAI withdrew three math manuscripts, including "Algebraicity of Weil classes on split abelian eightfolds," after a sign error invalidated a stabilization-trace cancellation argument. The withdrawal also affects two papers that depended on that construction, and the withdrawn papers now carry notices linking to archived manuscripts. The update also revised 14 other manuscripts with proof repairs and added six formalizations.

  15. PandailyAI score46

    Galbot and Tsinghua's LATENT Wins IROS Award for Humanoid Tennis Forehand

    A Galbot, Tsinghua University and collaborators paper won IROS 2026's Best Entertainment and Amusement Paper Award for LATENT, a humanoid tennis-return method trained on imperfect amateur motion-capture clips. In simulation, the full forehand policy succeeded on 96.52 percent of returns, versus 71.85 percent for PULSE. On a real Unitree G1, the paper reports 90.90 percent forehand success across 20 consecutive rallies, with motion capture still used rather than the robot's own cameras.

  16. PandailyAI score38

    Huawei Presents Experimental XMFS Shared-Memory Filesystem at LPC 2026

    Huawei engineers presented XMFS, an experimental Linux kernel prototype filesystem, at the Linux Plumbers Conference in Prague on October 5. It aims to let applications reach cross-node shared memory on CXL 3.0 or Huawei unified bus servers through standard POSIX file calls. The code exists only on openEuler, not in the mainline Linux kernel.

  17. Anthropic ResearchAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    Johns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    AIWhy it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

Oct 7

Oct 7Wed
  1. Khazix (数字生命卡兹克)AI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    OpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    AIWhy it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. Elvis SaraviaAI score67

    Tool-using multimodal models refuse harmful requests less often, NVIDIA study finds

    A NVIDIA study accepted at NeurIPS 2026 reports that multimodal models refuse harmful requests less reliably when they call tools. Refusal failures rise by up to 68.7% relative and by 17.7% on average across the models tested, including Claude Opus 4.6 and 4.7 and Gemini Agentic Vision. The authors attribute this to tool outputs crowding out the original harmful intent and to attention shifting toward describing tool results. Re-inserting the original request and image before the final response restores part of the lost refusals.

  3. vLLMAI score46

    vLLM-Omni technical report unifies serving for omni-modality generation

    The vLLM team released a technical report on vLLM-Omni, a unified serving runtime for omni-modality generation spanning multi-stage autoregressive pipelines, iterative diffusion, and stateful sessions. Current LLM servers and diffusion stacks each cover only one of these patterns, pushing deployments to stitch disjoint runtimes together. vLLM-Omni offers a shared control plane in which an orchestrator advances requests across stages, specialized engines handle compute, and a connector carries payloads.

  4. Apple Machine Learning ResearchAI score42

    Apple's Normalizing Trajectory Models generate images in four steps with exact likelihood

    Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow trained with exact likelihood. The model matches or outperforms strong image generation baselines on text-to-image benchmarks in just four sampling steps while retaining exact likelihood over the generative trajectory.

  5. Waymo BlogAI score42

    Sober Drivers Still Face Nearly 4x Nighttime Fatal Crash Risk, Waymo Study Finds

    Waymo research found that even fully sober human drivers face nighttime fatal crash risk 3.1 to 3.9 times higher than daytime risk, pointing to systemic hazards beyond impairment. The study used an exposure reconstruction model across the 50 most populous U.S. urban areas, showing removing alcohol-involved drivers lowers the average urban fatal crash rate by 23%, from 1.42 to 1.10 per 100 million miles.

  6. Google ResearchAI score23

    Join Alex Bie at the @COLM_conf Google booth #107 today at 5:00 PM for a walkthrough of ContinuousBench, a standardized benchmark designed to measure knowledge transfer in differentially private (DP) synthesis. Don't miss the chance to explore if DP synthetic data truly preserve information, or just style? @GoogleDeepMind Read the paper: https://arxiv.org/abs/2606.01849

    Join Alex Bie at the @COLM_conf Google booth #107 today at 5:00 PM for a walkthrough of ContinuousBench, a standardized benchmark designed to measure knowledge transfer in differentially private (DP) synthesis. Don't miss the chance to explore if DP synthetic data truly preserve information, or just style? @GoogleDeepMind Read the paper: https://arxiv.org/abs/2606.01849

  7. Google ResearchAI score42

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

  8. Google ResearchAI score62

    Google Research finds AI boosts patent drafting but junior lawyers' gains vanish without it

    A Google Research field experiment with 133 patent lawyers found AI tool access raised drafting scores by 0.34 to 0.38 standard deviations over three months. When the tool was removed for a redlining task, only senior lawyers kept an advantage of 0.45 SD, while junior lawyers showed no discernible improvement. The authors argue that tools which boost current output must not stop junior professionals from building the judgment that senior experts rely on.

    AIWhy it matters: The field experiment separates AI's short-term productivity gains from skill retained after the tool is removed, which matters for training junior professionals.

  9. NVIDIA AIAI score26

    An AI agent makes a mistake early in a task, then keeps going in the wrong direction. Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps. Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd

    An AI agent makes a mistake early in a task, then keeps going in the wrong direction. Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps. Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd

  10. Google ResearchAI score10

    LLM agents learn by interacting with environments, but static setups limit their growth. Today at 2:00 PM, join Zifeng Wang at the #COLM2026 Google booth (#107) to learn about EnvHarness, a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and agent adaptability.

    LLM agents learn by interacting with environments, but static setups limit their growth. Today at 2:00 PM, join Zifeng Wang at the #COLM2026 Google booth (#107) to learn about EnvHarness, a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and agent adaptability.

  11. Amazon ScienceAI score10

    Amazon Scholar and @UTAustin professor @mattlease explores how to harness AI for scientific discovery while assessing potential risks, drawing on work from @UTGoodSystems and @CosmicAI_Inst. Catch his Expo Talk at @COLM_conf Thursday at 1pm PT. #COLM2026

    Amazon Scholar and @UTAustin professor @mattlease explores how to harness AI for scientific discovery while assessing potential risks, drawing on work from @UTGoodSystems and @CosmicAI_Inst. Catch his Expo Talk at @COLM_conf Thursday at 1pm PT. #COLM2026

  12. Epoch AIAI score26

    We’re planning to periodically rerun InnovationEval with new, uncontaminated papers. We hope this will provide early signs if AI approaches automating AI R&D end-to-end, rather than performing individual tasks under human direction. Read more at our website: https://epoch.ai/publications/innovationeval

    We’re planning to periodically rerun InnovationEval with new, uncontaminated papers. We hope this will provide early signs if AI approaches automating AI R&D end-to-end, rather than performing individual tasks under human direction. Read more at our website: https://epoch.ai/publications/innovationeval

  13. Epoch AIAI score37

    Unfortunately, newer models have seen the original innovation during their training, making the task significantly easier. However, even with this information, GPT-6 Astra and Claude Fable 5.1 struggled to reimplement it.

    Unfortunately, newer models have seen the original innovation during their training, making the task significantly easier. However, even with this information, GPT-6 Astra and Claude Fable 5.1 struggled to reimplement it.

  14. Epoch AIAI score24

    AI performance was underwhelming. Neither model achieved anything close to the human-authored reference. They reused existing methods from the literature and tuned hyperparameters, but struggled to create anything new.

    AI performance was underwhelming. Neither model achieved anything close to the human-authored reference. They reused existing methods from the literature and tuned hyperparameters, but struggled to create anything new.

  15. Epoch AIAI score22

    We instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these could be improved by a recent human-authored post-training innovation: on-policy self-distillation (SDPO).

    We instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these could be improved by a recent human-authored post-training innovation: on-policy self-distillation (SDPO).