Skip to content

#Paper/Research

Oct 8

TodayOct 8Thu4 items
  1. Leiphone (雷峰网)AI score46

    IROS 2026 Best Paper goes to LT-Mem robot long-term memory study

    At IROS 2026 in Pittsburgh, the Best Paper Award went to Yumin Lee, Hyoseok Ju and Giseop Kim for LT-Mem, a volatility-aware spatio-temporal memory system for lifelong robot scene understanding. The Best Student Paper Award went to Pei-An Hsieh and colleagues for flatness-preserving residual learning enabling real-time tight quadrotor formation flight. Other honors included a humanoid tennis-skills paper and SteadyTray, a humanoid tray-transport study.

  2. TechCrunch · AIAI score65

    OpenAI's math solutions fall short of the field's standards, mathematicians say

    OpenAI released hundreds of claimed solutions to hard math problems but did not fully meet guidelines from the Advisory Group on Mathematics and Artificial Intelligence. Only 10 of 719 manuscripts included chain-of-thought releases, and just 42% of proofs were formalized. A Cambridge and King's College paper found discrepancies between a natural language proof and its Lean code for a Navier-Stokes-derived problem.

  3. Hacker News · AI (150+ points)AI score40

    OpenAI Withdraws Three Math Papers Over Sign Error in Weil Classes Proof

    OpenAI withdrew three math manuscripts, including "Algebraicity of Weil classes on split abelian eightfolds," after a sign error invalidated a stabilization-trace cancellation argument. The withdrawal also affects two papers that depended on that construction, and the withdrawn papers now carry notices linking to archived manuscripts. The update also revised 14 other manuscripts with proof repairs and added six formalizations.

Oct 7

Oct 7Wed
  1. Latent SpaceAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    OpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

Oct 6

Oct 6Tue
  1. Gizmodo · AIAI score62

    OpenAI Releases 377 Math Results on GitHub Amid Expert Concerns

    OpenAI released 377 new math results on GitHub, including one paper claiming a proof of the full Birch-Swinnerton-Dyer leading term formula for elliptic curves over the rationals under specific conditions. The results come from the same unreleased internal model that produced its earlier Navier-Stokes result, which conflicts with a September 29 recommendation from the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) to stop testing advanced math problems on proprietary models.

  2. Amazon ScienceAI score4

    Day 1 at @COLM_conf. We're here all week with live research talks, meet-the-scientist sessions, and conversations on what's next for language models. Stop by the Amazon booth. #COLM2026 More info: https://amzn.to/4jsMYKM

    Day 1 at @COLM_conf. We're here all week with live research talks, meet-the-scientist sessions, and conversations on what's next for language models. Stop by the Amazon booth. #COLM2026 More info: https://amzn.to/4jsMYKM

  3. Amazon ScienceAI score8

    We're at @COLM_conf this week! Stop by the Amazon booth for live talks on super weights in LLMs, multi-agent orchestration, responsible AI, and more, plus meet-the-scientist sessions all week. #COLM2026 Full schedule: https://amzn.to/4jsMYKM

    We're at @COLM_conf this week! Stop by the Amazon booth for live talks on super weights in LLMs, multi-agent orchestration, responsible AI, and more, plus meet-the-scientist sessions all week. #COLM2026 Full schedule: https://amzn.to/4jsMYKM

  4. ARC PrizeAI score14

    Announcing ARC Prize Research Summit 2026 Research Speaker @alexisfox is a researcher at Duke University exploring how AI can reason over longer horizons, and lead author of PRO-LONG, which uses programmatic memory to support long-horizon reasoning.

    Announcing ARC Prize Research Summit 2026 Research Speaker @alexisfox is a researcher at Duke University exploring how AI can reason over longer horizons, and lead author of PRO-LONG, which uses programmatic memory to support long-horizon reasoning.

Oct 5

Oct 5Mon
  1. IEEE Spectrum · AIAI score58

    Mathematicians Debate OpenAI's Navier-Stokes Claim and AI's Impact on the Field

    Mathematicians at the Heidelberg Laureate Forum discussed AI companies, including OpenAI, Anthropic, and Google, solving longstanding math problems. OpenAI announced it had solved the Navier-Stokes existence and smoothness problem, a claim the article says is still awaiting verification, and Harris criticized the company's conduct toward a mathematician. Researchers also warn that AI solutions may lack understandable methods and are changing how academics work.

Oct 2

Oct 2Fri

Oct 1

Oct 1Thu
  1. Sundar PichaiAI score60

    Google DeepMind's SynthID Bio watermarks AI-designed protein sequences

    Google DeepMind announced SynthID Bio, a family of watermarking methods for AI-generated biological designs. According to the quoted post, the team can embed an imperceptible signature directly into protein sequences without affecting their biological function. Sundar Pichai called it a big step forward for scientific integrity and biosecurity.

Sep 30

Sep 30Wed
  1. Amazon ScienceAI score7

    Amazon researchers are heading to @COLM_conf in San Francisco next week with accepted publications spanning agentic AI, LLM post-training, multi-agent systems, and automated reasoning. See our research and booth schedule: https://amzn.to/4dcLxw9 #COLM2026

    Amazon researchers are heading to @COLM_conf in San Francisco next week with accepted publications spanning agentic AI, LLM post-training, multi-agent systems, and automated reasoning. See our research and booth schedule: https://amzn.to/4dcLxw9 #COLM2026

  2. Google · Innovation & AIAI score46

    Google AI Flu Model Ranks First in CDC FluSight Hospitalization Forecasts

    A flu forecasting model built with Google AI ranked first among 39 eligible models in the CDC's FluSight 2025-26 season evaluation for predicting U.S. flu-related hospital admissions. The model was developed using Empirical Research Assistance (ERA), an AI tool that generates optimization algorithms, and ERA's underlying technology is now available to trusted testers.

Sep 25

Sep 25Fri
  1. Anthropic ResearchAI score67

    Claude computes a nine-loop physics amplitude that experts had not reached

    Anthropic researchers used Claude Science to compute the nine-loop six-particle amplitude in planar N=4 super Yang-Mills, a toy-model result that physicist Lance Dixon checked. The work reportedly cost roughly one or two thousand dollars, with about $100 of compute for the bootstrap calculation, and a similar result was reached by Song He's group.

    AIWhy it matters: The guest post shows a frontier physics calculation done with modest compute, which helps readers gauge what current AI can handle in research and what it still cannot.

Sep 23

Sep 23Wed
  1. Felix RiesebergAI score38

    You may have heard that my colleagues over in our experimental biology team have discovered their first major breakthrough, which is really exciting - they discovered a new enzyme system, array-associated reverse transcriptases (ART), that have never been reported by human scientists. I really love reading some of Claude's takes while it was assisting with this research though. Good Claude!

    You may have heard that my colleagues over in our experimental biology team have discovered their first major breakthrough, which is really exciting - they discovered a new enzyme system, array-associated reverse transcriptases (ART), that have never been reported by human scientists. I really love reading some of Claude's takes while it was assisting with this research though. Good Claude!

  2. Anthropic · YouTubeAI score65

    Anthropic launches a molecular biology lab where Claude hunts for unusual proteins

    Anthropic is introducing a molecular biology research group and lab to test whether Claude can help scientists find unusual proteins. Claude combs through large DNA datasets, flags uncharacterized proteins, and passes its most promising ideas to scientists, who test them at the bench. In one early program, Claude discovered a novel enzyme system with CRISPR-like repeats.

    AIWhy it matters: The source shows Claude being used in a wet-lab workflow, from scanning DNA datasets to flagging proteins for scientists to test at the bench.

Sep 15

Sep 15Tue

Sep 9

Sep 9Wed
  1. Cognition Blog (Devin, Windsurf)AI score82

    Cognition's Devin factors RSA-260 using a GPU lattice siever

    Cognition's Devin agent, directed by Eric Lu, factored the 260-digit RSA-260 number using a new GPU implementation of the general number field sieve built on CADO-NFS. The author estimates the run cost about 13.5 GPU-years, roughly $400k at market prices, and projects RSA-1024 factoring at around $30M, while RSA-2048 is not meaningfully affected.

    AIWhy it matters: The source gives a full cost breakdown and scaling estimates for RSA factoring on GPUs, showing how far the cost of breaking RSA-1024 has fallen.

Aug 17

Aug 17Mon
  1. Microsoft ResearchAI score12

    Microsoft is proud to be a part of SIGCOMM 2026. Check out our sessions, covering network traffic engineering, new load balancing strategy, adaptive photonic switching schedules, AI network simulation, and more. https://msft.it/6018aKgfp

    Microsoft is proud to be a part of SIGCOMM 2026. Check out our sessions, covering network traffic engineering, new load balancing strategy, adaptive photonic switching schedules, AI network simulation, and more. https://msft.it/6018aKgfp

Jul 15

Jul 15Wed

Jun 16

Jun 16Tue
  1. BAAIAI score38

    BAAI unveils WuJie physical-world AI architecture in 2026 report

    BAAI President Wang Zhongyuan announced a shift in AI from token prediction to physical state prediction in the institute's 2026 annual research report. The report unveiled the full-stack WuJie architecture spanning foundation models, autonomous agents, and hardware-software infrastructure, and noted that BAAI has open-sourced over 200 models with global downloads exceeding 1 billion.

May 19

May 19Tue

May 6

May 6Wed
  1. OpenAI Alignment Research BlogAI score62

    OpenAI finds accidental chain-of-thought grading in several RL runs but no clear monitorability loss

    OpenAI reports that its automated system found accidental chain-of-thought grading in RL runs for several released models, including GPT-5.4 Thinking and GPT-5.4 mini. Its analysis found no clear reduction in CoT monitorability, though the company says subtler effects cannot be ruled out. OpenAI says it still avoids grading CoTs during RL and has fixed the affected reward pathways.

    AIWhy it matters: The post shows how accidental chain-of-thought grading was detected and tested, giving a concrete method for checking monitorability risks in RL training.

Feb 25

Feb 25Wed