Skip to content
TodayOct 8Thu25 items
  1. 雷峰网 Leiphone46

    IROS 2026 papers show AI reintegrating with classical robotics rather than replacing it

    Of 1,933 IROS 2026 papers, Robot Learning/Embodied AI appears in about 809, while Navigation/Planning covers 564 and Perception/Vision 556. The article argues large models are being embedded into traditional planning, geometry, and control rather than replacing them. Vision-language-action models are shifting toward efficiency, 3D understanding, memory, and system integration.

  2. SiliconANGLE · AI60

    OpenAI publishes 722 AI-generated math papers, including Riemann hypothesis progress

    OpenAI has published 722 math papers generated by an unreleased AI model, posted to GitHub, spanning about 20 mathematical subfields. The model did not fully prove the Riemann hypothesis but proved the quasi-Riemann hypothesis, and it also produced theoretical computer science and partial differential equation results. Many papers include Lean files for computer verification, and OpenAI plans to release more of them.

  3. Midjourney34

    We're testing a new "thinking mode" for our image generation on our Alpha website (alpha dot midjourney dot com). We're finding it boosts prompt accuracy, typography and coherence. We'd love your help testing it on your images and telling us what you think. Thanks! <3

    We're testing a new "thinking mode" for our image generation on our Alpha website (alpha dot midjourney dot com). We're finding it boosts prompt accuracy, typography and coherence. We'd love your help testing it on your images and telling us what you think. Thanks! <3

  4. Midjourney Updates46

    Midjourney Tests Thinking Mode for Image Generation on Alpha Site

    Midjourney is testing a "Thinking Mode" on its Alpha website, where users can click "Rerun (Thinking)" in a job's lightbox to regenerate an image. The company says early tests show gains in prompt accuracy, typography, and coherence, and it is asking users to share feedback in its #ideas-and-features channel. It may later offer the mode broadly or as an option to add more thinking after a job.

  5. Artificial Analysis42

    Generating more output tokens doesn’t necessarily translate to a higher score. GPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.

    Generating more output tokens doesn’t necessarily translate to a higher score. GPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.

  6. Artificial Analysis34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    Artificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

  7. Ethan Mollick42

    Interesting to see, given the controversy over the OpenAI release of a series of proofs and what it means for the discipline of mathematics, that at least some of the OpenAI proofs seem to have kicked off extremely rapid iterative advances from a wide community of collaborators.

    Interesting to see, given the controversy over the OpenAI release of a series of proofs and what it means for the discipline of mathematics, that at least some of the OpenAI proofs seem to have kicked off extremely rapid iterative advances from a wide community of collaborators.

  8. Lewis Tunstall62

    Lewis Tunstall Shares a Physics Paper Proof Developed with OpenAI's Astra Model

    Lewis Tunstall quotes Kyle Cranmer's post about a paper by Nate Gunnarsson on a non-perturbative approach to chiral fermions in the Standard Model, extending Lüscher's abelian result. The paper's acknowledgments state that OpenAI's GPT-6 Astra model was essential, proposing refinement strategies, writing rewrites of the proof, and carrying out Lean verification.

  9. Stanford HAI22

    Stanford HAI leaders urge keeping people central to AI-driven research

    Stanford HAI associate directors Risa Wechsler and Russ Altman, speaking at a Stanford orientation, argued that AI agents can deepen scientific research but must be paired with interdisciplinary collaboration. They stressed rigorous, reproducible methods and clearly measured uncertainty, since convincing AI answers are not enough. They also said labs must weigh agent costs and preserve mentorship so that automation supports human participation in research.

  10. Google Research14

    Missed yesterday's demo on adaptive agent environments? Stop by the #COLM2026 Google booth #107 today at 11:00 AM PT to catch Zifeng Wang presenting EnvHarness — a flexible framework enabling co-evolution between LLM agents and their training environments.

    Missed yesterday's demo on adaptive agent environments? Stop by the #COLM2026 Google booth #107 today at 11:00 AM PT to catch Zifeng Wang presenting EnvHarness — a flexible framework enabling co-evolution between LLM agents and their training environments.

  11. 阮一峰 · 科技爱好者周刊42

    Weekly tech digest examines Jev decision model, which returns probabilities instead of text

    TypeSafe AI released Jev, a "decision model" that returns a floating-point probability rather than text, which can answer yes/no and multiple-choice questions and score content against criteria. The source cites two browser-extension examples: semantic Ctrl+F search and webpage quality scoring. Simon Willison's criticism is that Jev offers no explanation for its numbers.

  12. The Decoder46

    Ethereum researchers warn AI math advances could threaten crypto wallet signatures

    Ethereum researcher Justin Drake warned on X that AI-assisted math could, in the worst case, break the signature system used by crypto wallets within months, and urged a "bunker mode" in which users move funds to addresses that have never signed a transaction. Vitalik Buterin agreed but cautioned against moving too fast, saying he has lost more money to botched migrations than to hacks. No one has yet broken the current ECDSA signature scheme in practice.

  13. vLLM62

    vLLM v0.31.0 adds DeepSeek-V4.1-Flash support and new serving features

    vLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors. Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.

  14. 量子位 QbitAI44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    UniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  15. MarkTechPost45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    NVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

  16. Gergely Orosz48

    You can either hold crypto and keep being stressed out if a math breakthrough would drain your wallet; or someone stealing your keys would drain your wallet; or someone kidnapping you and forcing you to hand over your keys would drain your wallet Or you can just not hold crypto

    You can either hold crypto and keep being stressed out if a math breakthrough would drain your wallet; or someone stealing your keys would drain your wallet; or someone kidnapping you and forcing you to hand over your keys would drain your wallet Or you can just not hold crypto

Oct 7Wed
  1. 数字生命卡兹克88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    OpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. Andrew Curran52

    AI Labs Reportedly Test Internal Models Against Cryptographic Protocols

    Scott Aaronson reports, based on his sources, that some AI companies have begun discreetly investigating whether their latest internal models can break important cryptographic protocols and primitives. He notes that cryptography is conspicuously absent from OpenAI's list of 376 papers, and the quoted post adds that the US government has censored academic quantum cryptanalysis results.

  3. François Chollet44

    Chollet: Programming and math training don't boost general intelligence

    François Chollet compares AI progress to human learning, noting that 1980s research found programming training improves coding but does not transfer to general reasoning. He argues general intelligence is a fundamental brain property rather than a trainable skill, since domain practice improves only that domain. The post is framed as background for his question whether AI's jagged frontier, driven by math and code via RLVR, reflects general capability or continued human-data bottlenecks.

  4. Epoch AI67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    Epoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  5. Andrew Curran38

    Scott Aaronson on The Mathocalypse. 'The UGC proof invents a completely new bizarre code with a noise test. It's some crazy recursive construction. It's not the long code, not the short code - some alien craziness' I hear a lot of this today. Get used to alien craziness.

    Scott Aaronson on The Mathocalypse. 'The UGC proof invents a completely new bizarre code with a noise test. It's some crazy recursive construction. It's not the long code, not the short code - some alien craziness' I hear a lot of this today. Get used to alien craziness.

  6. Ethan Mollick60

    Mathematicians react to hundreds of AI-generated proofs released by OpenAI

    Ethan Mollick shares early first-hand accounts from mathematicians grappling with hundreds of AI proofs released by OpenAI. He highlights problems solved in ways no human has yet understood, raising questions about what it means to know something. The linked Scott Aaronson post quotes a researcher, Dana, describing the proofs as unclear and hard to read without AI help, with some possibly verified by a Lean certificate.