Skip to content

#Reasoning

Oct 8

TodayOct 8Thu9 items
  1. Artificial Analysis42

    Generating more output tokens doesn’t necessarily translate to a higher score. GPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.

    Generating more output tokens doesn’t necessarily translate to a higher score. GPT-6 Astra (max) scores 8.6% on ~81k output tokens per task, under half the ~180k of Grok 4.7 (xhigh). Three Claude models generated the most output tokens (~202k to ~562k per task) and score 2.8% to 6.4%.

  2. Ethan Mollick42

    Interesting to see, given the controversy over the OpenAI release of a series of proofs and what it means for the discipline of mathematics, that at least some of the OpenAI proofs seem to have kicked off extremely rapid iterative advances from a wide community of collaborators.

    Interesting to see, given the controversy over the OpenAI release of a series of proofs and what it means for the discipline of mathematics, that at least some of the OpenAI proofs seem to have kicked off extremely rapid iterative advances from a wide community of collaborators.

  3. Lewis Tunstall62

    Lewis Tunstall Shares a Physics Paper Proof Developed with OpenAI's Astra Model

    Lewis Tunstall quotes Kyle Cranmer's post about a paper by Nate Gunnarsson on a non-perturbative approach to chiral fermions in the Standard Model, extending Lüscher's abelian result. The paper's acknowledgments state that OpenAI's GPT-6 Astra model was essential, proposing refinement strategies, writing rewrites of the proof, and carrying out Lean verification.

  4. Stanford HAI22

    Stanford HAI leaders urge keeping people central to AI-driven research

    Stanford HAI associate directors Risa Wechsler and Russ Altman, speaking at a Stanford orientation, argued that AI agents can deepen scientific research but must be paired with interdisciplinary collaboration. They stressed rigorous, reproducible methods and clearly measured uncertainty, since convincing AI answers are not enough. They also said labs must weigh agent costs and preserve mentorship so that automation supports human participation in research.

  5. 阮一峰 · 科技爱好者周刊42

    Weekly tech digest examines Jev decision model, which returns probabilities instead of text

    TypeSafe AI released Jev, a "decision model" that returns a floating-point probability rather than text, which can answer yes/no and multiple-choice questions and score content against criteria. The source cites two browser-extension examples: semantic Ctrl+F search and webpage quality scoring. Simon Willison's criticism is that Jev offers no explanation for its numbers.

  6. The Decoder46

    Ethereum researchers warn AI math advances could threaten crypto wallet signatures

    Ethereum researcher Justin Drake warned on X that AI-assisted math could, in the worst case, break the signature system used by crypto wallets within months, and urged a "bunker mode" in which users move funds to addresses that have never signed a transaction. Vitalik Buterin agreed but cautioned against moving too fast, saying he has lost more money to botched migrations than to hacks. No one has yet broken the current ECDSA signature scheme in practice.

  7. Gergely Orosz48

    You can either hold crypto and keep being stressed out if a math breakthrough would drain your wallet; or someone stealing your keys would drain your wallet; or someone kidnapping you and forcing you to hand over your keys would drain your wallet Or you can just not hold crypto

    You can either hold crypto and keep being stressed out if a math breakthrough would drain your wallet; or someone stealing your keys would drain your wallet; or someone kidnapping you and forcing you to hand over your keys would drain your wallet Or you can just not hold crypto

Oct 7

Oct 7Wed
  1. Andrew Curran52

    AI Labs Reportedly Test Internal Models Against Cryptographic Protocols

    Scott Aaronson reports, based on his sources, that some AI companies have begun discreetly investigating whether their latest internal models can break important cryptographic protocols and primitives. He notes that cryptography is conspicuously absent from OpenAI's list of 376 papers, and the quoted post adds that the US government has censored academic quantum cryptanalysis results.

  2. François Chollet44

    Chollet: Programming and math training don't boost general intelligence

    François Chollet compares AI progress to human learning, noting that 1980s research found programming training improves coding but does not transfer to general reasoning. He argues general intelligence is a fundamental brain property rather than a trainable skill, since domain practice improves only that domain. The post is framed as background for his question whether AI's jagged frontier, driven by math and code via RLVR, reflects general capability or continued human-data bottlenecks.

  3. Andrew Curran38

    Scott Aaronson on The Mathocalypse. 'The UGC proof invents a completely new bizarre code with a noise test. It's some crazy recursive construction. It's not the long code, not the short code - some alien craziness' I hear a lot of this today. Get used to alien craziness.

    Scott Aaronson on The Mathocalypse. 'The UGC proof invents a completely new bizarre code with a noise test. It's some crazy recursive construction. It's not the long code, not the short code - some alien craziness' I hear a lot of this today. Get used to alien craziness.

  4. Ethan Mollick60

    Mathematicians react to hundreds of AI-generated proofs released by OpenAI

    Ethan Mollick shares early first-hand accounts from mathematicians grappling with hundreds of AI proofs released by OpenAI. He highlights problems solved in ways no human has yet understood, raising questions about what it means to know something. The linked Scott Aaronson post quotes a researcher, Dana, describing the proofs as unclear and hard to read without AI help, with some possibly verified by a Lean certificate.

  5. Marcus on AI62

    Marcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality

    Gary Marcus argues that OpenAI's math announcement omits the procedure, the model architecture, and the failure rate, so its generalizability cannot be assessed. He says it could be a step toward AGI or a Lean-based verification trick in a verifiable domain, and the initial report cannot distinguish the two. The post includes a quoted Terence Tao post that shares a satirical press release about a fictional film-endings repository.

  6. Mark Chen46

    The Navier-Stokes moment was always much more about the figure below than about the Navier-Stokes problem itself. The run represents a decade of mathematical progress in a week. Can't wait to point these tools at life sciences, building our next models, and alignment!

    The Navier-Stokes moment was always much more about the figure below than about the Navier-Stokes problem itself. The run represents a decade of mathematical progress in a week. Can't wait to point these tools at life sciences, building our next models, and alignment!

  7. Santiago18

    Prompt engineering is not dead. It was never alive in the first place. The better these models become, the less we'll need to resort to phrasing tricks or weird instructions. The endgame is models that understand the way we communicate. "Prompt engineering" will simply become "communication skills".

    Prompt engineering is not dead. It was never alive in the first place. The better these models become, the less we'll need to resort to phrasing tricks or weird instructions. The endgame is models that understand the way we communicate. "Prompt engineering" will simply become "communication skills".

  8. Exponential View72

    OpenAI's 722 machine-generated math results may split mathematics into two layers

    OpenAI released 722 mathematical manuscripts in 372 families, produced by an unreleased frontier model, with the average result taking the equivalent of three hours of ChatGPT Pro thinking. The author notes many results are verified in Lean but not all, and suggests mathematics could divide into vast machine-verified work and a compressed human 'effective theory' that people can actually understand.

Oct 6

Oct 6Tue
  1. Lewis Tunstall25

    This is the most important plot from the Beam release IMO. The Chinese models are great, but horribly token inefficient (try running an eval with max reasoning to feel the pain). I'm looking forward to a future where open models start competing on this axis!

    This is the most important plot from the Beam release IMO. The Chinese models are great, but horribly token inefficient (try running an eval with max reasoning to feel the pain). I'm looking forward to a future where open models start competing on this axis!

  2. will depue35

    i’m surprised all of the ai lab math results have been real so far: we haven’t found any with profound errors or a real bug yet, when you should expect so from humans. i assume at least a couple of these results today shouldnt survive scrutiny?

    i’m surprised all of the ai lab math results have been real so far: we haven’t found any with profound errors or a real bug yet, when you should expect so from humans. i assume at least a couple of these results today shouldnt survive scrutiny?

  3. Kevin Weil40

    AI x mathematics ftw. What an incredible release today from OpenAI 🤯 There's a lot of work to do to get models of the same caliber in the physical sciences, but just think of the possibilities as we get there. And we will get there.

    AI x mathematics ftw. What an incredible release today from OpenAI 🤯 There's a lot of work to do to get models of the same caliber in the physical sciences, but just think of the possibilities as we get there. And we will get there.

  4. will depue62

    Will DePue's list claims AI resolved dozens of famous open math problems

    A post by Will DePue titled "Fable 5.1's list" presents 100 mathematical results and says 59% were released today, 87% AI and 13% human. The list includes items attributed to OpenAI, Anthropic, Google DeepMind and human mathematicians, each marked by a colored indicator, and it describes many entries as formalized in Lean or as openai/math family numbers. The post supplies no independent verification of these claims.

  5. Microsoft Research36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    Microsoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

Oct 5

Oct 5Mon
  1. Mike Knoop62

    Dust pretrains transformers with zeroth-order optimization, approaching backprop results

    Dust is a zeroth-order method that pretrains transformers and sometimes matches or exceeds backprop given large compute. The authors report it is about 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. The post also cites the gradient-alignment result up to 1B tokens and the virtual population idea for scaling.