Skip to contentSkip to stories

Updated

#Reasoning

Showing low-relevance items too. Hide low-relevance items

Oct 3

Oct 3Sat
  1. Sebastian RaschkaAI score38

    Raschka's Reasoning from Scratch covers RLVR and GRPO implementation

    AISebastian Raschka released round six of his Reasoning from Scratch series, introducing Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO) with an implementation. The video covers accuracy and format rewards, DeepSeek-R1 training, and GRPO versus PPO, then walks through a training loop and evaluates checkpoints on MATH-500.

    Video from @rasbt's post

Oct 2

Oct 2Fri
  1. Design ArenaAI score40

    GPT-6 Astra hedges far more than Claude Opus 5.5 in reasoning summaries

    AIDesign Arena analyzed 324 thinking summaries and found OpenAI's GPT-6 Astra uses hedging words like "maybe," "might," and "it seems" about 20 times as often as Anthropic's Claude Opus 5.5. Opus usually weighs a few options and commits early, in about 4 out of 5 summaries versus 1 in 4 for Astra, which the post says works more like a designer while Opus works more like a builder.

    Video from @DesignArena's post
  2. MIT News · AIAI score14

    MIT's Cathy Wu Uses Reinforcement Learning to Tackle Transportation Challenges

    AIMIT associate professor Cathy Wu is applying machine learning and reinforcement learning (RL) to design safer, more efficient transportation systems. Her team found RL can train effectively on about 10 percent of related problems, and a selection algorithm improved training efficiency by up to 30 times. Her recent work estimates eco-driving measures could cut vehicle emissions by 11 to 22 percent.

  3. AI at MetaAI score22

    Muse Spark helps prove finite-time blow-up in a laser-inspired wave model

    AIWith help from Muse Spark, researchers proved that a wave in a laser-inspired model must blow up in finite time under the conditions studied. The result comes from a tug-of-war between one effect squeezing the wave inward and another spreading it out. The paper is titled finite-time blow-up of radial negative-energy solutions for the mass-critical biharmonic nonlinear Schrödinger equation.

    Image from @AIatMeta's post
  4. AI at MetaAI score61

    Meta shares six math papers from mathematician-AI collaborations on open problems

    AIAI at Meta says mathematicians used Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the standard meta.ai chat interface to find solutions to open problems. The company is sharing six resulting papers, each marking which passages were drafted primarily by humans or AI, with mathematicians guiding the work and a second group reviewing it.

  5. Redwood Research BlogAI score34

    Capabilities research pushes the safety-usefulness frontier too, not just safety research

    AIThe post argues that counting all research as safety work because it widens the safety-usefulness Pareto frontier is misleading. Safety research typically creates new safety options without boosting usefulness, while capabilities research typically raises usefulness at safety's expense, so developers tend to choose less safe points.

  6. Harrison ChaseAI score53

    Google Research's Cogentic uses multi-agent proof search to produce verified results

    AIGoogle Research's Cogentic is a multi-agent harness running on Gemini that searches for proofs of open theoretical computer science problems without expert hints. It runs rounds where an orchestrator launches provers, two adversarial verifiers must both accept each draft, and shared disk documents store attempts and verified lemmas. The system produced new results on five open problems in online learning, auction theory, and mechanism design, each checked by domain experts.

  7. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  8. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

    Image from @huggingface's post
  9. MIT Technology Review · AIAI score62

    AlphaGo's move 37 shows why LLMs do not truly reason, an AlphaGo team member argues

    AIThore Graepel, a core member of the AlphaGo team, argues that current large language models do not truly reason, despite chain-of-thought gains in math and coding. He says they lack an explicit, inspectable epistemic state, keep knowledge and reasoning intertwined in their weights, and often produce post-hoc explanations. He proposes systems that maintain an auditable epistemic state and evaluate each step by how much it resolves uncertainty.

Oct 1

Oct 1Thu
  1. François CholletAI score62

    Chollet Argues Reasoning Models Differ from Base LLMs by Inductive Program Prediction

    AIFrançois Chollet argues the key difference between base LLMs and modern LRMs is a shift from transductive answer prediction to inductive prediction of the program or reasoning chain behind an answer. He says this enables test-time induction and substantial fluid intelligence in LRMs, which he claims base LLMs largely lack. He cites ARC 1 results: base LLMs remain around 10-15%, while LRMs of the same size or smaller saturated the benchmark in 2025.

  2. Alexander DoriaAI score54

    SYNTH paper proposes fully synthetic single-stage training for reasoning models

    AIThe SYNTH paper, titled It's All Training, presents a fully synthetic single-stage pipeline for training workable reasoning models with high data efficiency. The authors argue this approach does not separate training into pretraining, mid-training, or post-training stages. The image shows the paper's abstract, which describes a pipeline built from a 58,000-article Wikipedia-based synthetic corpus and models named Baguettotron-600M and Baguettotron-MoE.

    Image from @Dorialexander's post
  3. Amazon ScienceAI score34

    Amazon Science Explains Graph-Centric Agentic AI for Network Root Cause Analysis

    AIAmazon Science describes a graph-centric approach in which a network digital twin graph and cascaded graph algorithms, orchestrated by an agentic AI layer, identify root causes in complex network failures. The approach was demonstrated with NTT DOCOMO at the Mobile World Conference, achieving root cause analysis in minutes on commercial networks. The article traces how graphs evolved from topology models to active reasoning substrates for agents.

  4. Anthropic ResearchAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

Sep 30

Sep 30Wed
  1. Apple Machine Learning ResearchAI score36

    RLTL;DR: Self-Improvement Through Internalized Self-Generated Feedback

    AIApple researchers introduced RLTL;DR, a reinforcement learning method in which an agent writes its own one-line insight after each failed attempt and learns to map tasks to those insights. On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at 0% to 1% Pass@1, while RLTL;DR reached 14–31% with insights in context and 12–13% without them at evaluation. A compact variant, SFTL;DR, trained on just 4k task-insight tuples recovered nearly the full performance of RLTL;DR.