Skip to content

#Reasoning

Oct 6

Oct 6Tue
  1. ARC PrizeAI score38

    Grok 4.7 from @SpaceXAI on ARC-AGI (Verified): - ARC-AGI-3: 1.8%, $2.7k (standard harness), 10.0%, $4.8k (provider adapter harness) - ARC-AGI-2: 61.4%, $2.01/task - ARC-AGI-1: 90.2%, $0.64/task Grok 4.7 scores higher than Grok 4.6 on ARC-AGI-1, but lower on ARC-AGI-2 and 3.

    Grok 4.7 from @SpaceXAI on ARC-AGI (Verified): - ARC-AGI-3: 1.8%, $2.7k (standard harness), 10.0%, $4.8k (provider adapter harness) - ARC-AGI-2: 61.4%, $2.01/task - ARC-AGI-1: 90.2%, $0.64/task Grok 4.7 scores higher than Grok 4.6 on ARC-AGI-1, but lower on ARC-AGI-2 and 3.

  2. Microsoft ResearchAI score36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    Microsoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

  3. ARC PrizeAI score25

    ARC Prize finds DeepSeek V4.1 Flash high reasoning gains no clear edge

    On ARC-AGI-1, DeepSeek V4.1 Flash scored 88.5% at high reasoning versus 90.5% at low, with high using 35% more output tokens without consistently better answers. On ARC-AGI-2, the reported per-task cost of max reasoning ($0.129) appears slightly lower than high ($0.133), but after excluding incomplete tasks caused by API issues, max is about 4.5% more expensive per task.

  4. ARC PrizeAI score46

    DeepSeek V4.1 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 72.9%, $0.13/task - ARC-AGI-1: 94.5%, $0.07/task DeepSeek V4.1 Flash beats V4 Flash's best scores by 5.5 points on ARC-AGI-1 and 11.5 on ARC-AGI-2, but costs about 250% more per task.

    DeepSeek V4.1 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 72.9%, $0.13/task - ARC-AGI-1: 94.5%, $0.07/task DeepSeek V4.1 Flash beats V4 Flash's best scores by 5.5 points on ARC-AGI-1 and 11.5 on ARC-AGI-2, but costs about 250% more per task.

  5. Sophia YangAI score26

    Reinforcement learning at scale: - Autoscaling actor fleet → tens of thousands of rollouts in parallel, async training - Built for long trajectories: millions of tokens per rollout, multiple compactions, low staleness - New methods at both stages cut off-policy drift → stable long-horizon RL - 3k GPUs → ~33B tokens/day, ~16B trainable after filtering + masking - Rewards climb across representative envs as the policy learns harder tasks ↓

    Reinforcement learning at scale: - Autoscaling actor fleet → tens of thousands of rollouts in parallel, async training - Built for long trajectories: millions of tokens per rollout, multiple compactions, low staleness - New methods at both stages cut off-policy drift → stable long-horizon RL - 3k GPUs → ~33B tokens/day, ~16B trainable after filtering + masking - Rewards climb across representative envs as the policy learns harder tasks ↓

  6. Latent SpaceAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    Reflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

Oct 5

Oct 5Mon
  1. Apple Machine Learning ResearchAI score23

    RISED uses rubrics to guide multi-environment LLM agent training and data selection

    Apple researchers introduce RISED, a framework that uses rubrics to guide data selection and policy supervision when training one LLM agent across multiple interactive environments. An LLM judge tags rollouts with a shared rubric vocabulary, positive rubrics provide privileged context for an on-policy self-distillation teacher, and negative rubrics steer generation away from recurring failures. The authors report that RISED achieves the highest mean pass rate across environments and ranks first or second in each environment, across model backbones.

  2. Mike KnoopAI score62

    Dust pretrains transformers with zeroth-order optimization, approaching backprop results

    Dust is a zeroth-order method that pretrains transformers and sometimes matches or exceeds backprop given large compute. The authors report it is about 1,000 to 10,000x more compute efficient than EGGROLL, the state-of-the-art ES method, for training transformers. The post also cites the gradient-alignment result up to 1B tokens and the virtual population idea for scaling.

  3. Sophia YangAI score62

    Reflection AI's Beam open model has 501B total parameters and 23B active

    Sophia Yang congratulated Reflection AI on Beam, a 501B-parameter open model with 23B active per token. She attributes its efficiency to an RL length penalty that discourages unnecessary tokens and a sparse MoE architecture. Reflection says full weights will be released this month, and the quoted post reports training over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks.

  4. Clément DelangueAI score72

    Reflection AI announces Beam, a 501B-parameter agentic open model

    Reflection AI introduced Beam, an agentic open model with 501B total parameters and 23B active parameters, trained end-to-end from scratch. The quoted announcement says it targets frontier reasoning efficiency and coding and agentic tasks, with full weights due this month. Clément Delangue, Hugging Face's CEO, reposted it with a welcome to the Reflection organization on Hugging Face.

    AIWhy it matters: The quoted announcement names Beam's parameter scale, active-parameter count, and coding and agentic focus, which helps readers gauge where it fits among open models.

  5. GoodfireAI score10

    Tue 11am (Franciscan A) Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior Tue 4:30pm (Imperial Ballroom) Do SAEs Capture Concept Manifolds? Wed 4:30pm (Franciscan B) Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts

    Tue 11am (Franciscan A) Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior Tue 4:30pm (Imperial Ballroom) Do SAEs Capture Concept Manifolds? Wed 4:30pm (Franciscan B) Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts

  6. Import AIAI score47

    Import AI 475 Covers Swarm Scaling, Google DeepMind's SynthID Bio, and AI Science Labs

    Toby Ord argues that AI agent swarms trade extra tokens for faster completion, needing about twice the total tokens of a single agent for the same performance with four agents, but in half the wall-clock time. He notes swarm scaling shows diminishing returns, with 10x agents yielding roughly 3x to 5x the performance of 10x tokens on one agent. A CSAIP poll found 61% of Americans think voluntary AI industry commitments are "not enough."

  7. Ethan MollickAI score17

    I've been complaining about some OpenAI choices tonight, so, on the plus side, I will say there is still no equivalent to GPT-6 Pro (just as there wasn't with previous Pro models). It both does really hard tasks in one shot & does a surprisingly good job communicating the results

    I've been complaining about some OpenAI choices tonight, so, on the plus side, I will say there is still no equivalent to GPT-6 Pro (just as there wasn't with previous Pro models). It both does really hard tasks in one shot & does a surprisingly good job communicating the results

Oct 3

Oct 3Sat
  1. François CholletAI score22

    Chollet: Computation alone doesn't make AI models conscious

    François Chollet argues that the claim AI models are likely conscious because they are computation is as flawed as saying a rock is likely alive because it is made of atoms. He says static input-output programs lack properties associated with consciousness, such as information integration, interoception, temporal binding, and embodiment. He adds that humanity has not created a conscious program and sees no signs of being close, so any future case should rest on evidence and consciousness science.

  2. IndexTeam (Bilibili) · new models on Hugging FaceAI score29

    Index-Homura-2B-FP4 released as NVFP4 quantized translation model

    IndexTeam released Index-Homura-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Homura-2B multilingual translation model, which supports 150 languages. The quantized checkpoint shows a 5.73% perplexity increase over the BF16 original (3.5011 to 3.7017) on a fixed corpus, and its zh-en and en-zh outputs are semantically equivalent under greedy decoding. Full FP4 acceleration requires an NVIDIA Blackwell GPU, while the source recommends the FP8 build for Hopper and Ampere hardware.

  3. CohereAI score22

    For fine-tuning SFT, we focused on extending machine translation capabilities by prepping data focused on post-editing, error detection, and terminology. We also went through multiple steps of reinforcement learning and DPO targeted towards errors the model exhibited. (2/6)

    For fine-tuning SFT, we focused on extending machine translation capabilities by prepping data focused on post-editing, error detection, and terminology. We also went through multiple steps of reinforcement learning and DPO targeted towards errors the model exhibited. (2/6)

  4. Sebastian RaschkaAI score38

    Raschka's Reasoning from Scratch covers RLVR and GRPO implementation

    Sebastian Raschka released round six of his Reasoning from Scratch series, introducing Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO) with an implementation. The video covers accuracy and format rewards, DeepSeek-R1 training, and GRPO versus PPO, then walks through a training loop and evaluates checkpoints on MATH-500.

Oct 2

Oct 2Fri
  1. MIT News · AIAI score14

    MIT's Cathy Wu Uses Reinforcement Learning to Tackle Transportation Challenges

    MIT associate professor Cathy Wu is applying machine learning and reinforcement learning (RL) to design safer, more efficient transportation systems. Her team found RL can train effectively on about 10 percent of related problems, and a selection algorithm improved training efficiency by up to 30 times. Her recent work estimates eco-driving measures could cut vehicle emissions by 11 to 22 percent.

  2. AI at MetaAI score21

    6️⃣ Non-Associative Algebra: The team worked with Muse Spark to find an exception to a proposed rule about mathematical structures inspired by biology. They went further by developing an alternative characterization, which the researchers checked and refined. Read the paper: https://ai.meta.com/research/publications/on-solvable-evolution-algebras-and-a-conjecture-by-garcia-martinez-and-perez-rodriguez/

    6️⃣ Non-Associative Algebra: The team worked with Muse Spark to find an exception to a proposed rule about mathematical structures inspired by biology. They went further by developing an alternative characterization, which the researchers checked and refined. Read the paper: https://ai.meta.com/research/publications/on-solvable-evolution-algebras-and-a-conjecture-by-garcia-martinez-and-perez-rodriguez/

  3. AI at MetaAI score42

    5️⃣ Arithmetic Physics: Researchers working with Muse Spark connected an idea from number theory with a calculation in string theory. They proved the link works in more cases than previously known, building on ideas from the 1980s. Read the paper: https://ai.meta.com/research/publications/string-two-point-function-height-function-on-a-curve/

    5️⃣ Arithmetic Physics: Researchers working with Muse Spark connected an idea from number theory with a calculation in string theory. They proved the link works in more cases than previously known, building on ideas from the 1980s. Read the paper: https://ai.meta.com/research/publications/string-two-point-function-height-function-on-a-curve/

  4. AI at MetaAI score22

    4️⃣ Optimization: When can you replace a hard math problem with a simpler one without losing anything? Researchers worked with Muse Spark to prove a clear rule for when a particular simplification captures the original exactly, and when it leaves a gap. Read the paper: https://ai.meta.com/research/publications/tightness-of-the-cycle-based-relaxation-for-completed-length-three-alpha-cycles/

    4️⃣ Optimization: When can you replace a hard math problem with a simpler one without losing anything? Researchers worked with Muse Spark to prove a clear rule for when a particular simplification captures the original exactly, and when it leaves a gap. Read the paper: https://ai.meta.com/research/publications/tightness-of-the-cycle-based-relaxation-for-completed-length-three-alpha-cycles/

  5. AI at MetaAI score18

    3️⃣ Group Theory: Researchers disproved a proposed rule about mathematical structures that describe symmetry by finding one counterexample. Muse Spark generated the search code that found it, and the team verified the result and completed the proof. Read the paper: https://ai.meta.com/research/publications/semiabelian-groups-need-not-be-monomial/

    3️⃣ Group Theory: Researchers disproved a proposed rule about mathematical structures that describe symmetry by finding one counterexample. Muse Spark generated the search code that found it, and the team verified the result and completed the proof. Read the paper: https://ai.meta.com/research/publications/semiabelian-groups-need-not-be-monomial/

  6. AI at MetaAI score22

    2️⃣ Differential Equations: Imagine a tug-of-war between one effect squeezing a wave inward and another spreading it out. Can the wave keep concentrating forever? With help from Muse Spark, researchers proved that, under the conditions studied, a wave in a laser-inspired model must "blow up" in finite time. Read the paper: https://ai.meta.com/research/publications/finite-time-blow-up-of-radial-negative-energy-solutions-for-the-mass-critical-biharmonic-nonlinear-schrodinger-equation/

    2️⃣ Differential Equations: Imagine a tug-of-war between one effect squeezing a wave inward and another spreading it out. Can the wave keep concentrating forever? With help from Muse Spark, researchers proved that, under the conditions studied, a wave in a laser-inspired model must "blow up" in finite time. Read the paper: https://ai.meta.com/research/publications/finite-time-blow-up-of-radial-negative-energy-solutions-for-the-mass-critical-biharmonic-nonlinear-schrodinger-equation/

  7. AI at MetaAI score40

    1️⃣ Probability: Mathematicians worked with Muse Spark to answer a question about fitting random points onto the surface of a stretched sphere. For the setting studied, they proved a sharp cutoff between when an exact fit is likely and when it is unlikely. Read the paper: https://ai.meta.com/research/publications/the-strict-threshold-for-gaussian-ellipsoid-fitting/

    1️⃣ Probability: Mathematicians worked with Muse Spark to answer a question about fitting random points onto the surface of a stretched sphere. For the setting studied, they proved a sharp cutoff between when an exact fit is likely and when it is unlikely. Read the paper: https://ai.meta.com/research/publications/the-strict-threshold-for-gaussian-ellipsoid-fitting/

  8. AI at MetaAI score61

    Meta shares six math papers from mathematician-AI collaborations on open problems

    AI at Meta says mathematicians used Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the standard meta.ai chat interface to find solutions to open problems. The company is sharing six resulting papers, each marking which passages were drafted primarily by humans or AI, with mathematicians guiding the work and a second group reviewing it.

  9. Redwood Research BlogAI score34

    Capabilities research pushes the safety-usefulness frontier too, not just safety research

    The post argues that counting all research as safety work because it widens the safety-usefulness Pareto frontier is misleading. Safety research typically creates new safety options without boosting usefulness, while capabilities research typically raises usefulness at safety's expense, so developers tend to choose less safe points.

  10. Harrison ChaseAI score53

    Google Research's Cogentic uses multi-agent proof search to produce verified results

    Google Research's Cogentic is a multi-agent harness running on Gemini that searches for proofs of open theoretical computer science problems without expert hints. It runs rounds where an orchestrator launches provers, two adversarial verifiers must both accept each draft, and shared disk documents store attempts and verified lemmas. The system produced new results on five open problems in online learning, auction theory, and mechanism design, each checked by domain experts.

  11. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    Liquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  12. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    Hugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    AIWhy it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

  13. Thomas WolfAI score40

    Very thoughtful piece from Kevin Buzzard (perfect IMO score, number theorist, pioneer of formal maths in Lean) if math is not just about “human understanding,” then what is it about? and if ai capabilities keep growing exponentially what happens since “mathematics is infinite”?

    Very thoughtful piece from Kevin Buzzard (perfect IMO score, number theorist, pioneer of formal maths in Lean) if math is not just about “human understanding,” then what is it about? and if ai capabilities keep growing exponentially what happens since “mathematics is infinite”?

  14. MIT Technology Review · AIAI score62

    AlphaGo's move 37 shows why LLMs do not truly reason, an AlphaGo team member argues

    Thore Graepel, a core member of the AlphaGo team, argues that current large language models do not truly reason, despite chain-of-thought gains in math and coding. He says they lack an explicit, inspectable epistemic state, keep knowledge and reasoning intertwined in their weights, and often produce post-hoc explanations. He proposes systems that maintain an auditable epistemic state and evaluate each step by how much it resolves uncertainty.