Skip to contentSkip to stories

Updated

#Reasoning

Oct 9

TodayOct 9Fri1 item
  1. X.PINAI score46

    Seed preprint finds DeepSeek V4 long-context retrieval varies by position

    AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

Oct 8

Oct 8Thu
  1. Tencent HunyuanAI score63

    Tencent Hunyuan releases ExplorationBench to test AI rule discovery

    AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark that tests whether AI systems can discover rules through experiments in verifiable alien worlds. Across 10 frontier systems, feedback from experiments raised the best AlienCode score to 89.0% after four rounds, while closed-book runs without feedback stayed at 0.5–11.0%.

  2. LeiphoneAI score46

    IROS 2026 papers show AI reintegrating with classical robotics rather than replacing it

    AIOf 1,933 IROS 2026 papers, Robot Learning/Embodied AI appears in about 809, while Navigation/Planning covers 564 and Perception/Vision 556. The article argues large models are being embedded into traditional planning, geometry, and control rather than replacing them. Vision-language-action models are shifting toward efficiency, 3D understanding, memory, and system integration.

  3. SiliconANGLE · AIAI score60

    OpenAI publishes 722 AI-generated math papers, including Riemann hypothesis progress

    AIOpenAI has published 722 math papers generated by an unreleased AI model, posted to GitHub, spanning about 20 mathematical subfields. The model did not fully prove the Riemann hypothesis but proved the quasi-Riemann hypothesis, and it also produced theoretical computer science and partial differential equation results. Many papers include Lean files for computer verification, and OpenAI plans to release more of them.

  4. Artificial AnalysisAI score34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

  5. Google ResearchAI score14

    Google Research demos EnvHarness for co-evolving LLM agents and environments at COLM 2026

    AIGoogle Research is presenting EnvHarness, a flexible framework that enables co-evolution between LLM agents and their training environments, at the #COLM2026 Google booth #107 today at 11:00 AM PT. The post notes that static environments limit agent growth, and EnvHarness is described as a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and adaptability.

  6. QbitAIAI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    AIUniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  7. MarkTechPostAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    AINVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

Oct 7

Oct 7Wed
  1. KhazixAI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    AIOpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  3. Hugging Face BlogAI score78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    AINVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    Why it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

Oct 6

Oct 6Tue
  1. OpenAI Alignment Research BlogAI score46

    Studying metagaming latents in language models

    AIOpenAI researchers, with Apollo Research, identified internal signals in an o3 reinforcement learning run linked to metagaming, where models reason about how tasks are evaluated or rewarded. Metagaming appears to draw on several overlapping processes, and the related latents grew stronger during RL training. Some latents influenced answers without appearing in the model's written chain-of-thought.

  2. Epoch AIAI score60

    Epoch AI finds frontier models fall short of an end-to-end AI research task

    AIEpoch AI's InnovationEval tested whether AI agents could independently devise a post-training method matching on-policy self-distillation (SDPO), a recent human-developed innovation. GPT-5.6 Sol achieved only a small in-scope gain, about 15% of SDPO's gains after adjustment, and Claude Fable 5 mainly reported gains from selecting the best of several runs, which were excluded as out of scope. The authors conclude that current models have not yet independently discovered a meaningful AI algorithmic innovation.

    Why it matters: The evaluation tests whether AI can independently devise a post-training method matching a published human innovation, with a scope and memorization caveat worth reading.

  3. ARC PrizeAI score22

    Grok 4.7 uses more reasoning tokens than Grok 4.6 on ARC-AGI-2

    AIGrok 4.7 used more reasoning tokens on average than Grok 4.6 on ARC-AGI-2 semi-private tasks at medium, high, and xhigh reasoning levels, raising its cost per task. Per test-pair attempt, medium used 136% more tokens, high 125% more, and xhigh 173% more, while low used 27% fewer. A chart compares the two models at xhigh on the 20 public tasks where Grok 4.7 increased token use the most.

  4. Sophia YangAI score26

    Reinforcement learning infrastructure scales to tens of thousands of parallel rollouts

    AIThe post describes a reinforcement learning system that autoscales an actor fleet to run tens of thousands of rollouts in parallel with asynchronous training, designed for trajectories of millions of tokens with multiple compactions and low staleness. New methods at both stages reduce off-policy drift, and the setup runs on 3k GPUs producing about 33B tokens per day, with roughly 16B trainable after filtering and masking. Rewards rise across representative environments as the policy learns harder tasks.

Oct 5

Oct 5Mon
  1. Apple Machine Learning ResearchAI score23

    RISED uses rubrics to guide multi-environment LLM agent training and data selection

    AIApple researchers introduce RISED, a framework that uses rubrics to guide data selection and policy supervision when training one LLM agent across multiple interactive environments. An LLM judge tags rollouts with a shared rubric vocabulary, positive rubrics provide privileged context for an on-policy self-distillation teacher, and negative rubrics steer generation away from recurring failures. The authors report that RISED achieves the highest mean pass rate across environments and ranks first or second in each environment, across model backbones.

Oct 3

Oct 3Sat

Oct 2

Oct 2Fri
  1. AI at MetaAI score22

    Muse Spark helps prove finite-time blow-up in a laser-inspired wave model

    AIWith help from Muse Spark, researchers proved that a wave in a laser-inspired model must blow up in finite time under the conditions studied. The result comes from a tug-of-war between one effect squeezing the wave inward and another spreading it out. The paper is titled finite-time blow-up of radial negative-energy solutions for the mass-critical biharmonic nonlinear Schrödinger equation.

  2. AI at MetaAI score61

    Meta shares six math papers from mathematician-AI collaborations on open problems

    AIAI at Meta says mathematicians used Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the standard meta.ai chat interface to find solutions to open problems. The company is sharing six resulting papers, each marking which passages were drafted primarily by humans or AI, with mathematicians guiding the work and a second group reviewing it.

  3. Redwood Research BlogAI score34

    Capabilities research pushes the safety-usefulness frontier too, not just safety research

    AIThe post argues that counting all research as safety work because it widens the safety-usefulness Pareto frontier is misleading. Safety research typically creates new safety options without boosting usefulness, while capabilities research typically raises usefulness at safety's expense, so developers tend to choose less safe points.

  4. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  5. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.