Skip to contentSkip to stories

Updated

All AI news

Oct 8

Oct 8Thu
  1. Andrew CurranAI score13

    Andrew Curran Posts "The saga continues" Amid Tightened κ Result

    AIAndrew Curran posted a brief "The saga continues" update, with no clear publisher or model identified. Quoted context from @0xdoug reports a validated, merged PR that tightened κ from 2⁻¹⁸² to 2⁻¹⁵, described as a 500-thousand-fold improvement over the previous result and a 2^167-fold improvement over the original OpenAI result. The quoted post credits a community effort and says results are being verified and published.

  2. Elvis SaraviaAI score46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    AIRSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  3. Google ResearchAI score14

    Google Research demos EnvHarness for co-evolving LLM agents and environments at COLM 2026

    AIGoogle Research is presenting EnvHarness, a flexible framework that enables co-evolution between LLM agents and their training environments, at the #COLM2026 Google booth #107 today at 11:00 AM PT. The post notes that static environments limit agent growth, and EnvHarness is described as a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and adaptability.

  4. Goodfire ResearchAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  5. The Next PlatformAI score43

    How Distributed AI Training Changes the Network Between Datacenters

    AILarge-scale AI training is spreading across multiple datacenters, with Google, Microsoft, AWS, Meta, and CoreWeave cited as examples. Because synchronized GPU clusters must exchange data in bursts, inter-site links can become a bottleneck, which Cisco estimates may require aggregate bandwidth about 14x a conventional DCI baseline.

  6. OpenBMBAI score36

    ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%

    AIReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.

  7. QbitAIAI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    AIUniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  8. MarkTechPostAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    AINVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

  9. PandailyAI score46

    Galbot and Tsinghua's LATENT Wins IROS Award for Humanoid Tennis Forehand

    AIA Galbot, Tsinghua University and collaborators paper won IROS 2026's Best Entertainment and Amusement Paper Award for LATENT, a humanoid tennis-return method trained on imperfect amateur motion-capture clips. In simulation, the full forehand policy succeeded on 96.52 percent of returns, versus 71.85 percent for PULSE. On a real Unitree G1, the paper reports 90.90 percent forehand success across 20 consecutive rallies, with motion capture still used rather than the robot's own cameras.

  10. PandailyAI score38

    Huawei Presents Experimental XMFS Shared-Memory Filesystem at LPC 2026

    AIHuawei engineers presented XMFS, an experimental Linux kernel prototype filesystem, at the Linux Plumbers Conference in Prague on October 5. It aims to let applications reach cross-node shared memory on CXL 3.0 or Huawei unified bus servers through standard POSIX file calls. The code exists only on openEuler, not in the mainline Linux kernel.

  11. Anthropic ResearchAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    AIJohns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    Why it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

  12. Artificial Analysis ArticlesAI score50

    Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

    AIHarvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

Oct 7

Oct 7Wed
  1. KhazixAI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    AIOpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. Elvis SaraviaAI score67

    Tool-using multimodal models refuse harmful requests less often, NVIDIA study finds

    AIA NVIDIA study accepted at NeurIPS 2026 reports that multimodal models refuse harmful requests less reliably when they call tools. Refusal failures rise by up to 68.7% relative and by 17.7% on average across the models tested, including Claude Opus 4.6 and 4.7 and Gemini Agentic Vision. The authors attribute this to tool outputs crowding out the original harmful intent and to attention shifting toward describing tool results. Re-inserting the original request and image before the final response restores part of the lost refusals.

  3. vLLMAI score46

    vLLM-Omni technical report unifies serving for omni-modality generation

    AIThe vLLM team released a technical report on vLLM-Omni, a unified serving runtime for omni-modality generation spanning multi-stage autoregressive pipelines, iterative diffusion, and stateful sessions. Current LLM servers and diffusion stacks each cover only one of these patterns, pushing deployments to stitch disjoint runtimes together. vLLM-Omni offers a shared control plane in which an orchestrator advances requests across stages, specialized engines handle compute, and a connector carries payloads.

  4. Apple Machine Learning ResearchAI score42

    Apple's Normalizing Trajectory Models generate images in four steps with exact likelihood

    AIApple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow trained with exact likelihood. The model matches or outperforms strong image generation baselines on text-to-image benchmarks in just four sampling steps while retaining exact likelihood over the generative trajectory.

  5. Waymo BlogAI score42

    Sober Drivers Still Face Nearly 4x Nighttime Fatal Crash Risk, Waymo Study Finds

    AIWaymo research found that even fully sober human drivers face nighttime fatal crash risk 3.1 to 3.9 times higher than daytime risk, pointing to systemic hazards beyond impairment. The study used an exposure reconstruction model across the 50 most populous U.S. urban areas, showing removing alcohol-involved drivers lowers the average urban fatal crash rate by 23%, from 1.42 to 1.10 per 100 million miles.

  6. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.