Skip to contentSkip to stories

Updated

#Paper/Research

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 5

Oct 5Mon
  1. Redwood Research BlogAI score62

    Frontier models give different decision theory answers depending on who is asking

    AIRedwood Research reports that Claude Fable 5.1 almost always names FDT or FDT/UDT when no academic cue is given, but names CDT about 30% to 100% of the time when the prompt signals mainstream academic philosophy. Similar shifts appear on moral realism, p-zombie conceivability, P(doom), and AGI timelines, which the author treats as a form of sycophancy or audience awareness. The post recommends caution when interpreting attitude evals where no human consensus exists, and notes the effect is weaker in other models tested.

  2. GitHub Blog · AI & MLAI score63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    AIGitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    Why it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

  3. MIT Technology Review · AIAI score30

    Enterprise AI agents need organizational knowledge to reach production, survey finds

    AIA survey of 300 data, AI, and technology executives found only 34% of organizations' agentic AI projects reach production, with legacy systems, security concerns, and missing knowledge context as main obstacles. Production leaders, who advance 61% of projects beyond pilot, show stronger semantic knowledge capabilities. Most firms plan to invest in retrieval pipelines, AI-ready APIs, retrieval-augmented generation, and knowledge graphs.

  4. Clément DelangueAI score62

    Hugging Face turns 10 coding harnesses into RL environments via a capture proxy

    AIHugging Face says a capture proxy lets reinforcement learning train open models inside unmodified coding harnesses such as Claude Code, Codex, and OpenCode. The proxy records the exact token IDs and logprobs vLLM samples and hands them to TRL for training. On LFM2.5-2.6B, training in four harnesses at once raised OpenCode results from 34% to 58%, while SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.

  5. IEEE Spectrum · AIAI score58

    Mathematicians Debate OpenAI's Navier-Stokes Claim and AI's Impact on the Field

    AIMathematicians at the Heidelberg Laureate Forum discussed AI companies, including OpenAI, Anthropic, and Google, solving longstanding math problems. OpenAI announced it had solved the Navier-Stokes existence and smoothness problem, a claim the article says is still awaiting verification, and Harris criticized the company's conduct toward a mathematician. Researchers also warn that AI solutions may lack understandable methods and are changing how academics work.

Oct 4

Oct 4Sun
  1. Apple Machine Learning ResearchAI score22

    Apple Study Examines How Users Negotiate Ontological Boundaries in Personal Sensing Systems

    AIApple and Stanford researchers built two open-ended probes using a Wizard of Oz technique so participants could train personalized machine learning systems on phenomena they defined themselves. In a week-long exploratory study, participants identified four sites where ontological boundaries were negotiated: the boundaries of a phenomenon, the subject as part of relations, signal versus noise, and the objectivity of data. The paper offers starting points for supporting boundary negotiation through design.

Oct 3

Oct 3Sat
  1. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

  2. Sebastian RaschkaAI score38

    Raschka's Reasoning from Scratch covers RLVR and GRPO implementation

    AISebastian Raschka released round six of his Reasoning from Scratch series, introducing Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO) with an implementation. The video covers accuracy and format rewards, DeepSeek-R1 training, and GRPO versus PPO, then walks through a training loop and evaluates checkpoints on MATH-500.

Oct 2

Oct 2Fri
  1. Baseten BlogAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

  2. AI at MetaAI score22

    Muse Spark helps prove finite-time blow-up in a laser-inspired wave model

    AIWith help from Muse Spark, researchers proved that a wave in a laser-inspired model must blow up in finite time under the conditions studied. The result comes from a tug-of-war between one effect squeezing the wave inward and another spreading it out. The paper is titled finite-time blow-up of radial negative-energy solutions for the mass-critical biharmonic nonlinear Schrödinger equation.

  3. AI at MetaAI score61

    Meta shares six math papers from mathematician-AI collaborations on open problems

    AIAI at Meta says mathematicians used Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the standard meta.ai chat interface to find solutions to open problems. The company is sharing six resulting papers, each marking which passages were drafted primarily by humans or AI, with mathematicians guiding the work and a second group reviewing it.

  4. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  5. Google ResearchAI score60

    Google's TEE-based federated learning system adds verifiable privacy guarantees

    AIGoogle announces a next-generation federated learning system that uses Trusted Execution Environments to provide verifiable, auditable data anonymization. The system publishes access policies to a public transparency log and is deployed in Gboard, which has launched English and Japanese next-word prediction models with stronger privacy guarantees and improved accuracy. Training time has also sped up significantly because computation moved to the server and is parallelized across many machines.

    Why it matters: The post shows how Trusted Execution Environments make federated learning's privacy claims externally verifiable, rather than relying on trust in the server operator.

  6. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

Oct 1

Oct 1Thu
  1. Sundar PichaiAI score60

    Google DeepMind's SynthID Bio watermarks AI-designed protein sequences

    AIGoogle DeepMind announced SynthID Bio, a family of watermarking methods for AI-generated biological designs. According to the quoted post, the team can embed an imperceptible signature directly into protein sequences without affecting their biological function. Sundar Pichai called it a big step forward for scientific integrity and biosecurity.

  2. Latent.SpaceAI score60

    Recursive Language Models explained by MIT's Alex Zhang on coding agents

    AIA Latent.Space podcast episode features MIT researcher Alex Zhang explaining recursive language models (RLMs). He discusses why Claude Code, Codex, and Pi are basically the same, and how RLMs use code, context offloading, and recursive subagents to generalize across tasks. The episode also covers OpenAI's 10,000-agent, 130B-output-token experiment and academia's freedom to pursue ambitious research bets.

  3. Apple Machine Learning ResearchAI score28

    Language Discrimination Narrows Multilingual Speech Model Gap, Study Finds

    AIResearchers Maureen de Seyssel, Jie Chi, and Zakaria Aldeneh found that strengthening language discrimination during pretraining reduces the performance gap between multilingual and monolingual HuBERT speech models. In a controlled English/French setting, phone-ABX error fell from 11.6% to 10.4%, close to the monolingual 10.8%, while lexical sWUGGY scores rose from 52.1% to 56.7%. The gains were largest when language discrimination was introduced in the first training iteration.

  4. Epoch AIAI score62

    Epoch AI estimates how many concurrent AI agents 2025–27 memory shipments could run

    AIEpoch AI estimates that high-bandwidth memory shipped in 2025–27 could eventually support about 30–170 million concurrent frontier-model agents once fully deployed and allocated. Using DeepSeek V4 Pro serving benchmarks, the estimate rises to about 1.9 billion concurrent agents. The authors compare the implied API-equivalent spending of $2.6–5.3 trillion per year with projected developer revenue of roughly $1 trillion by end-2027, suggesting demand may lag supply.

    Why it matters: The analysis converts HBM shipment data into concurrent agent capacity and compares it with projected API revenue, showing where compute buildout may outpace demand.

  5. Apple Machine Learning ResearchAI score34

    Limits of Confidence-Based Sampling in Discrete Diffusion Models

    AIApple Machine Learning Research reports that discrete diffusion steps match the training distribution only when simultaneously written token positions are conditionally independent given already-fixed tokens. The authors show that per-position distributions cannot determine such dependence, and on the synthetic ScanAndAdd task, confidence-ranked groups of two or more positions were dependent and produced a generated distribution 29 times the sampling-noise floor in total variation.

  6. PyTorch BlogAI score38

    TLX-Optimized Jagged Flash Attention Beats FA4 on Blackwell B200 for Meta GEM

    AIMeta's Jagged Flash Attention kernel, built with TLX on NVIDIA Blackwell B200, outperforms FlashAttention-4 (May 2026 version) on GEM's jagged shapes by about 13% on the forward pass and about 50% on the backward pass. The TLX attention kernel is roughly 3.2K lines of Triton-level code, about 3× shorter than FA4's ~10K-line CuteDSL kernels. The benchmarks use bfloat16 on B200.

  7. AnthropicAI score38

    Harvard physicist builds toolkit to match Claude with science calculations

    AIHarvard physicist Matthew Schwartz argues that LLMs are poorly matched to science when used as human-style collaborators, so he built a toolkit for exact quantitative calculations. Working with Claude, the approach surfaced connections to ecology, population genetics, and a dozen other fields, with domain experts steering it toward interesting questions.

  8. Goodfire ResearchAI score60

    Goodfire proposes protein embedding monitors for biosecurity risks in AI agents

    AIGoodfire Research developed sequence-aware monitors using protein language model embeddings to flag concerning biological sequences in dual-use AI agent tasks. On a custom benchmark, the monitors outperformed frontier model safeguards with fewer refusals on benign requests, and they held up better against paraphrasing and fragmentation attacks. The paraphrase results rely on in-silico estimates and do not establish whether the redesigned proteins keep biological activity, and the monitors run in milliseconds per sequence.

    Why it matters: The post gives a concrete benchmark setup and fragmentation results, showing how sequence embeddings can separate dual-use biology requests that task-based safeguards handle poorly.

  9. Prime IntellectAI score34

    Qwen3.6 reward rises 2.8x via GRPO on Hosted Training

    AIPrime Intellect reports that after about 100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain. Qwen3.5, trained the same way, reached 0.356, suggesting the method works across model families. Both post-trained models finished well ahead of other open models and narrowed the gap to Claude Opus 4.8, with Qwen3.6 activating only 3B parameters per token.