Skip to contentSkip to stories

Updated

#Open source/Repo

Oct 8

Oct 8Thu
  1. Artificial AnalysisAI score28

    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs.

    AICost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

  2. Elvis SaraviaAI score46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    AIRSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  3. PandailyAI score38

    Huawei Presents Experimental XMFS Shared-Memory Filesystem at LPC 2026

    AIHuawei engineers presented XMFS, an experimental Linux kernel prototype filesystem, at the Linux Plumbers Conference in Prague on October 5. It aims to let applications reach cross-node shared memory on CXL 3.0 or Huawei unified bus servers through standard POSIX file calls. The code exists only on openEuler, not in the mainline Linux kernel.

Oct 7

Oct 7Wed
  1. KhazixAI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    AIOpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. vLLMAI score46

    vLLM-Omni technical report unifies serving for omni-modality generation

    AIThe vLLM team released a technical report on vLLM-Omni, a unified serving runtime for omni-modality generation spanning multi-stage autoregressive pipelines, iterative diffusion, and stateful sessions. Current LLM servers and diffusion stacks each cover only one of these patterns, pushing deployments to stitch disjoint runtimes together. vLLM-Omni offers a shared control plane in which an orchestrator advances requests across stages, specialized engines handle compute, and a connector carries payloads.

Oct 6

Oct 6Tue

Oct 5

Oct 5Mon
  1. GitHub Blog · AI & MLAI score63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    AIGitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    Why it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

  2. Clément DelangueAI score62

    Hugging Face turns 10 coding harnesses into RL environments via a capture proxy

    AIHugging Face says a capture proxy lets reinforcement learning train open models inside unmodified coding harnesses such as Claude Code, Codex, and OpenCode. The proxy records the exact token IDs and logprobs vLLM samples and hands them to TRL for training. On LFM2.5-2.6B, training in four harnesses at once raised OpenCode results from 34% to 58%, while SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.

Oct 2

Oct 2Fri
  1. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  2. Google ResearchAI score60

    Google's TEE-based federated learning system adds verifiable privacy guarantees

    AIGoogle announces a next-generation federated learning system that uses Trusted Execution Environments to provide verifiable, auditable data anonymization. The system publishes access policies to a public transparency log and is deployed in Gboard, which has launched English and Japanese next-word prediction models with stronger privacy guarantees and improved accuracy. Training time has also sped up significantly because computation moved to the server and is parallelized across many machines.

    Why it matters: The post shows how Trusted Execution Environments make federated learning's privacy claims externally verifiable, rather than relying on trust in the server operator.

Oct 1

Oct 1Thu

Sep 26

Sep 26Sat

Sep 25

Sep 25Fri

Sep 21

Sep 21Mon
  1. Microsoft ResearchAI score50

    Microsoft Research open-sources RetroChimera, a retrosynthesis model published in Nature

    AIMicrosoft Research published RetroChimera, a retrosynthesis framework that combines the R-SMILES 2 Transformer model and the NeuralLoc graph neural network through learned ensembling to propose synthesis routes for small molecules. In blind tests, PhD-level chemists preferred its individual reaction predictions over those from preceding models and recorded literature reactions. The implementation and weights are open-sourced for researchers developing new medicinal molecules and materials.

Sep 17

Sep 17Thu

Aug 27

Aug 27Thu
  1. LMSYS OrgAI score47

    MiniMax-H3 gets up to 6.24x speedup on 8×H200 GPUs

    AIMiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

Jul 15

Jul 15Wed
  1. Liquid AI NewsletterAI score38

    Liquid AI Releases Antidoom and IFStruct to Fix Reasoning Loops and Schema Errors

    AILiquid AI released Antidoom, an open-source method that retrains a single overtrained token to eliminate "doom loops" in small reasoning models. On LFM2.5-2.6B and Qwen3.5-4B, loop rates fell from 10.2% to 1.4% and from 22.9% to 1%, respectively. The company also released IFStruct, an open-source benchmark measuring whether model outputs satisfy a schema, where LFM2.5-350M rose from 21.10% to 44.90% after training.

Apr 23

Apr 23Thu
  1. OpenAI Alignment Research BlogAI score44

    OpenAI Open-Sources Chain-of-Thought Monitorability Evaluation Datasets and Code

    AIOpenAI is releasing a subset of datasets, reference code, and the g-mean 2 metric for evaluating chain-of-thought monitorability. The release includes most datasets from its monitorability suite, while some evaluations relying on private or restricted data were omitted. The company says it will keep reporting monitorability results in future frontier reasoning model system cards.

Mar 26

Mar 26Thu

Mar 17

Mar 17Tue
  1. BAAIAI score46

    BAAI unveils RoboBrain-Dex, dexterous manipulation trained on human egocentric data

    AIBAAI has released RoboBrain-Dex, a dexterous manipulation model for embodied intelligence trained on large-scale, diverse human egocentric data rather than massive robot teleoperation datasets. The company says this approach yields strong generalization, marking a shift from small-data, weakly generalizing methods toward big-data robotic manipulation. The code has been open-sourced on GitHub.

Mar 13

Mar 13Fri
  1. Berkeley AI ResearchAI score34

    SPEX and ProxySPEX Identify Influential LLM Interactions at Scale with Fewer Ablations

    AIBerkeley AI Research introduces SPEX, a signal-processing framework that identifies influential interactions in LLMs using far fewer ablations than exhaustive analysis. A hierarchy-based extension, ProxySPEX, matches SPEX performance with around 10x fewer ablations. The methods apply to feature, data, and model component attribution.

Feb 19

Feb 19Thu

Dec 12, 2025

Dec 12, 2025Fri
  1. Apple · new models on Hugging FaceAI score46

    Apple's SHARP Turns a Single Photo into a 3D Scene in Under a Second

    AIApple has released SHARP, a model that generates a 3D Gaussian representation of a scene from a single photograph in less than a second on a standard GPU. The output renders in real time as high-resolution photorealistic views of nearby camera positions, with metric absolute scale, and the paper reports reductions of 25–34% in LPIPS and 21–43% in DISTS versus the best prior model.

Aug 11, 2021

Aug 11, 2021Wed