Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 21

Sep 21Mon
  1. Xiaomi MiMo · new models on Hugging FaceAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    AIXiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  2. Xiaomi MiMo · new models on Hugging FaceAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

  3. howie.seriousAI score34

    Agrees with critique that GPT-6 Astra lags on open-ended tasks

    AIResponding to a post by ScarletKc, howie.serious simply agrees with the claim that GPT-6 Astra struggles with open-ended, exploratory work that lacks a fixed correct answer. The main post is a one-word endorsement (), while the quoted post argues GPT models excel at verifiable, goal-defined tasks and that Claude Fable handles open-ended exploration better.

Sep 20

Sep 20Sun
  1. xAI News (Grok)AI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

Sep 19

Sep 19Sat
  1. StepFunAI score20

    StepFun's Step 5 Preview targets finance tasks with FinStepBench evaluations

    AIStepFun says it is focusing Step 5 Preview on finance, judging it on verifying reliable information, reconciling conflicting reports, stating assumptions, and producing consistent, reproducible valuations. The post says the model is evaluated on FinStepBench, covering LiveSearch, CorporateValuation, and DeepResearch, and on FrontierFinance across six investment use cases.

    Image from @StepFun_ai's post
  2. Sebastian RaschkaAI score36

    Raschka's Inference Scaling Part 1: Sampling for Better Accuracy

    AISebastian Raschka starts a series on inference scaling by modifying text generation with temperature scaling, top-p filtering, and multinomial sampling to produce diverse outputs. He says this enables self-consistency and best-of-N approaches that improve answer accuracy by more than 2x. The video covers chain-of-thought prompting, a MATH-500 evaluation, and accuracy versus compute tradeoffs.

    Video from @rasbt's post
  3. Sebastian RaschkaAI score42

    Muon reduces memorization compared with AdamW in nanoGPT training experiments

    AIMuon appears to outperform AdamW because it suppresses memorization, according to WeightWatcher experiments on a single-head nanoGPT model across five seeds. At 10,000 steps, teacher-forced recall of planted sequences was about 62% for AdamW versus under 1% for Muon. The author notes that some Muon layers also show α < 2, so α alone does not explain memorization and individual layers and their ESDs should be examined.

Sep 18

Sep 18Fri
  1. Google ResearchAI score22

    Google Research releases MilleMiglia, a public middle-mile logistics benchmark

    AIGoogle Research has introduced MilleMiglia, a standardized benchmark for optimizing middle-mile logistics, the segment that moves goods across hundreds of miles overnight. The benchmark uses spatial clustering and gravity models to simulate realistic middle-mile delivery scenarios. It addresses the difficulty of optimizing these networks without public data.

    Image from @GoogleResearch's post
  2. SemiAnalysisAI score52

    Engram offloading to DRAM beats SSD for DeepSeek-V4.1-Flash serving on B200

    AISemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.

Sep 17

Sep 17Thu
  1. AnthropicAI score38

    Anthropic and Adaptyv Bio launch protein design competition with 5,000 validated designs

    AIAnthropic is partnering with Adaptyv Bio on a protein design competition in which over 5,000 designs will be experimentally validated. Anthropic is providing up to $1 million in Claude credits plus funding for experimental validation alongside Adaptyv, while Modal contributes up to $250,000 in compute and Twist Bioscience supplies DNA.

  2. Sierra BlogAI score38

    Sierra Achieves AIUC-1 Certification for Its AI Agent Platform

    AISierra has become AIUC-1 certified after an independent audit by Schellman and testing by the Artificial Intelligence Underwriting Company (AIUC), a new standard for AI agents that tests resistance to manipulation and unauthorized access. Schellman found that Sierra met all applicable AIUC-1 requirements, and the technical evaluations recur at least quarterly with a full audit each year. The certification complements Sierra's existing SOC 2 Type II, ISO 27001, and ISO 42001 attestations.

  3. SenseTimeAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

    Image from @SenseTime_AI's post
  4. Ai2 (Allen Institute for AI)AI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    AINortheastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

Sep 16

Sep 16Wed
  1. Matei ZahariaAI score58

    Databricks reports engineering measurements from rolling out Astra to 3,500 engineers

    AIDatabricks rolled out Astra to all of its engineers and reported internal measurements. Engineers given Astra increased overall coding spend by around 60% compared with baseline, and Astra outperformed prior top models on highly complex system design tasks, while the gain on medium or low complexity tasks was unclear.

  2. Matei ZahariaAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceAI score44

    Tencent releases WeVisDoc-4B, a document parser that leads OmniDocBench v1.6

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, converts page images into structured Markdown with LaTeX formulas and HTML tables. It scores 95.38 Overall on OmniDocBench v1.6 and a mean Overall of 75.54 across three PureDocBench tracks, ranking first among compared end-to-end parsers in all four reported settings. The model is available on Hugging Face and runs through vLLM, which requires version 0.11.1 or later.

  2. Tencent · new models on Hugging FaceAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  3. Jazzyear · InsightsAI score67

    HiDream's vivago R1 agent targets five-minute AI video delivery

    AIHiDream.ai launched vivago R1, a content creation agent, globally, with a domestic version upgrade. The company says R1 can output five-minute high-quality videos through agent planning, with a claimed 85% usable-output rate and support for multi-round extensions. It also released HiDream-O1-Video-1.0, a native omni-modal video model supporting single shots of 5 to 20 seconds at 1080p.

  4. Elad GilAI score38

    Periodic Labs' open model Neon reportedly beats GPT-6 Astra on materials benchmark

    AIPeriodic Labs says it used 1,300 H200 GPUs and months of its lab data to mid-train and RL an open-source model called Neon, which it claims surpasses GPT-6 Astra on its analysis benchmark. The company says it is focusing first on hard materials science problems, including superconductors, magnets, and semiconductor materials.

  5. Jason WeiAI score40

    Jason Wei says wet-lab data lets a specialized model beat GPT-6 Astra

    AIJason Wei argues that specialized, often private wet-lab data can let a task-specific model outperform a general frontier model on scientific tasks. He cites Neon, an open-source model that Liam Fedus says was mid-trained and RL-tuned on experimental data using 1,300 H200s to surpass GPT-6 Astra on an analysis benchmark. The post frames this data as a potential moat as work moves toward the frontier of science.

  6. Google AI StudioAI score72

    Google launches Gemini 3.8 Live and Extended Thinking voice models

    AIGoogle introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents that reason and speak simultaneously. The Extended Thinking version scores 82.6 on Artificial Analysis' Speech to Speech Quality Index and 97.7% on Big Bench Audio, while 3.8 Live targets scale and cost efficiency. Developers can access both through the Gemini API in Google AI Studio, and enterprise and consumer rollouts vary by product.

    Why it matters: The source names the two models, their access paths, and specific benchmark results, showing how the voice agent capabilities differ between the two tiers.

  7. RadixArkAI score42

    Periodic Labs builds Neon on SGLang and Miles for 2.5x faster inference

    AIPeriodic Labs chose SGLang and Miles to build Neon, an open-source model it says surpasses GPT-6 Astra on its analysis benchmark after mid-training and RL on 1,300 H200s. RadixArk says Periodic extended both frameworks for scientific RL at trillion-parameter scale, delivering more efficient training, lower memory use, and 2.5x faster inference. The work has been contributed back to both projects.

  8. Sebastian RaschkaAI score28

    GPT-5.6 Astra and Qwen3.8 Max take different Paint approaches

    AIIn a Paint recreation test, GPT-5.6 Astra built the image from layered geometric shapes, while Qwen3.8 Max worked pixel by pixel. Qwen's output looks closer to the original, but Raschka argues this single example does not show either model generalizes better or has stronger computer-use or visual understanding, and it illustrates how benchmarks comparing only final results can be misleading.

    Video from @rasbt's post
  9. Tencent HyAI score38

    EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle

    AITencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.

    Image from @TencentHunyuan's post

Sep 14

Sep 14Mon
  1. Intern Large ModelsAI score62

    Intern-S2-397B: Shanghai AI Lab releases open multimodal model for scientific research

    AIIntern Large Models introduces Intern-S2-397B, a multimodal foundation model built for long-horizon scientific research and scientific agents. The post reports leading open-source results on IMO-Proof and AdvancedMathBench, and says the model reaches the level of Gemini 3.1 Pro on those tasks. It is now supported by vLLM and SGLang, with weights on Hugging Face and ModelScope and a chat demo available.

    Image from @intern_lm's post

Sep 13

Sep 13Sun
  1. inclusionAI (Ant Ling) · new models on Hugging FaceAI score36

    SingProbe adds a streaming guardrail to Step-3.7-Flash without a separate safety model

    AIinclusionAI released Step-3.7-Flash-singprobe, an 8.13M-parameter probe that reuses Step-3.7-Flash hidden states to score query intent, response unsafety, and hallucination risk at every generated token. The probe adds less than 0.5% decode-time overhead and reports 0.9858 R-AUC and 0.9295 T-AUC on streaming safety benchmarks. It is supported through SGLang and vLLM integration branches and loads from Hugging Face by checkpoint ID.