Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Sep 18

Sep 18Fri
  1. WanOfficialAI score16

    Wan3.0 ranks #2 in Overall Video on OpenArt Arena

    AIWan3.0 placed second in the Overall Video category on OpenArt Arena, according to its developers. The post highlights support for 30-second takes, native audio, and any reference input for creators.

    Image from @Alibaba_Wan's post

Sep 17

Sep 17Thu
  1. AnthropicOfficialAI score38

    Anthropic and Adaptyv Bio launch protein design competition with 5,000 validated designs

    AIAnthropic is partnering with Adaptyv Bio on a protein design competition in which over 5,000 designs will be experimentally validated. Anthropic is providing up to $1 million in Claude credits plus funding for experimental validation alongside Adaptyv, while Modal contributes up to $250,000 in compute and Twist Bioscience supplies DNA.

  2. Sierra BlogOfficialAI score38

    Sierra Achieves AIUC-1 Certification for Its AI Agent Platform

    AISierra has become AIUC-1 certified after an independent audit by Schellman and testing by the Artificial Intelligence Underwriting Company (AIUC), a new standard for AI agents that tests resistance to manipulation and unauthorized access. Schellman found that Sierra met all applicable AIUC-1 requirements, and the technical evaluations recur at least quarterly with a full audit each year. The certification complements Sierra's existing SOC 2 Type II, ISO 27001, and ISO 42001 attestations.

  3. SenseTimeOfficialAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

    Image from @SenseTime_AI's post
  4. Ai2 (Allen Institute for AI)OfficialAI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    AINortheastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

Sep 16

Sep 16Wed
  1. Matei ZahariaXAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

  2. Bryan CatanzaroXAI score13

    Bryan Catanzaro to speak at GTC Berlin on open models

    AINVIDIA's Bryan Catanzaro, VP of Applied Deep Learning Research, will present at GTC Berlin on building open models developers can inspect, adapt, and deploy. The post is a conference invitation, with GTC Berlin set for October 20–22, 2026, and no new model or product announced.

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceOfficialAI score44

    Tencent releases WeVisDoc-4B, a document parser that leads OmniDocBench v1.6

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, converts page images into structured Markdown with LaTeX formulas and HTML tables. It scores 95.38 Overall on OmniDocBench v1.6 and a mean Overall of 75.54 across three PureDocBench tracks, ranking first among compared end-to-end parsers in all four reported settings. The model is available on Hugging Face and runs through vLLM, which requires version 0.11.1 or later.

  2. Tencent · new models on Hugging FaceOfficialAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  3. Jazzyear · InsightsNewsAI score67

    HiDream's vivago R1 agent targets five-minute AI video delivery

    AIHiDream.ai launched vivago R1, a content creation agent, globally, with a domestic version upgrade. The company says R1 can output five-minute high-quality videos through agent planning, with a claimed 85% usable-output rate and support for multi-round extensions. It also released HiDream-O1-Video-1.0, a native omni-modal video model supporting single shots of 5 to 20 seconds at 1080p.

  4. Jason WeiXAI score40

    Jason Wei says wet-lab data lets a specialized model beat GPT-6 Astra

    AIJason Wei argues that specialized, often private wet-lab data can let a task-specific model outperform a general frontier model on scientific tasks. He cites Neon, an open-source model that Liam Fedus says was mid-trained and RL-tuned on experimental data using 1,300 H200s to surpass GPT-6 Astra on an analysis benchmark. The post frames this data as a potential moat as work moves toward the frontier of science.

  5. Google AI StudioOfficialAI score72

    Google launches Gemini 3.8 Live and Extended Thinking voice models

    AIGoogle introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents that reason and speak simultaneously. The Extended Thinking version scores 82.6 on Artificial Analysis' Speech to Speech Quality Index and 97.7% on Big Bench Audio, while 3.8 Live targets scale and cost efficiency. Developers can access both through the Gemini API in Google AI Studio, and enterprise and consumer rollouts vary by product.

    Why it matters: The source names the two models, their access paths, and specific benchmark results, showing how the voice agent capabilities differ between the two tiers.

  6. RadixArkOfficialAI score42

    Periodic Labs builds Neon on SGLang and Miles for 2.5x faster inference

    AIPeriodic Labs chose SGLang and Miles to build Neon, an open-source model it says surpasses GPT-6 Astra on its analysis benchmark after mid-training and RL on 1,300 H200s. RadixArk says Periodic extended both frameworks for scientific RL at trillion-parameter scale, delivering more efficient training, lower memory use, and 2.5x faster inference. The work has been contributed back to both projects.

  7. Sebastian RaschkaXAI score28

    GPT-5.6 Astra and Qwen3.8 Max take different Paint approaches

    AIIn a Paint recreation test, GPT-5.6 Astra built the image from layered geometric shapes, while Qwen3.8 Max worked pixel by pixel. Qwen's output looks closer to the original, but Raschka argues this single example does not show either model generalizes better or has stronger computer-use or visual understanding, and it illustrates how benchmarks comparing only final results can be misleading.

    Video from @rasbt's post
  8. Tencent HyOfficialAI score38

    EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle

    AITencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.

    Image from @TencentHunyuan's post

Sep 14

Sep 14Mon
  1. Intern Large ModelsOfficialAI score62

    Intern-S2-397B released in BF16 and FP8 under Apache 2.0

    AIShanghai AI Laboratory's Intern Large Models announced Intern-S2-397B, available in BF16 and FP8 under Apache 2.0. The post reports 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both, and says it was jointly trained across 20+ scientific domains with long-horizon agent RL.

  2. Intern Large ModelsOfficialAI score62

    Intern-S2-397B: Shanghai AI Lab releases open multimodal model for scientific research

    AIIntern Large Models introduces Intern-S2-397B, a multimodal foundation model built for long-horizon scientific research and scientific agents. The post reports leading open-source results on IMO-Proof and AdvancedMathBench, and says the model reaches the level of Gemini 3.1 Pro on those tasks. It is now supported by vLLM and SGLang, with weights on Hugging Face and ModelScope and a chat demo available.

    Image from @intern_lm's post

Sep 13

Sep 13Sun
  1. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score36

    SingProbe adds a streaming guardrail to Step-3.7-Flash without a separate safety model

    AIinclusionAI released Step-3.7-Flash-singprobe, an 8.13M-parameter probe that reuses Step-3.7-Flash hidden states to score query intent, response unsafety, and hallucination risk at every generated token. The probe adds less than 0.5% decode-time overhead and reports 0.9858 R-AUC and 0.9295 T-AUC on streaming safety benchmarks. It is supported through SGLang and vLLM integration branches and loads from Hugging Face by checkpoint ID.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score40

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.5-397B-A17B

    AIinclusionAI has released Qwen3.5-397B-A17B-singprobe, an intrinsic streaming guardrail built on Qwen/Qwen3.5-397B-A17B that scores query intent, response unsafety, and hallucination risk at every generated token using the base model's hidden states. The probe has 8.13M parameters, taps layers 18, 38, and 58, and adds less than 0.5% decode-time overhead. Training code is available at inclusionAI/SingProbe, and the probe runs through SGLang or vLLM integration branches.

  3. Fireworks AI BlogOfficialAI score52

    Fireworks adds DeepSeek-V4.1-Flash, matching GPT-6 Astra coding accuracy at 1/15th the cost

    AIFireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.

  4. Sebastian RaschkaXAI score35

    Raschka's Reasoning from Scratch Round 3 Builds a Math Verifier

    AISebastian Raschka's third "Reasoning from Scratch" video covers building a math verifier for evaluating language models and for later reinforcement learning with verifiable rewards (RLVR) training. The walkthrough covers extracting final answers from boxed outputs, normalizing them, checking mathematical equivalence, and running evaluation on the MATH-500 dataset.

    Video from @rasbt's post
  5. Mike KnoopXAI score50

    Mike Knoop argues intelligence is capped at optimal decision-making

    AIMike Knoop argues intelligence can be measured as the ratio of a decision's quality to the optimal decision, capped at 100%. He says Astra is already 80% optimal on ARC v3 speedruns and identifies horizontal data acquisition and efficiency/cost as the most plausible near-term areas for RSI. Background from @mhmazur reports that GPT-6 Astra scored 100% on the 25 ARC-AGI-3 public games using 6,485 actions versus a human baseline of 17,135.

Sep 12

Sep 12Sat
  1. Mike KnoopXAI score46

    Mike Knoop urges keeping AI research open amid slowdown proposals

    AIMike Knoop says he sees a path to an ARC-AGI-4 benchmark focused on open-ended invention, which he calls the gating capability between zero-sum automation and positive-sum innovation. He argues that coordinated slowdown efforts would likely apply to everyone, including open-source work, and cites chain of thought and the transformer as inventions that grew out of open science research. He concludes the research frontier must stay open to keep humanity on a positive-sum path.

  2. Epoch AI · The Epoch BriefOfficialAI score60

    Epoch Brief covers Huawei chips, Nvidia's GDP effect, and GPT-6 Astra benchmarks

    AIEpoch AI's newsletter reports that Huawei is far behind Nvidia and is unlikely to catch up this decade due to export controls. It also finds official US GDP statistics understate growth by about 0.3 percentage points over the past year, and that GPT-6 Astra set new records on Epoch's evaluations, including the Epoch Capabilities Index.

    Why it matters: The newsletter bundles several analyses of AI chips, GDP measurement, and benchmarks, so it helps readers scan the research agenda behind each finding.

  3. The Algorithmic BridgeBlogAI score52

    AI's Math Breakthroughs Could Starve Mathematics of the Hard Problems It Needs

    AIAlberto Romero argues that AI solving Millennium Prize problems in 2026 threatens mathematics through success, not failure. He draws on Terence Tao's view that struggle shapes mathematicians, and that proof abundance without hard problems could leave fields depleted, like overplanted farmland.

Sep 11

Sep 11Fri
  1. Baseten BlogOfficialAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

  2. VercelOfficialAI score26

    Tailscale's Aperture model router is built on Vercel AI Gateway

    AITailscale offers instant access to hundreds of models for any user in a secure tailnet through its customer-facing model router, Aperture. Aperture is built on Vercel's AI Gateway and offers zero data retention, zero markup with free BYOK, and cost and usage data on every response.

  3. Redwood Research BlogBlogAI score62

    Prompt tuning lifts CoT controllability scores on open models

    AIRedwood Research reports that better prompt templates raise chain-of-thought controllability scores on the CoTControl eval for open-source reasoning models by roughly 2-3x or more. For example, GPT-OSS-120B rose from 5.5% to 15% in the zero-shot setting. The author concludes that current CoT controllability numbers may underestimate what models can do, though the finding does not significantly undermine the view that current models probably cannot consistently evade CoT monitoring.

  4. BAAIOfficialAI score46

    BAAI unveils AREX, a 122B MoE research agent for hard search

    AIBAAI introduced AREX, a research agent built on a 122B-parameter mixture-of-experts model with 10B active parameters. It drafts candidate answers, checks each constraint, and revisits unresolved points rather than running one long search. The post says AREX performs on hard search benchmarks comparable to GPT-5.4.

    Video from @BAAIBeijing's post

Sep 10

Sep 10Thu
  1. Ai2 · new models on Hugging FaceOfficialAI score34

    AstaBrief-8B-SFT: Ai2's 8B model for cited scientific research reports

    AIAi2 released AstaBrief-8B-SFT, an 8B intermediate supervised fine-tuning checkpoint built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. On the ScholarQA-CS2 test set of 100 computer science questions, it scored an average of 83.7 versus 77.3 for base Qwen3-8B, with citation recall at 71.3 versus 64.6. The model is licensed under Apache 2.0 for research and educational use.

  2. Greg BrockmanXAI score36

    GPT-6 Astra tops frontier models in antibody developability prediction

    AIOpenAI's GPT-6 Astra is reported as the best-performing frontier model for antibody developability prediction in one benchmark, outperforming other frontier models tested on properties such as aggregation and stability. The post also says Astra built an interactive antibody visualization in about one hour.