Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 13

Sep 13Sun
  1. inclusionAI (Ant Ling) · new models on Hugging FaceAI score40

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.5-397B-A17B

    AIinclusionAI has released Qwen3.5-397B-A17B-singprobe, an intrinsic streaming guardrail built on Qwen/Qwen3.5-397B-A17B that scores query intent, response unsafety, and hallucination risk at every generated token using the base model's hidden states. The probe has 8.13M parameters, taps layers 18, 38, and 58, and adds less than 0.5% decode-time overhead. Training code is available at inclusionAI/SingProbe, and the probe runs through SGLang or vLLM integration branches.

  2. Fireworks AI BlogAI score52

    Fireworks adds DeepSeek-V4.1-Flash, matching GPT-6 Astra coding accuracy at 1/15th the cost

    AIFireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.

Sep 12

Sep 12Sat
  1. Epoch AI · The Epoch BriefAI score60

    Epoch Brief covers Huawei chips, Nvidia's GDP effect, and GPT-6 Astra benchmarks

    AIEpoch AI's newsletter reports that Huawei is far behind Nvidia and is unlikely to catch up this decade due to export controls. It also finds official US GDP statistics understate growth by about 0.3 percentage points over the past year, and that GPT-6 Astra set new records on Epoch's evaluations, including the Epoch Capabilities Index.

    Why it matters: The newsletter bundles several analyses of AI chips, GDP measurement, and benchmarks, so it helps readers scan the research agenda behind each finding.

Sep 11

Sep 11Fri
  1. Baseten BlogAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

Sep 10

Sep 10Thu
  1. Ai2 · new models on Hugging FaceAI score34

    AstaBrief-8B-SFT: Ai2's 8B model for cited scientific research reports

    AIAi2 released AstaBrief-8B-SFT, an 8B intermediate supervised fine-tuning checkpoint built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. On the ScholarQA-CS2 test set of 100 computer science questions, it scored an average of 83.7 versus 77.3 for base Qwen3-8B, with citation recall at 71.3 versus 64.6. The model is licensed under Apache 2.0 for research and educational use.

  2. Cognition Blog (Devin, Windsurf)AI score66

    Cognition releases SWE-2, a coding model trained with cost-penalized RL

    AICognition introduces SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less. The post attributes the gains to an RL algorithm that trains all reasoning-effort levels in one run, with cost penalties tuned to the base model's Pareto frontier. SWE-2 is available starting today in Devin Desktop and CLI, with rollout to Devin Web and Fusion.

    Why it matters: The post explains how the cost penalty and length-weighted baseline are derived, which helps readers judge the tradeoffs in coding model post-training.

  3. Amazon ScienceAI score55

    Research agents avoid overfitting when their winning strategies compress into few tokens

    AIAmazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.

  4. DeepSeek API NewsAI score72

    DeepSeek releases V4.1-Flash with native multimodal support and API updates

    AIDeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native multimodal visual understanding. The API now serves it under the model name deepseek-flash, while V4 Flash and V4 Flash Vision Exp were retired and routed to V4.1 Flash. API prices were reduced with the release, and V4 Pro remains available after September 14, 2026.

    Why it matters: The release lists benchmark results alongside API model-name changes and retirements, so developers can check both capability claims and migration steps.

Sep 9

Sep 9Wed
  1. DeepSeek · new models on Hugging FaceAI score78

    DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

    AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

    Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

  2. Fireworks AI BlogAI score60

    Genspark's Gen-1 Slides matches Opus 5 decks at about one-tenth the cost per deck

    AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.

    Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.

  3. Cognition Blog (Devin, Windsurf)AI score82

    Cognition's Devin factors RSA-260 using a GPU lattice siever

    AICognition's Devin agent, directed by Eric Lu, factored the 260-digit RSA-260 number using a new GPU implementation of the general number field sieve built on CADO-NFS. The author estimates the run cost about 13.5 GPU-years, roughly $400k at market prices, and projects RSA-1024 factoring at around $30M, while RSA-2048 is not meaningfully affected.

    Why it matters: The source gives a full cost breakdown and scaling estimates for RSA factoring on GPUs, showing how far the cost of breaking RSA-1024 has fallen.

Sep 8

Sep 8Tue
  1. Google Developers BlogAI score36

    Google Developers Blog outlines behavioral evals for guarding AI coding agents against regressions

    AIGoogle Developers Blog argues that teams building AI coding agents should replace end-to-end benchmark scores with behavioral evaluations that test discrete, observable actions. Examples include asking clarifying questions on underspecified prompts, running a local validator before marking a build change complete, and consulting live search for current information. The post recommends fast, deterministic unit-style checks, outcome-based LLM-as-a-judge checks for complex tasks, and batch runs that track aggregate pass rates over time.

  2. Google DeepMindAI score74

    Google DeepMind launches AlphaGenome Atlas to predict 9 billion DNA variant effects

    AIGoogle DeepMind has introduced AlphaGenome Atlas, a platform with predicted molecular effects for 9 billion single-nucleotide variants in the human genome. It is free for academic research through a web portal, and the AlphaGenome Variant Impact score condenses predictions from AlphaGenome and AlphaMissense into one number for ranking variants. The source says collaborators used it to identify variants in unsolved rare disease cases and to find rare non-coding variants linked to traits.

    Why it matters: The source details how precomputed variant predictions, a single impact score, and linked feature attributions make genome-wide mutation effects searchable for researchers without coding skills.

Sep 7

Sep 7Mon
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score45

    openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning

    AIOpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning. The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline. The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.

Sep 6

Sep 6Sun
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score62

    OpenBMB releases MiniCPM5-2B, a 2B open-source model with open training data

    AIOpenBMB has released MiniCPM5-2B, a dense 2B Transformer built for on-device and resource-constrained deployment, with an average score of 53.9 in its comparison set. The release also opens the UltraData datasets behind it, including UltraX, UltraData-Code, UltraData-SFT-Agent-2609 and UltraData-RL-2609, and includes GGUF, MLX, GPTQ and DSpark variants for common runtimes.

    Why it matters: The release pairs a 2B model with open training datasets and reports per-benchmark comparisons against named same-size and larger models, letting readers check the claims directly.

Sep 4

Sep 4Fri
  1. Tencent · new models on Hugging FaceAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

  2. Tencent · new models on Hugging FaceAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

Sep 3

Sep 3Thu
  1. Google DeepMind · The KeywordAI score72

    Google DeepMind releases WeatherNext 3, a global weather model with hourly satellite-based forecasts

    AIGoogle DeepMind and Google Research introduced WeatherNext 3, which generates hourly global forecasts at up to 5-kilometer resolution using live geostationary satellite data. The company reports that precipitation forecasts improved by up to 60% against IMERG in medium-range evaluations, and that longer-range precipitation forecasts are up to 50% more accurate. The model is now available across Search, Gemini, Google Maps, Google Maps Platform Weather API, Google Earth Engine, BigQuery, and Google Cloud Storage.

    Why it matters: The post explains how training on live satellite data and station observations changes resolution and update frequency, with precipitation accuracy gains reported against named baselines.

Sep 2

Sep 2Wed
  1. ARC PrizeAI score77

    OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private

    AIOpenAI's GPT-6 Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K under the Standard harness, and 99.9% for $19K under the Provider Adapter harness. The authors say Astra used fewer actions than the human baseline on 96.0% of levels, and they note it is not claimed to be AGI.

    Why it matters: The report pairs benchmark scores with replays of the model's notation and tool use, showing how it solved unfamiliar environments rather than only that it did.

  2. NVIDIA · new models on Hugging FaceAI score67

    NVIDIA releases Nemotron-3-Labs-Ultra-Math-RL for mathematical proof reasoning

    AINVIDIA has published Nemotron-3-Labs-Ultra-Math-RL on Hugging Face, a 550B total, 55B active parameter model for solving difficult math problems and identifying proof mistakes. The model is part of an ensemble that reached gold-medal level at the International Mathematical Olympiad 2026, and it is available for commercial and non-commercial use under the OpenMDW-1.1 license. Deployment is designed for NVIDIA Blackwell or Hopper GPUs, with a recommended minimum of 8× B200 on a single node and a context length of up to 1M tokens.

    Why it matters: The release details the model's math-proof role, its 550B total and 55B active parameters, and its vLLM deployment requirements for teams weighing adoption.

Sep 1

Sep 1Tue
  1. Anthropic · YouTubeAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

  2. Anthropic · YouTubeAI score72

    Anthropic releases Claude Fable 5.1 for complex, long-running tasks

    AIAnthropic has released Claude Fable 5.1, an upgrade to its most capable model class, and says it is available everywhere today. The company reports that at lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost. It is described as strong at complex multi-step work, such as long proofs and contracts with hundreds of cross-references, and at fixing root causes in software issues.

    Why it matters: The source reports cost and effort-level tradeoffs for long-running tasks, helping readers judge whether the upgrade changes their workloads or budgets.

  3. Ai2 · new models on Hugging FaceAI score22

    Ai2 Releases Supplemental ACE2S-SHiELD+ Ablation Checkpoints on Hugging Face

    AIAi2 has published supplemental checkpoints for its ACE2S-SHiELD+ climate model on Hugging Face, covering four ablation configurations that test random CO2 data and energy conservation. Each configuration includes two random-seed models, and the repository recommends the main ACE2S-SHiELD+ checkpoint for most uses. The checkpoints are licensed under Apache 2.0 for research and educational use.

  4. Microsoft AI BlogAI score34

    Microsoft Publishes 2026 Responsible AI Transparency Report on Governance and Agentic AI Risks

    AIMicrosoft published its 2026 Responsible AI Transparency Report, its third annual edition, detailing updates to its governance and risk management. The company re-engineered its Responsible AI Standard to adapt to evolving technical risks and regulatory requirements, and is extending controls such as agent identities, tool permissions, and action monitoring to agentic AI systems.

  5. Ai2 (Allen Institute for AI)AI score56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    AIAi2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

  6. OpenBMB (MiniCPM) · new models on Hugging FaceAI score49

    MiniCPM5-2B-Midtrain: OpenBMB releases mid-training checkpoint of 2B-class model

    AIOpenBMB released MiniCPM5-2B-Midtrain, a BF16 mid-training checkpoint taken before SFT in the MiniCPM5-2B series, on Hugging Face and ModelScope. The series is a 2B dense Transformer with 2,516,756,480 total parameters and a 131,072-token context length, and the final MiniCPM5-2B reports an average score of 53.9 against 51.1 for the best larger comparison model. The release also includes GGUF, MLX, and GPTQ variants, along with the UltraData datasets.

Aug 31

Aug 31Mon
  1. Liquid AI NewsletterAI score46

    Liquid AI launches Pipette, an open-source benchmark for on-device foundation models

    AILiquid AI and Artificial Analysis released Pipette, an open-source benchmark platform for foundation models on edge devices, covering over 1,000 configurations across 30+ models. It measures five on-device metrics, including throughput, latency, context scaling, and memory use, on macOS, Windows, iOS, and Android. Liquid AI also said its updated LFM2.5 Q4_0 checkpoints, trained with Quantization-Aware Distillation, retain roughly 97% of BF16 baseline performance and suffer 73.4% less quality loss than standard post-training Q4_0 quantization.

  2. DeepSeek · new models on Hugging FaceAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    AIDeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    Why it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP Releases Core-Embed 8B for Compositional Multimodal Retrieval

    AIAlibaba NLP has released core-emb-8b, an MLLM-based multimodal embedding model that distills a reranker's compositional judgments to distinguish attribute-object bindings such as "a white plate and a black chair" versus "a black plate and a white chair." The 8B dense embedding model, built on the Qwen3-VL-based VL-Emb backbone, scores 0.666 total average on compositional benchmarks, 5.7 points above its backbone. It is part of a family that also includes 2B embedding and reranker models.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    AIAlibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.

  3. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score36

    Alibaba's core-reranker-2b Model Targets Compositional Image-Text Relevance Scoring

    AIAlibaba NLP released core-reranker-2b, a 2B-parameter multimodal relevance-scoring model built on Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image pairs. The Core-Reranker family also includes an 8B variant, and Core-Reranker-8B reports an 82.7% total average on compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, 10.7 points above Jina-Reranker. Usage details are provided in the source, including loading through the GitHub repository wrapper classes.

Aug 27

Aug 27Thu
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score65

    OpenBMB releases MiniCPM5-2B-SFT, a 2B open model with SFT-only checkpoint

    AIOpenBMB released MiniCPM5-2B-SFT, an SFT-only BF16 checkpoint taken before RL and OPD, within its MiniCPM5-2B series. The model is a 2B dense Transformer built for on-device and local deployment, with 131,072-token context and the same training recipe as the final release.

    Why it matters: The source gives concrete benchmark averages against same-size and larger models, plus released training data and multiple deployment formats, useful for judging a compact on-device model.

  2. OpenBMB (MiniCPM) · new models on Hugging FaceAI score57

    OpenBMB releases MiniCPM5-2B, a 2B-class open model with open training data

    AIOpenBMB released MiniCPM5-2B, a dense 2B Transformer for on-device and resource-constrained deployment, alongside its training datasets. The source reports a 53.9 average across its comparison set and strong results in coding, math, long-context, tool use, and agentic tasks. This page is the pre-training base checkpoint, with BF16 weights and GGUF, MLX, GPTQ, and LiteRT-LM variants listed separately.

  3. Tencent · new models on Hugging FaceAI score80

    Tencent open-sources Hy4 preview, a 770B-parameter MoE model

    AITencent's Hy Team released Hy4 preview, a Mixture-of-Experts model with 770B total parameters and 49B activated per token, with a 1M context length. Hugging Face hosts the Instruct model and an FP8 quantized version under the Apache License 2.0, with vLLM and SGLang deployment instructions provided.

    Why it matters: The model card gives architecture, activated parameters, and vLLM and SGLang deployment recipes, useful for judging whether the release fits your serving setup.

  4. Qwen · new models on Hugging FaceAI score62

    Qwen-Drive-1.0 releases open weights for driving VQA, perception, and planning

    AIQwen has published Qwen-Drive-1.0-4B on Hugging Face, a vision-language model for autonomous driving built on Qwen3.5-4B. The release includes a BEV perception head and two Planning Experts, planner-sft and planner-rl, with code and an inference example in the linked GitHub repository.

    Why it matters: The source gives concrete benchmark results and a runnable setup, letting readers judge how a driving VLM with planning and perception heads compares with existing systems.

Aug 26

Aug 26Wed
  1. Amazon ScienceAI score46

    Dependence-Aware Aggregation Improves LLM-as-a-Judge Accuracy by 9% to 14%

    AIAmazon researchers proposed a dependence-aware method for aggregating LLM judges' votes, using an Ising model to account for correlated errors among judges. The approach outperformed a weighted majority-vote baseline by 9% to 14% on standard metrics across three binary tasks, including relevance classification, where it reached 0.912 accuracy versus 0.820. The method is unsupervised, learning from judge outputs without human reference labels.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Fireworks AI BlogAI score46

    DeepSeek V4 Pro Solves Security Tasks at Half the Cost Per Success

    AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.