Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Sep 10

Sep 10Thu
  1. DeepSeekOfficialAI score38

    DeepSeek unveils 552B MoE model with asymmetric encoder-decoder design

    AIDeepSeek has introduced a 552B-parameter MoE model built on a new Causal Encoder–Decoder architecture, activating 8B parameters for input and 16B for output. The company says new pre-training methods and larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.

    Image from @deepseek_ai's post
  2. DeepSeek API NewsOfficialAI score72

    DeepSeek releases V4.1-Flash with native multimodal support and API updates

    AIDeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native multimodal visual understanding. The API now serves it under the model name deepseek-flash, while V4 Flash and V4 Flash Vision Exp were retired and routed to V4.1 Flash. API prices were reduced with the release, and V4 Pro remains available after September 14, 2026.

    Why it matters: The release lists benchmark results alongside API model-name changes and retirements, so developers can check both capability claims and migration steps.

Sep 9

Sep 9Wed
  1. DeepSeek · new models on Hugging FaceOfficialAI score78

    DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

    AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

    Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

  2. Fireworks AI BlogOfficialAI score60

    Genspark's Gen-1 Slides matches Opus 5 decks at about one-tenth the cost per deck

    AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.

    Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.

  3. TinkerOfficialAI score28

    Tinker and OpenResearch automate auditing of self-distillation methods

    AITinker says it and OpenResearch let agents test dozens of competing published post-training methods automatically, with compute cost forecast to within a dollar. The main post cites a grant-supported effort, while the quoted alphaXiv post says agents reproduced SDFT's continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds.

  4. METROfficialAI score31

    METR plans investigation into AI misalignment incidents and propensities

    AIMETR says its planned investigation will cover all questions raised in its recently updated post on how independent researchers could study AI propensities after misalignment incidents. The post defines misalignment incidents as cases where an AI agent autonomously took sophisticated, sustained actions violating human intent.

  5. Cognition Blog (Devin, Windsurf)OfficialAI score82

    Cognition's Devin factors RSA-260 using a GPU lattice siever

    AICognition's Devin agent, directed by Eric Lu, factored the 260-digit RSA-260 number using a new GPU implementation of the general number field sieve built on CADO-NFS. The author estimates the run cost about 13.5 GPU-years, roughly $400k at market prices, and projects RSA-1024 factoring at around $30M, while RSA-2048 is not meaningfully affected.

    Why it matters: The source gives a full cost breakdown and scaling estimates for RSA factoring on GPUs, showing how far the cost of breaking RSA-1024 has fallen.

Sep 8

Sep 8Tue
  1. Google Developers BlogOfficialAI score36

    Google Developers Blog outlines behavioral evals for guarding AI coding agents against regressions

    AIGoogle Developers Blog argues that teams building AI coding agents should replace end-to-end benchmark scores with behavioral evaluations that test discrete, observable actions. Examples include asking clarifying questions on underspecified prompts, running a local validator before marking a build change complete, and consulting live search for current information. The post recommends fast, deterministic unit-style checks, outcome-based LLM-as-a-judge checks for complex tasks, and batch runs that track aggregate pass rates over time.

  2. InferactOfficialAI score42

    Inferact reports open models hit 130K tokens/GPU-sec on agentic workloads

    AIInferact says months of vLLM tuning for agentic workloads, validated on SemiAnalysis's AgentX benchmark, let open-source models reach up to 130K tokens per GPU-second. The company claims this is 106 times cheaper than Opus 5 API pricing. The work is described as part of a vLLM blog post covering architecture, framework, and runtime optimizations.

  3. Noam BrownXAI score36

    OpenAI model reportedly delivers a huge step up over today's LLMs

    AINoam Brown says a plot shows the model OpenAI used is a major advance beyond today's LLMs, and that no one relied on Levent's or Tristan's prompts. He was responding to questions about how Navier-Stokes might be achieved and the attention given to those prompts. Background from Sebastien Bubeck describes coordination over Euler and Navier-Stokes results, including a disputed suggestion about Levent's authorship.

    Image from @polynoamial's post
  4. BAAIOfficialAI score34

    Robot models excel at single moves but fail chained tasks

    AIBAAI reports that robot models trained on individual skills such as grasping, placing, pulling, and opening performed poorly when asked to chain them into full tasks without extra practice. The best score was 16.7%, and some models scored zero. The post's example notes a robot may open a drawer yet get stuck on the handle.

    Image from @BAAIBeijing's post
  5. BAAIOfficialAI score46

    Top embodied models hit 98% in sim but drop sharply on hardware

    AIIn simulation, the best embodied AI models complete easy tabletop tasks about 98% of the time. On physical Franka robots, their success rate falls to 24%–72% of simulated performance, and to 13%–60% in a dual-arm setup. Models that look tied in simulation can differ by 30 points on real hardware.

    Image from @BAAIBeijing's post
  6. BAAIOfficialAI score43

    FlagEval-Robo tests 12 open-weight embodied AI models across simulation and real robots

    AIBAAI introduces FlagEval-Robo, an open dual-track evaluation suite linking simulation with real-world execution. The team post-trained and stress-tested 12 leading open-weight embodied AI models under strictly aligned conditions. The post raises whether high benchmark scores reflect physical reality, though it does not yet report specific results.

    Image from @BAAIBeijing's post
  7. Google DeepMindOfficialAI score74

    Google DeepMind launches AlphaGenome Atlas to predict 9 billion DNA variant effects

    AIGoogle DeepMind has introduced AlphaGenome Atlas, a platform with predicted molecular effects for 9 billion single-nucleotide variants in the human genome. It is free for academic research through a web portal, and the AlphaGenome Variant Impact score condenses predictions from AlphaGenome and AlphaMissense into one number for ranking variants. The source says collaborators used it to identify variants in unsolved rare disease cases and to find rare non-coding variants linked to traits.

    Why it matters: The source details how precomputed variant predictions, a single impact score, and linked feature attributions make genome-wide mutation effects searchable for researchers without coding skills.

Sep 7

Sep 7Mon
  1. Tencent HyOfficialAI score44

    Tencent Hy4 preview upgraded to cut overthinking and token use

    AITencent Hunyuan says its Hy4 preview has been upgraded to reduce long thinking and over-verification on complex tasks, which users had flagged. The company reports the same task quality with fewer turns and lower input and output tokens, confirmed by benchmark and human evaluation. The upgrade is live for all users, and Tencent says it will keep iterating based on feedback.

  2. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score45

    openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning

    AIOpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning. The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline. The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.

Sep 6

Sep 6Sun
  1. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score62

    OpenBMB releases MiniCPM5-2B, a 2B open-source model with open training data

    AIOpenBMB has released MiniCPM5-2B, a dense 2B Transformer built for on-device and resource-constrained deployment, with an average score of 53.9 in its comparison set. The release also opens the UltraData datasets behind it, including UltraX, UltraData-Code, UltraData-SFT-Agent-2609 and UltraData-RL-2609, and includes GGUF, MLX, GPTQ and DSpark variants for common runtimes.

    Why it matters: The release pairs a 2B model with open training datasets and reports per-benchmark comparisons against named same-size and larger models, letting readers check the claims directly.

Sep 5

Sep 5Sat
  1. AI at MetaOfficialAI score46

    AIRA₃ cuts GPU kernel latency 27% and reaches Kaggle gold level

    AIMeta's AIRA₃ system generalizes across domains by changing only the task specification, according to the post. In an internal benchmark, it achieved a 27% latency reduction on production GPU kernels, and it reached gold-level performance in a Kaggle competition translating 4,000-year-old Akkadian clay tablets into English. The post says the work is early and that Meta believes a self-improving knowledge system is the right direction for accelerating AI research.

  2. AI at MetaOfficialAI score38

    AIRA₃ ensemble places 8th with gold-medal results in live competition

    AIMeta's AIRA₃ entered the live competition with an ensemble of models, and the 8th-ranked gold-medal entry combined GPT 5.5 (w/ OpenCode) and Claude 4.8 (w/ ClaudeCode). Post-hoc testing found Muse Spark 1.2 (w/ MuseCode) also reached gold-medal level, while Muse Spark 1.1 (w/ OpenCode) and GLM 5.2 (w/ OpenCode) reached silver-medal level, all graded on the same private test set.

    Image from @AIatMeta's post

Sep 4

Sep 4Fri
  1. John SchulmanXAI score34

    Schulman praises metric and dataset for training models to explain behavior

    AIJohn Schulman says a metric for explanation quality, centered on counterfactual simulatability, enables hillclimbing, and praises Adam et al. for a more diverse and realistic dataset and pipeline. He notes that models can be trained to write better post-hoc explanations of their own behavior, as described in a linked thread by @a_karvonen. That thread reports training on thousands of self-explanations of in-the-wild behaviors, with generalization to held-out evals.

  2. Lewis Tunstall @ COLM 🌉XAI score60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

  3. Lewis Tunstall @ COLM 🌉XAI score22

    Research Preference Models Rank AI Research Ideas to Save Compute

    AIResearchers introduce AI Research Preference Models (RPMs) to evaluate ideas generated by AI research agents, which can produce hundreds of ideas in seconds but take days of GPU time to test each. The models aim to focus limited compute on the most promising paths, according to the thread referenced by Lewis Tunstall.

  4. Tencent · new models on Hugging FaceOfficialAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

  5. Tencent · new models on Hugging FaceOfficialAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

Sep 3

Sep 3Thu
  1. TinkerOfficialAI score51

    Bespoke Labs post-trains Inkling on one code repo and reports broader coding gains

    AIBespoke Labs post-trained the Inkling base model on a single GitHub repository using supervised fine-tuning and GRPO reinforcement learning. The post reports a 57-point improvement on the held-out fontTools evaluation over the base model, along with gains on Terminal-Bench 2.1 and SWE-bench Lite. It also says the post-trained model uses about 40% fewer tokens.

    Image from @tinkerapi's post
  2. Google DeepMind · The KeywordOfficialAI score72

    Google DeepMind releases WeatherNext 3, a global weather model with hourly satellite-based forecasts

    AIGoogle DeepMind and Google Research introduced WeatherNext 3, which generates hourly global forecasts at up to 5-kilometer resolution using live geostationary satellite data. The company reports that precipitation forecasts improved by up to 60% against IMERG in medium-range evaluations, and that longer-range precipitation forecasts are up to 50% more accurate. The model is now available across Search, Gemini, Google Maps, Google Maps Platform Weather API, Google Earth Engine, BigQuery, and Google Cloud Storage.

    Why it matters: The post explains how training on live satellite data and station observations changes resolution and update frequency, with precipitation accuracy gains reported against named baselines.

Sep 2

Sep 2Wed
  1. TinkerOfficialAI score44

    Lightning Rod's new work shows scoring rules reshape LLM forecaster profiles

    AILightning Rod, working with Philip Tetlock and Ville Satopää, post-trained five versions of the same LLM that differed only in the scoring rule used as the RL reward. The versions reached similar aggregate scores but had very different bias, information, and noise (BIN) profiles, so a good Brier score alone does not show whether a forecaster can distinguish likely from unlikely events.

  2. ARC PrizeOfficialAI score77

    OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private

    AIOpenAI's GPT-6 Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K under the Standard harness, and 99.9% for $19K under the Provider Adapter harness. The authors say Astra used fewer actions than the human baseline on 96.0% of levels, and they note it is not claimed to be AGI.

    Why it matters: The report pairs benchmark scores with replays of the model's notation and tool use, showing how it solved unfamiliar environments rather than only that it did.

  3. The Register · AINewsAI score39

    AI Models Misidentify Mushrooms in Test, Sometimes Calling Deadly Species Edible

    AIPiotr Migdał tested 16 AI models on 1,040 mushroom photos covering 55 species, and the best, Gemini-3.8-flash, was correct on its first guess only 65 percent of the time. Dangerous mistakes were common, with the death cap called edible 16 percent of the time, and Qwen3.8-27b wrongly labeled poisonous mushrooms edible 36 percent of the time. Migdał warns users not to eat any mushroom because an AI says it is safe.

  4. NVIDIA · new models on Hugging FaceOfficialAI score67

    NVIDIA releases Nemotron-3-Labs-Ultra-Math-RL for mathematical proof reasoning

    AINVIDIA has published Nemotron-3-Labs-Ultra-Math-RL on Hugging Face, a 550B total, 55B active parameter model for solving difficult math problems and identifying proof mistakes. The model is part of an ensemble that reached gold-medal level at the International Mathematical Olympiad 2026, and it is available for commercial and non-commercial use under the OpenMDW-1.1 license. Deployment is designed for NVIDIA Blackwell or Hopper GPUs, with a recommended minimum of 8× B200 on a single node and a context length of up to 1M tokens.

    Why it matters: The release details the model's math-proof role, its 550B total and 55B active parameters, and its vLLM deployment requirements for teams weighing adoption.