Updated
#Eval/Benchmark
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
Oct 9
ARC Prize@arcprizeAI score22
OpenBMB@OpenBMBAI score28MiniCPM5-2B runs at 37 tok/s on iPhone Air
AIOpenBMB reports that its MiniCPM5-2B model runs at 37 tokens per second on an iPhone Air. The post presents this as evidence that small open multimodal models can run on mobile devices without a cloud GPU, with NobodyWho noting the model is available in its Chat app.
MarkTechPostAI score44 Underdog Releases Saluki 27B, a 2-Bit Qwen3.8-27B That Beats the Original at Tool Calling
AIUnderdog has released Saluki 27B under Apache 2.0, a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB, versus 54 GB for the full BF16 model. On Underdog Bench, Saluki scores 88 against 84 for the full model, and it raises parallel tool-call accuracy to 42 from 35. It runs on stock llama.cpp, but math and reasoning drop sharply, with AIME 2025 at 79.2 versus 96.7.
IThome · AIAI score46 JetBrains Releases Mellum2.1 Coding Model With Near-Double Qwen3.5-9B Throughput
AIJetBrains released Mellum2.1, a 12B mixture-of-experts coding model with 2.5B active parameters under Apache 2.0, emphasizing agentic programming. Under high load, its inference throughput in tokens is nearly twice that of Qwen3.5-9B in JetBrains' comparison, and multi-token prediction (MTP) speeds single-request responses by about 1.6x. The model is available on Hugging Face for local or private-infrastructure deployment, with GGUF and vLLM MTP support announced for later.
IThome · AIAI score55 Odyssey-3 world model scores 66.1 on Physics-IQ Verified benchmark
AIOdyssey announced the Odyssey-3 series of foundation world models, with Odyssey-3 Pro scoring 66.1 on the Physics-IQ Verified video-to-video benchmark, the highest recorded on that leaderboard. The series includes a standard version balancing physical accuracy and generation cost, and a Pro version with stronger physics prediction. The preview supports first-person and third-person navigation and lets users move the camera, take actions, or trigger events while the model predicts environmental changes in real time.
vLLM@vllm_projectAI score42vLLM Semantic Router team releases Decision 2.0 multi-question classification models
AIThe vLLM Semantic Router team has released Decision 2.0, which answers multiple questions about one input in a single forward pass and outputs per-option probabilities. The post presents this as useful for routing and classification. A quoted post from Xunzhuo Liu says Decision 2.0 includes six open decision models ranging from 0.6B to 27B parameters, each topping same-size open models on the Jev Decision Index 0.3.
Nace AI@NaceAIAI score22NDI 1.0 document processing model launches for coding agents at 90% lower cost
AINACE introduces NDI 1.0, a document processing model for coding agents that it says is 90% cheaper and ranks first on the Parse Index. The company says it offers native MCP, SDK, and CLI integration for Claude Code, Codex, Hermes, OpenClaw, and PI, and supports 50 languages. NACE also states the model was trained on over 15M financial files and offers $25 in free API credits to developers.

Oct 8
PandailyAI score57 Shanghai AI Lab Open-Sources Intern-Decision Small Models for Structured Decisions
AIShanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.
QbitAIAI score52 Claude Haiku 5.5 launches with higher benchmark scores and new migration requirements
AIAnthropic released Claude Haiku 5.5, which the article says outperforms DeepSeek V4.1 Flash and GLM-5.3-Flash on official benchmarks and matches GPT-6 Luna on price. On OSWorld 2.1, its Low effort tier scores 42.0% at $0.07 per task, versus 15.7% at $1.45 for Haiku 4.5 at Max. Migrating from Haiku 4.5 requires changes to thinking configuration, sampling parameters, assistant prefill, and the computer-use tool version.
SiliconANGLE · AIAI score63 Anthropic releases Claude Haiku 5.5 small model and halves Sonnet 5.5 cache read prices
AIAnthropic released Claude Haiku 5.5, a small model priced at about a quarter of Haiku 4.5's running cost. Anthropic also halved Sonnet 5.5 cache read prices from 20 cents to 10 cents per million tokens.
Xiaomi MiMoAI score44 Xiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support
AIXiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.
MiniMax (official)@MiniMax_AIAI score34MiniMax H3 nears closed-source SOTA on physics in open video world models
AIMiniMax says its open-source H3 model is almost on par with closed-source state-of-the-art video world models on physics. The claim is supported by a quoted benchmark, World Models' Last Exam in Physics, where eight leading models scored at most 57.76/100 across 40 physics tasks, and free-fall videos averaged only 26.61/100 on composite scores.
Artificial Analysis@ArtificialAnlysAI score38Grok Imagine Video 1.5 Lite leads in architecture, consumer, and knowledge-work use cases
AIArtificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier in Architecture & Real Estate, Consumer, and Productivity & Knowledge Work use cases. It sits furthest from the frontier in Live-Action Film and Frontier use cases. Against Grok Imagine Video 1.5, Lite matches it in Social Media & Creator Content and trails it on the other nine use cases.

Artificial Analysis@ArtificialAnlysAI score31Grok Imagine Video 1.5 Lite leads on quality and speed benchmark
AIAmong 12 models on AA-Video-T2V-Silent v2.0, Grok Imagine Video 1.5 Lite is the only one that is both fastest and highest quality, with no model beating it on both measures. It generates a 10-second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher but takes 94 seconds for a 5-second clip, while Vidu Q3 Turbo is 9 seconds faster on a 5-second 720p clip yet scores well below it.

Artificial Analysis@ArtificialAnlysAI score46Grok Imagine Video 1.5 Lite outranks Veo 3.1 at a third of the cost
AIGrok Imagine Video 1.5 Lite ranks #17 on AA-Video-T2V v2.0, two places above Google's Veo 3.1. At 1080p with audio, it costs $0.14 per second versus $0.40 per second for Veo 3.1. Compared with Grok Imagine Video 1.5, Lite is 44% cheaper at 1080p but ranks six places lower.

Artificial Analysis@ArtificialAnlysAI score42Grok Imagine Video 1.5 Lite ranks #17 in video arena at lower cost
AISpaceXAI's Grok Imagine Video 1.5 Lite ranks #17 on both AA-Video-T2V v2.0 leaderboards, ahead of Google's Veo 3.1 at about a third of its price. It is the fastest model at its quality level in Artificial Analysis benchmarks, with a median of 60.5 seconds for a 10-second 1080p clip, and it costs $0.14 per second at 1080p, 56% of Grok Imagine Video 1.5's $0.25 per second.

Epoch AI@EpochAIResearchAI score40OpenAI halves GPT-6.1 Sol cached-input price, speeds long prompts
AIEpoch AI reports that OpenAI halved GPT-6.1 Sol's cached-input price compared with GPT-6 Sol. Its measurements also show the model handles long prompts faster. The post raises, but does not confirm, a new architecture for GPT-6.1 Sol.

Arena.ai@arenaAI score31Claude Haiku 5.5 scores 1587 on Arena, up 257 points from Haiku 4.5
AIAnthropic's Claude Haiku 5.5 reached 1587 points on the Arena WebDev leaderboard, a 257-point gain over Claude Haiku 4.5's 1330. The post says it also improved significantly across all categories.

Artificial Analysis@ArtificialAnlysAI score38GPT-6 Sol (Daybreak Blue) tops Artificial Analysis Cyber Index
AIGPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model and ranks #1 on the Index. Compared with the publicly available GPT-6 Sol, it shows its largest gains on CyberGym-E2E, the benchmark where the most safety refusals are observed.

Sherwin Wu@sherwinwuPickAI score60Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%
AISherwin Wu, an OpenAI employee, says the updated Harvey LAB-AA v1.1 benchmark, announced by Artificial Analysis with Harvey, is more useful than the original LAB results. The new Hallucination-Gated All-Pass Rate credits a task only when every rubric criterion passes and no material hallucination appears. Grok 4.7 (xhigh) leads at 9.4%, while GPT-6 Astra (max) at 8.6% has very few material hallucinations.
Why it matters: The update adds a hallucination gate to a legal benchmark, showing that models with high all-pass rates can rank much lower once material errors count.
Boris Power@BorisMPowerAI score46OpenAI's GPT-6.1-Sol leads new Arena Alignment Index for agents
AIThe Arena Alignment Index, built from over 90K real-world agent sessions across 27 models, ranks OpenAI's GPT-6.1-Sol first with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7. GPT-6.1-Sol also posted the lowest observed rates across the index's three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. The index's authors report that newer models consistently outperform their predecessors across all four labs, suggesting broad progress in agent safety.
Alexander Doria@DorialexanderAI score46LightOnOCR-3 claims state-of-the-art OCR performance under 1B parameters
AILightOn has released LightOnOCR-3, a family of OCR models in 0.8B and 4B versions that it says lead benchmarks including OlmOCR-Bench and ParseBench, with the 0.8B model positioned as the sub-1B option. The models recognize text, handwriting, images, charts and document structure in one pass, process documents up to twice as fast as LightOnOCR-2, and are released under the Apache 2.0 license.

🚨 AI News | TestingCatalog@testingcatalogAI score49Odyssey launches Odyssey-3 world model with public research preview
AIOdyssey has launched Odyssey-3, its most powerful foundation world model, with a public research preview. Odyssey-3 Pro scored 66.1 on Physics-IQ Verified video-to-video with best-of-8 sampling, the highest reported result. The model generates environments from prompts and predicts changes in real time as users move through scenes.

elvis@omarsar0AI score42Odyssey-3 Pro tops Physics-IQ Verified and shows robotic error recovery.
AIOdyssey released Odyssey-3 Pro, which sets a new top score on Physics-IQ Verified, a benchmark where models continue videos of real physics experiments. In robotics, a robot arm with tens of hours of demonstrations recovered from a missed grasp, a behavior absent from those demos.

Santiago@svpinoAI score46Odyssey 3 Pro world model tops Physics-IQ and goes live
AIOdyssey 3 Pro, a world model, is now live as a research preview and ranks first on the Physics-IQ Verified video-to-video benchmark. The post says it can learn from visual observations and map that knowledge to physical controls for robots, cars, video games, and drones. Odyssey-3, the model launched alongside it, is described as free to try.

Odyssey@odysseymlAI score22Odyssey-3 Pro sets new Physics-IQ video-to-video benchmark record
AIOdyssey-3 Pro achieved a score of 66.1 on Physics-IQ Verified's video-to-video benchmark, the highest reported score so far. Physics-IQ evaluates physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.

Odyssey@odysseymlAI score40Odyssey launches Odyssey-3, a foundation world model setting Physics-IQ state of the art
AIOdyssey has launched Odyssey-3, which it describes as its most powerful foundation world model yet. The company says the model sets a new state of the art on Physics-IQ and can power robots, train AIs, and generate interactive experiences, and it is available to try for free today.

Aravind Srinivas@AravSrinivasAI score22Perplexity Decider ranks first on DecisionBench at lowest cost
AIPerplexity's Decider V1.1 ranked first on DecisionBench while also having the lowest cost, according to a post highlighting the result. The benchmark results cited include 949 shared text cases, 93.9% accuracy, a 534 ms median latency, and $0.016 per 1k decisions.
MarkTechPostAI score58 JetBrains releases Mellum2.1, a 12B MoE open model for coding agents
AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.
Arena.ai@arenaAI score55Arena raises $200M Series B and launches Alignment Index for AI agents
AIArena announced a $200M Series B at a $3.1B valuation and released its Alignment Index, a benchmark measuring agent safety and alignment. The index is built from 90K+ real-world agent sessions across 27 models and tracks Unauthorized Action, False Attribution, and Deceptive Completion. OpenAI's GPT-6.1-Sol leads with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7.

Sophia Yang@sophiamyangPickAI score60LeiphoneAI score62 Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding
AIAnthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.
Aravind Srinivas@AravSrinivasAI score40Perplexity open-weights clinical decision model beats Jev on cost and accuracy
AIPerplexity's pplx-decider v1.1 scored 643 of 669 on clinical decisions versus 628 for Jev, according to a quoted post. The quoted post says it costs 42% less per decision at about the same speed and that its weights can be downloaded and run inside hospitals.
OpenRouter@OpenRouterAI score29StepFun's Step 5 Preview beats Kimi K3 and GLM-5.3 on six benchmarks
AIStepFun's Step 5 Preview outperforms Kimi K3 and GLM-5.3 on six of eight benchmarks in its own evaluation. Reported scores include 67.7 on DeepSWE, 80.5 on ProgramBench, 49.0 on StepCodeBench, and 66.4 on FrontierFinance.

Understanding AI (Timothy B. Lee)AI score67 TypeSafe AI's Jev returns probabilities over fixed answers instead of text
AITypeSafe AI released Jev, a model that answers yes/no, multiple-choice, or rating questions by outputting the estimated probability of each option. The author notes this design lets the model be served faster and more cheaply than LLMs and fits ordinary if-statement logic, and says he used it to flag spam comments on his blog in place of Gemini 3 Flash.
OpenRouter · New modelsAI score54 StepFun releases Step 5 Preview, a 600B-parameter agentic model
AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.
JetBrains AI BlogPickAI score62 JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning
AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.
Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.
The DecoderAI score72 Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna
AIAnthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.
Latent SpaceAI score73 Claude Haiku 5.5 launches at GPT-6 Luna pricing with 1M context
AIAnthropic released Claude Haiku 5.5, priced the same as OpenAI's GPT-6 Luna, with a 1M-token context window. Artificial Analysis scored it 43 on its Intelligence Index, slightly ahead of GPT-6 Luna at 38, but it uses about 3x more output tokens at max effort.
MarkTechPostAI score65 Perplexity releases pplx-embed-v2-late, a 0.6B edge model and 9B model
AIPerplexity has released pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models in 0.6B and 9B sizes that retrieve text, images and rendered PDF pages in a shared embedding space. Both are available on Hugging Face under the MIT license, while a hosted API endpoint is planned but not yet live.
