Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri7 items
  1. MarkTechPostAI score44

    Underdog Releases Saluki 27B, a 2-Bit Qwen3.8-27B That Beats the Original at Tool Calling

    AIUnderdog has released Saluki 27B under Apache 2.0, a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB, versus 54 GB for the full BF16 model. On Underdog Bench, Saluki scores 88 against 84 for the full model, and it raises parallel tool-call accuracy to 42 from 35. It runs on stock llama.cpp, but math and reasoning drop sharply, with AIME 2025 at 79.2 versus 96.7.

  2. IThome · AIAI score46

    JetBrains Releases Mellum2.1 Coding Model With Near-Double Qwen3.5-9B Throughput

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts coding model with 2.5B active parameters under Apache 2.0, emphasizing agentic programming. Under high load, its inference throughput in tokens is nearly twice that of Qwen3.5-9B in JetBrains' comparison, and multi-token prediction (MTP) speeds single-request responses by about 1.6x. The model is available on Hugging Face for local or private-infrastructure deployment, with GGUF and vLLM MTP support announced for later.

  3. IThome · AIAI score55

    Odyssey-3 world model scores 66.1 on Physics-IQ Verified benchmark

    AIOdyssey announced the Odyssey-3 series of foundation world models, with Odyssey-3 Pro scoring 66.1 on the Physics-IQ Verified video-to-video benchmark, the highest recorded on that leaderboard. The series includes a standard version balancing physical accuracy and generation cost, and a Pro version with stronger physics prediction. The preview supports first-person and third-person navigation and lets users move the camera, take actions, or trigger events while the model predicts environmental changes in real time.

  4. vLLMAI score42

    vLLM Semantic Router team releases Decision 2.0 multi-question classification models

    AIThe vLLM Semantic Router team has released Decision 2.0, which answers multiple questions about one input in a single forward pass and outputs per-option probabilities. The post presents this as useful for routing and classification. A quoted post from Xunzhuo Liu says Decision 2.0 includes six open decision models ranging from 0.6B to 27B parameters, each topping same-size open models on the Jev Decision Index 0.3.

  5. Nace AIAI score22

    NDI 1.0 document processing model launches for coding agents at 90% lower cost

    AINACE introduces NDI 1.0, a document processing model for coding agents that it says is 90% cheaper and ranks first on the Parse Index. The company says it offers native MCP, SDK, and CLI integration for Claude Code, Codex, Hermes, OpenClaw, and PI, and supports 50 languages. NACE also states the model was trained on over 15M financial files and offers $25 in free API credits to developers.

    Video from @NaceAI's post

Oct 8

Oct 8Thu
  1. PandailyAI score57

    Shanghai AI Lab Open-Sources Intern-Decision Small Models for Structured Decisions

    AIShanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.

  2. QbitAIAI score52

    Claude Haiku 5.5 launches with higher benchmark scores and new migration requirements

    AIAnthropic released Claude Haiku 5.5, which the article says outperforms DeepSeek V4.1 Flash and GLM-5.3-Flash on official benchmarks and matches GPT-6 Luna on price. On OSWorld 2.1, its Low effort tier scores 42.0% at $0.07 per task, versus 15.7% at $1.45 for Haiku 4.5 at Max. Migrating from Haiku 4.5 requires changes to thinking configuration, sampling parameters, assistant prefill, and the computer-use tool version.

  3. Xiaomi MiMoAI score44

    Xiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support

    AIXiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.

  4. MiniMax (official)AI score34

    MiniMax H3 nears closed-source SOTA on physics in open video world models

    AIMiniMax says its open-source H3 model is almost on par with closed-source state-of-the-art video world models on physics. The claim is supported by a quoted benchmark, World Models' Last Exam in Physics, where eight leading models scored at most 57.76/100 across 40 physics tasks, and free-fall videos averaged only 26.61/100 on composite scores.

  5. Artificial AnalysisAI score38

    Grok Imagine Video 1.5 Lite leads in architecture, consumer, and knowledge-work use cases

    AIArtificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier in Architecture & Real Estate, Consumer, and Productivity & Knowledge Work use cases. It sits furthest from the frontier in Live-Action Film and Frontier use cases. Against Grok Imagine Video 1.5, Lite matches it in Social Media & Creator Content and trails it on the other nine use cases.

    Image from @ArtificialAnlys's post
  6. Artificial AnalysisAI score31

    Grok Imagine Video 1.5 Lite leads on quality and speed benchmark

    AIAmong 12 models on AA-Video-T2V-Silent v2.0, Grok Imagine Video 1.5 Lite is the only one that is both fastest and highest quality, with no model beating it on both measures. It generates a 10-second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher but takes 94 seconds for a 5-second clip, while Vidu Q3 Turbo is 9 seconds faster on a 5-second 720p clip yet scores well below it.

    Image from @ArtificialAnlys's post
  7. Artificial AnalysisAI score42

    Grok Imagine Video 1.5 Lite ranks #17 in video arena at lower cost

    AISpaceXAI's Grok Imagine Video 1.5 Lite ranks #17 on both AA-Video-T2V v2.0 leaderboards, ahead of Google's Veo 3.1 at about a third of its price. It is the fastest model at its quality level in Artificial Analysis benchmarks, with a median of 60.5 seconds for a 10-second 1080p clip, and it costs $0.14 per second at 1080p, 56% of Grok Imagine Video 1.5's $0.25 per second.

    Video from @ArtificialAnlys's post
  8. Sherwin WuAI score60

    Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%

    AISherwin Wu, an OpenAI employee, says the updated Harvey LAB-AA v1.1 benchmark, announced by Artificial Analysis with Harvey, is more useful than the original LAB results. The new Hallucination-Gated All-Pass Rate credits a task only when every rubric criterion passes and no material hallucination appears. Grok 4.7 (xhigh) leads at 9.4%, while GPT-6 Astra (max) at 8.6% has very few material hallucinations.

    Why it matters: The update adds a hallucination gate to a legal benchmark, showing that models with high all-pass rates can rank much lower once material errors count.

  9. Boris PowerAI score46

    OpenAI's GPT-6.1-Sol leads new Arena Alignment Index for agents

    AIThe Arena Alignment Index, built from over 90K real-world agent sessions across 27 models, ranks OpenAI's GPT-6.1-Sol first with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7. GPT-6.1-Sol also posted the lowest observed rates across the index's three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. The index's authors report that newer models consistently outperform their predecessors across all four labs, suggesting broad progress in agent safety.

  10. Alexander DoriaAI score46

    LightOnOCR-3 claims state-of-the-art OCR performance under 1B parameters

    AILightOn has released LightOnOCR-3, a family of OCR models in 0.8B and 4B versions that it says lead benchmarks including OlmOCR-Bench and ParseBench, with the 0.8B model positioned as the sub-1B option. The models recognize text, handwriting, images, charts and document structure in one pass, process documents up to twice as fast as LightOnOCR-2, and are released under the Apache 2.0 license.

    Image from @Dorialexander's post
  11. 🚨 AI News | TestingCatalogAI score49

    Odyssey launches Odyssey-3 world model with public research preview

    AIOdyssey has launched Odyssey-3, its most powerful foundation world model, with a public research preview. Odyssey-3 Pro scored 66.1 on Physics-IQ Verified video-to-video with best-of-8 sampling, the highest reported result. The model generates environments from prompts and predicts changes in real time as users move through scenes.

    Image from @testingcatalog's post
  12. SantiagoAI score46

    Odyssey 3 Pro world model tops Physics-IQ and goes live

    AIOdyssey 3 Pro, a world model, is now live as a research preview and ranks first on the Physics-IQ Verified video-to-video benchmark. The post says it can learn from visual observations and map that knowledge to physical controls for robots, cars, video games, and drones. Odyssey-3, the model launched alongside it, is described as free to try.

    Image from @svpino's post
  13. MarkTechPostAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  14. Arena.aiAI score55

    Arena raises $200M Series B and launches Alignment Index for AI agents

    AIArena announced a $200M Series B at a $3.1B valuation and released its Alignment Index, a benchmark measuring agent safety and alignment. The index is built from 90K+ real-world agent sessions across 27 models and tracks Unauthorized Action, False Attribution, and Deceptive Completion. OpenAI's GPT-6.1-Sol leads with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7.

    Video from @arena's post
  15. LeiphoneAI score62

    Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding

    AIAnthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.

  16. Understanding AI (Timothy B. Lee)AI score67

    TypeSafe AI's Jev returns probabilities over fixed answers instead of text

    AITypeSafe AI released Jev, a model that answers yes/no, multiple-choice, or rating questions by outputting the estimated probability of each option. The author notes this design lets the model be served faster and more cheaply than LLMs and fits ordinary if-statement logic, and says he used it to flag spam comments on his blog in place of Gemini 3 Flash.

  17. OpenRouter · New modelsAI score54

    StepFun releases Step 5 Preview, a 600B-parameter agentic model

    AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.

  18. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  19. The DecoderAI score72

    Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna

    AIAnthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.