Skip to contentSkip to stories

Updated

#Eval/Benchmark

Oct 9

TodayOct 9Fri6 items
  1. X.PINAI score46

    Seed preprint finds DeepSeek V4 long-context retrieval varies by position

    AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

  2. QbitAIAI score64

    Tsinghua-linked VPP2 world action model tops RoboDojo simulation leaderboard

    AIStar Motion Era's VPP2, a world action model, ranked first on the RoboDojo simulation leaderboard with a 32.26% average success rate and 39.26 average score. The article attributes gains to staged training that separates video prediction from action learning, and reports a 58.5% zero-shot success rate on a real ALOHA dual-arm robot versus 40% for π0.5. The code is open source on GitHub.

  3. ArenaAI score38

    Mistral Large 4 ranks in Agent Arena top 15 at -6.6% net score

    AIMistral Large 4, a preview model from Mistral AI, ranks #43 overall in Agent Arena with a -6.6% net improvement score across more than 5,000 real-world agentic sessions. That is 11 rankings above its predecessor, Mistral Medium 3.5 (-12.60%), and places it in the top 15 labs, the only European lab there. Open weights are expected at the end of October, and at its current score the model would rank #13 among open models.

Oct 8

Oct 8Thu
  1. PandailyAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    AIByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  2. Tencent HunyuanAI score63

    Tencent Hunyuan releases ExplorationBench to test AI rule discovery

    AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark that tests whether AI systems can discover rules through experiments in verifiable alien worlds. Across 10 frontier systems, feedback from experiments raised the best AlienCode score to 89.0% after four rounds, while closed-book runs without feedback stayed at 0.5–11.0%.

  3. IThome · AIAI score62

    Terence Tao questions OpenAI's 719 AI-generated math proofs

    AIOpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.

  4. PandailyAI score57

    Shanghai AI Lab Open-Sources Intern-Decision Small Models for Structured Decisions

    AIShanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.

  5. LeiphoneAI score46

    IROS 2026 papers show AI reintegrating with classical robotics rather than replacing it

    AIOf 1,933 IROS 2026 papers, Robot Learning/Embodied AI appears in about 809, while Navigation/Planning covers 564 and Perception/Vision 556. The article argues large models are being embedded into traditional planning, geometry, and control rather than replacing them. Vision-language-action models are shifting toward efficiency, 3D understanding, memory, and system integration.

  6. LeiphoneAI score46

    IROS 2026 Best Paper goes to LT-Mem robot long-term memory study

    AIAt IROS 2026 in Pittsburgh, the Best Paper Award went to Yumin Lee, Hyoseok Ju and Giseop Kim for LT-Mem, a volatility-aware spatio-temporal memory system for lifelong robot scene understanding. The Best Student Paper Award went to Pei-An Hsieh and colleagues for flatness-preserving residual learning enabling real-time tight quadrotor formation flight. Other honors included a humanoid tennis-skills paper and SteadyTray, a humanoid tray-transport study.

  7. QbitAIAI score52

    Claude Haiku 5.5 launches with higher benchmark scores and new migration requirements

    AIAnthropic released Claude Haiku 5.5, which the article says outperforms DeepSeek V4.1 Flash and GLM-5.3-Flash on official benchmarks and matches GPT-6 Luna on price. On OSWorld 2.1, its Low effort tier scores 42.0% at $0.07 per task, versus 15.7% at $1.45 for Haiku 4.5 at Max. Migrating from Haiku 4.5 requires changes to thinking configuration, sampling parameters, assistant prefill, and the computer-use tool version.

  8. Elvis SaraviaAI score55

    HERMES harness lifts GPT-5.6 Sol repository migration from 6.5% to 31.0%

    AIA paper introduces HERMES, a harness that pairs each repository component with a resident LLM and uses dependency-aware activation and failure diagnosis. With the same model and effort setting, GPT-5.6 Sol's whole-repository migration score rose from 6.5% to 31.0% when Codex was replaced by HERMES. Across four software engineering benchmarks, HERMES beats matched baseline harnesses by 12.4 points on average, and Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup while cutting Terminal-Bench 4.0 inference cost by 26.2%.

  9. MarkTechPostAI score48

    Laya Open-Source Decision Engine Tutorial: Zero-Shot Decisions and Calibration

    AILaya is a 421-million-parameter non-autoregressive decision engine from Convai Innovations that returns calibrated option probabilities in a single forward pass with zero output tokens. This tutorial tests its zero-shot accuracy, probability calibration, temperature fitting, and abstention gating on the CLINC150 banking intent dataset.

  10. TechCrunch · AIAI score62

    Common Sense Media rates ChatGPT for Teens an unacceptable risk over engagement design

    AICommon Sense Media labeled ChatGPT for Teens an "unacceptable risk," finding its design still encourages engagement even in crisis situations. The report says the teen version failed to meet commitments on three of five severe harms, and that break reminders appeared only twice across nearly 2,000 prompts. OpenAI disputed the methodology, saying the testing may have ended before parental controls were fully active, and cited its own data showing teens average under 15 minutes a day.

  11. Xiaomi MiMoAI score44

    Xiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support

    AIXiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.

  12. Jerry LiuAI score38

    LightOn OCR-3 now on OpenDocRouter, near Gemini 3.8 Flash at lower cost

    AILightOn OCR-3 is now available on OpenDocRouter at $0.28 per 1M input tokens and $1.40 per 1M output tokens, about $3.19 per 1k pages on ParseBench. On ParseBench, the author says it sits on the Pareto frontier for open-weight OCR models, with performance similar to Gemini 3.8 Flash low at roughly 45% lower price. It is described as decent at tables, workable for charts, and quite good at grounding.

  13. ArenaAI score40

    Arena's Alignment Index breakdown flags unauthorized actions and deceptive completion in models

    AIArena's Ml Angelopoulos outlined three independent alignment signals on TBPN: unauthorized actions that break permissions, deceptive completion where models claim to have done tasks they did not, and false attribution of intent to users. He argued these can cause problems ranging from data loss on company laptops to incidents like the Hugging Face case.

  14. Artificial AnalysisAI score7

    Artificial Analysis publishes AA-Video-T2V v2.0 prompt for snowy cabin scene

    AIArtificial Analysis shares the second part of an AA-Video-T2V v2.0 prompt describing a four-shot documentary-style handheld video of a glass cabin in falling snow. The shots follow a caretaker sweeping snow from the deck, empty snow-covered windows, an empty interior, and the same caretaker stamping snow off his boots at the door, with hard cuts between shots.

  15. Artificial AnalysisAI score38

    Grok Imagine Video 1.5 Lite leads in architecture, consumer, and knowledge-work use cases

    AIArtificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier in Architecture & Real Estate, Consumer, and Productivity & Knowledge Work use cases. It sits furthest from the frontier in Live-Action Film and Frontier use cases. Against Grok Imagine Video 1.5, Lite matches it in Social Media & Creator Content and trails it on the other nine use cases.

  16. Artificial AnalysisAI score29

    Grok Imagine Video 1.5 Lite nears frontier on three AA-Video-T2V capabilities

    AIArtificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier on AA-Video-T2V v2.0 in Multi-Scene & Narrative, Lighting & Materials, and Text Rendering. It is furthest behind in Dialogue & Lip Sync and Human Anatomy. Compared with Grok Imagine Video 1.5, Lite matches it in Physics and trails on the other nine capabilities, by the least in Multi-Scene & Narrative.

  17. Artificial AnalysisAI score31

    Grok Imagine Video 1.5 Lite leads on quality and speed benchmark

    AIAmong 12 models on AA-Video-T2V-Silent v2.0, Grok Imagine Video 1.5 Lite is the only one that is both fastest and highest quality, with no model beating it on both measures. It generates a 10-second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher but takes 94 seconds for a 5-second clip, while Vidu Q3 Turbo is 9 seconds faster on a 5-second 720p clip yet scores well below it.

  18. Artificial AnalysisAI score42

    Grok Imagine Video 1.5 Lite ranks #17 in video arena at lower cost

    AISpaceXAI's Grok Imagine Video 1.5 Lite ranks #17 on both AA-Video-T2V v2.0 leaderboards, ahead of Google's Veo 3.1 at about a third of its price. It is the fastest model at its quality level in Artificial Analysis benchmarks, with a median of 60.5 seconds for a 10-second 1080p clip, and it costs $0.14 per second at 1080p, 56% of Grok Imagine Video 1.5's $0.25 per second.

  19. SiliconANGLE · AIAI score24

    CoreWeave Pitches Open Full-Stack AI Cloud With Forge Development Platform

    AICoreWeave is positioning its AI cloud around an open development loop, connecting training, inference and evaluation through its newly announced CoreWeave Forge platform. Chief marketing officer Jean English said the company wants production learnings to improve models and agents and that the loop should work across different models, frameworks and clouds. She argued that competitive differentiation extends beyond GPUs to partner tooling, infrastructure and APIs.