Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 29

Sep 29Tue
  1. ModelScopeAI score44

    Intern-Decision multimodal models scale structured decisions at 0.8B–4B

    AIShanghai AI Laboratory's Intern-Decision family of 0.8B, 2B, and 4B multimodal models averages 79.38, 84.68, and 90.02 across seven decision benchmarks. Intern-Decision-4B scores 88.74, surpassing Jev while achieving better probability calibration. Reported mean latency is 33.98, 33.28, and 44.16 ms, versus 109.70 ms for Jev in the same local HF setup.

    Image from @ModelScope2022's post
  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score40

    InternLM releases AdvancedMathBench-AutoVerifier to grade natural-language math proofs

    AIInternLM's AutoVerifier, built on Qwen3_5MoeForConditionalGeneration with about 68 GiB of weights across 40 safetensors shards, evaluates natural-language mathematical proofs, explains errors, and identifies the earliest incorrect step. It serves as the automatic grader for AdvancedMathBench's ProverBench, which accepts a proof only when all eight judgments report -1. The model is a learned grader rather than a formal proof checker and can make errors.

  3. Artificial Analysis ArticlesAI score62

    Artificial Analysis open-sources AA-AgentPerf-Local for benchmarking local AI agents

    AIArtificial Analysis has open-sourced AA-AgentPerf-Local, a tool that replays recorded agent trajectories to measure inference speed on laptops and workstations. Initial results cover NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with the RTX 5090 fastest for models that fit its 32 GB. The source states the tool and leaderboard will expand to more hardware, frameworks, and models.

    Why it matters: The source gives per-system completion times and memory bandwidth figures, letting readers compare local hardware for running agentic workloads.

  4. Artificial Analysis ArticlesAI score78

    GPT-6.1 Sol replaces GPT-6 Sol with near-Astra intelligence at lower cost

    AIArtificial Analysis reports that GPT-6.1 Sol replaces GPT-6 Sol after seven days and scores 1 point below GPT-6 Astra on the Intelligence Index. At max effort it costs $0.72 per Intelligence Index task, compared with $3.26 for GPT-6 Astra and $1.05 for GPT-6 Sol. Its pricing matches GPT-6 Sol at $2/$10 per million input/output tokens, but it uses about 10-30% more output tokens.

    Why it matters: The source compares GPT-6.1 Sol against GPT-6 Sol, GPT-5.6 Sol, and GPT-6 Astra on cost per task and token use, helping readers weigh performance against price.

  5. Anthropic ResearchAI score80

    Anthropic says GLM-5.3 gives attackers cyber capabilities with weak safeguards

    AIAnthropic reports that Zhipu AI's GLM-5.3 can autonomously build end-to-end cyber exploits and is released without meaningful safeguards against misuse. In its simulated tests, attackers bypassed the model's safeguards 64% to 100% of the time using simple techniques, while the same attacks failed against safeguarded Claude models. Anthropic also cites an NIST CAISI assessment calling GLM-5.3 the most cyber-capable open-weight model released to date.

    Why it matters: The report shows how open-weight safeguards fail under simple bypasses, offering concrete test figures for judging misuse risk in released models.

Sep 28

Sep 28Mon
  1. ModelScopeAI score44

    Audio8 ASR Infinite enables unlimited-length streaming speech transcription with bounded memory

    AIAudio8 ASR Infinite transcribes Chinese and English audio of unlimited length using a rolling KV Cache that avoids accumulated drift. At a 480 ms delay, it reports 1.75 CER on AISHELL-1, 2.89 on AISHELL-4, and 3.04/6.81 WER on LibriSpeech test-clean/test-other. The preview release is under Apache 2.0, with deployment through an adapted vLLM stack.

    Video from @ModelScope2022's post
  2. Ali GhodsiAI score62

    Databricks finds Opus 5.5 cheaper and better, GPT-6 Luna 20x cheaper per task

    AIDatabricks tested recent AI models across 2,400 engineers and found Opus 5.5 offers the highest quality mid-tier performance, with about 20% lower same-task costs than Opus 4.8. The company is now encouraging Opus 5.5 as a default model for coding, and reports that GPT-6 Luna is at least 20 times cheaper per task than Opus 5.5, roughly matching Opus 4.6 on one difficult evaluation suite. The Luna findings are preliminary.

  3. ReplicateAI score28

    Pruna's P-Video-2-Pro video model now runs on Replicate

    AIReplicate has added P-Video-2-Pro, the latest video model from Pruna AI, which sits on the edge of the preference-speed and preference-price Pareto frontiers. Design Arena ranks its Quality and Speed variants tied for #2 on the Image to Video leaderboard with an Elo of 1325, with the Quality version generating in 8.0 seconds and the Speed version in 4.5 seconds.

  4. Epoch AI · The Epoch BriefAI score62

    Epoch AI finds AI cost per benchmark score falling 13× per year

    AIEpoch AI estimates that the cheapest cost of reaching a given benchmark score has fallen about 13× per year over the past five years, faster than DNA sequencing, compute, lithium batteries, or electricity. Its example: a 75% GPQA Diamond score that cost about 30 cents per question with o3 in January 2025 cost $0.0004 per question with GPT-5.6 Luna under 18 months later. The authors caution that benchmarks are imperfect proxies for market prices, and the decline rate slows over time.

    Why it matters: The source compares AI price declines with other transformative technologies using benchmark-based cost estimates, giving readers a measured sense of how fast cost per capability is falling.

  5. ModelScopeAI score43

    Jina-OCR-v1 parses full pages into Markdown at 2.57 pages per second

    AIJina-OCR-v1, a 3.4B-parameter MoE model that activates 570M parameters per token, converts entire document pages into structured Markdown at 2.57 pages per second. It scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, 7.4 points above DeepSeek-OCR on the latter, and delivers the highest throughput among 14 evaluated systems at concurrency 32. The model is released under CC BY-NC 4.0, so commercial use requires permission.

    Image from @ModelScope2022's post

Sep 27

Sep 27Sun
  1. Xiaomi MiMo · new models on Hugging FaceAI score44

    Xiaomi releases MiMo-V2.6-Flash-MOPD, an upgraded MoE model with 1M context

    AIXiaomi has released MiMo-V2.6-Flash-MOPD on Hugging Face, an upgrade of the MiMo-V2.6-Flash-RL checkpoint that fuses several domain-specialized teachers into one model. The sparse MoE model has 309B total and 15B activated parameters, a 1M-token context length, and supports text, image, video, and audio inputs. The checkpoint targets tool-call repetition, a failure mode where the model repeatedly issues the same or similar tool calls without making progress.

  2. Xiaomi MiMoAI score62

    Xiaomi MiMo Explains Fixing Tool-Call Repetition in MiMo-V2.6 Models

    AIXiaomi MiMo reports that tool-call repetition in MiMo-V2.6 reached over 0.05% of responses across agent harnesses, causing stalled agents and wasted context. The team traced the cause to an RL flooding penalty set at 32 calls per turn, which missed smaller excess behavior, and replaced the approach with a specialized teacher distilled via MOPD. Repetition rates for both Pro and Flash dropped substantially, at roughly $90,000 versus an estimated $2.31 million for the alternative fix.

    Why it matters: The post traces an agent failure to a reward blind spot and compares the costs of two fixes, offering a transferable debugging method for RL-trained tool-calling models.

Sep 26

Sep 26Sat
  1. Sebastian RaschkaAI score30

    Raschka's Reasoning from Scratch Covers Log-Probability Scoring and Self-Refinement

    AISebastian Raschka's fifth Reasoning from Scratch video explains log-probability scoring and self-refinement for LLMs. It covers token probabilities, PyTorch implementation, numerical stability, and a self-refinement loop evaluated on MATH-500, with the log-probability concept linked to cross-entropy loss in pre-training and distillation.

    Video from @rasbt's post
  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score45

    Intern-Decision-4B: Multimodal structured decision model from Qwen3.5-4B

    AIShanghai AI Lab's InternLM released Intern-Decision-4B, a multimodal structured decision model fine-tuned from Qwen3.5-4B, which returns answer distributions for multiple questions in one forward pass. On its benchmark table it scores an average of 90.02 with a Brier score of 0.347 and an ECE of 0.065, and per-query latency averages 44.16 ms on a single RTX 4090. The model is available with a Python DecisionEngine inference interface.

  3. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score44

    Intern-Decision-2B: Structured Multi-Question Decision Model Fine-Tuned from Qwen3.5-2B

    AIShanghai AI Lab's InternLM released Intern-Decision-2B, a multimodal structured decision model fine-tuned from Qwen3.5-2B that returns calibrated answer distributions for multiple questions in one forward pass. It averages 84.68 across listed benchmarks with a 0.437 Brier score and 33.28 ms mean latency on a single RTX 4090. Model weights, a Python DecisionEngine API, and GitHub code are available, with support for up to 16 questions and eight images.

  4. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score46

    Intern-Decision-0.8B: InternLM's structured decision model on Hugging Face

    AIInternLM released Intern-Decision-0.8B, a multimodal structured decision model fine-tuned from Qwen3.5-0.8B that scores answers to multiple questions in one forward pass. The model reports a 79.38 average score and a 33.98 ms mean latency on a single RTX 4090, with 0.8B, 2B, and 4B sizes available. It is accessed through a Python DecisionEngine API that returns calibrated probabilities rather than generating free-form text.

Sep 25

Sep 25Fri
  1. LMSYS OrgAI score38

    SGLang adds multi-item scoring for faster decision model serving

    AISGLang's /v1/score endpoint returns scores for exact requested labels such as Yes/No or A/B/C, and its multi-item scoring (MIS) computes shared context once while keeping candidates isolated. On Qwen3-8B, 16-candidate p95 latency dropped from 54.1 ms with Generate to 20.6 ms with MIS. On Qwen3-0.6B, MIS p95 stayed under about 100 ms as load rose, versus seconds for Generate and SIS.

    Image from @lmsysorg's post

Sep 24

Sep 24Thu
  1. ModelScopeAI score23

    NeoHorse-Jev-4B open model turns app states into structured decisions

    AIModelScope has released NeoHorse-Jev-4B, a compact open model that converts application states into structured decisions and probabilities. It scores 77.70 across six text decision benchmark groups, ranking first among four open-weight models with complete results in the comparison. Its prefill-only inference supports Choice, Noul, and Score primitives, accepts text or a single image with text, and is available under Apache 2.0 for deployment via vLLM, SGLang, Python, CLI, or HTTP.

    Video from @ModelScope2022's post
  2. vLLMAI score42

    TileRT and vLLM hit 469 tok/s on GLM-5.3 with MI355X

    AIThe TileRT and AMD teams reached 469 tok/s single-user decode for GLM-5.3 on 8× MI355X using vLLM. The setup disaggregates work, with vLLM handling prefill and TileRT handling latency-critical decode through vLLM's V1 connector interface. SemiAnalysis's AgentX benchmark reports the configuration at 470 TPS on GLM 5.3 (FP8), over 40% faster than GB300 TRTLLM using FP4.

  3. Google ResearchAI score60

    Google Research details four agentic frameworks for coherent long-form video generation

    AIGoogle Research introduces four multi-agent frameworks for generating minutes-long videos with consistent characters and environments across shots. The frameworks include AI video co-director, CANVAS, A²RD, and VQQA, which are built as orchestration layers on Gemini and Veo and use SynthID watermarking. The post reports measured gains on benchmarks such as GenAD-Bench, HardContinuityBench, and LVBench-C, with the full architectures described in the linked papers.

    Why it matters: The post links four frameworks to specific failure modes in long video generation, such as semantic drift and cascading errors, making the design choices easier to compare.

  4. Goodfire ResearchAI score48

    Steering Along Manifolds Beats Linear Steering for Controlling Llama's Days-of-Week Behavior

    AIGoodfire Research shows that steering Llama-3.1 8B along the curved representation manifold of weekdays produces output probabilities that follow the model's natural cyclic behavior, shifting probability mass smoothly from Monday to Tuesday to Friday. Linear steering along a straight vector, by contrast, cuts across the behavior manifold and yields noisy off-target tokens, some not days of the week at all. The authors argue that representation geometry and behavior geometry are linked bidirectionally.

  5. LangChain BlogAI score50

    LangSmith Engine v2 adds red teaming and pre-validated agent fixes

    AILangChain released LangSmith Engine v2, an in-platform agent that scans production traces to detect agent issues and validates proposed fixes before human review. Engine v2 adds Red Teaming, currently in Private Beta for LangSmith Deployment users, which tests agents for weaknesses such as hallucinations and system-prompt violations before they reach production. Engine v2 is available in SaaS deployments for LangSmith Plus and Enterprise plans, with Self-Hosted support and BYOK for Engine coming later.

  6. Anthropic ResearchAI score60

    Anthropic study finds Claude agent trading limited by preference understanding

    AIAnthropic ran a controlled book-swapping market with 201 employees and Claude-powered agents, which reached 0.55 efficiency against a 0.89 optimum. Agents matched participants' own rankings on 61% of book pairs, and about 85% of the shortfall came from imprecise preference representation rather than the trading floor design. Stronger models produced more efficient markets than weaker ones, while instructions mattered less.

    Why it matters: The study separates agent misunderstanding of user preferences from negotiation failure, showing which failure mode limits outcomes in agent-run markets.

  7. LangChain BlogAI score44

    LangSmith Launches Trajectories for Readable, Chronological Agent Session Views

    AILangChain has launched Trajectories in LangSmith, a chronological, conversational view that aggregates human, AI, and tool messages across an agent and its subagents. Trajectories work with traces from LangChain, LangGraph, Deep Agents, OpenAI and Claude agent SDKs, and coding agents like Codex, Claude Code, and Cursor. The feature is available now on all plans in the US.

Sep 23

Sep 23Wed
  1. Philipp SchmidAI score62

    Gemini 3.8 Flash TTS guide shows how to create and reuse your own voice

    AIGemini 3.8 Flash TTS and Flash-Lite TTS are now available in the Gemini API and AI Studio, with a new feature to replicate your own voice or create one from a sentence. The guide shows recording two clips, one of 15-20 seconds of natural speech and one reading a required consent sentence, then creating a reusable voice ID. It also explains that input text is now spoken word for word, so delivery belongs in speech_metadata.style and short sounds inline.

  2. Dario AmodeiAI score76

    Claude Helps Discover a Possible New Gene Editing Enzyme System

    AIAnthropic announced that Claude, working mostly on its own, identified a previously unknown enzyme system in bacteriophage DNA that may represent a new gene editing mechanism. Claude read literature and genome data, proposed experiments, and Anthropic's team carried them out. The function and biotechnological utility of the system remain unclear.

    Why it matters: The post pairs a Claude-led discovery with the lab workflow used to verify it, showing how AI and humans split the research work in biology.

  3. Black Forest LabsAI score62

    Black Forest Labs releases FLUX 3 Action, a robot policy model

    AIBlack Forest Labs says its FLUX 3 Action, a single-step 7B checkpoint, outperforms every other open policy on RoboLab. It processes each second of robot motion 1.45× to 1.66× faster than Pi0.5, and uses a 2.13-second action horizon versus Pi0.5's 1 second. The company adds that its guidance-distilled checkpoint raises the state-of-the-art RoboLab success rate while running 2.85× to 3.15× faster than the previous leading open WAM.

    Image from @bfl_ai's post
  4. Black Forest LabsAI score67

    Black Forest Labs releases FLUX 3 Action, an open 7B world action model for robots

    AIBlack Forest Labs says FLUX 3 Action is an open-weights 7B world action model that ranks first on the RoboLab benchmark. The company says it outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. The model predicts video and actions together, and the company is releasing the weights, code, fine-tuning recipe, benchmarks, and examples. It also integrated the model into Hugging Face's LeRobot with NVIDIA, with edge deployment on NVIDIA Jetson.

    Why it matters: The release pairs benchmark results with the trade-off it claims to remove between world action model performance and VLA speed, which is useful context for robotics teams weighing open models.

    Video from @bfl_ai's post
  5. Google AI StudioAI score62

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle introduces Gemini 3.8 Flash TTS for creative voice design and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient speech generation. Flash TTS supports voice creation from natural language prompts across more than 100 languages and dialects, and both models are rolling out today in the Gemini API and Google AI Studio, with enterprise access coming soon via Gemini Enterprise.