Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Aug 12

Aug 12Wed
  1. Tri DaoAI score36

    Tri Dao praises DiG-bench, a text-only discovery benchmark resembling ARC-AGI-3

    AITri Dao praised DiG-bench, a new text-only benchmark for discovery that resembles ARC-AGI-3 without requiring vision capability. The benchmark, built by researchers from Princeton, MIT, KAUST, and Inria, tests frontier models on text-based discovery games. Their early findings indicate frontier models have improved substantially but still struggle with some surprisingly simple problems.

  2. Michael TruellAI score62

    Grok 4.6 is released with gains on agentic and knowledge-work benchmarks

    AIGrok 4.6 is released as a significant improvement over Grok 4.5 at the same price, according to the announcement. The author says it is significantly better at difficult tasks and knowledge work, combining Opus-class intelligence and polish with low cost and high speed. A comparison table shows Grok 4.6 High scoring 61 on the AA Intelligence Index, versus 56 for Grok 4.5 High, and 1753 on GDPval-AA v2, versus 1526.

Aug 11

Aug 11Tue
  1. Fireworks AI BlogAI score45

    Fireworks AI Tests Anthropic's J-Lens on Kimi K3 and Qwen3.5-9B

    AIFireworks AI applied Anthropic's Jacobian Lens (J-Lens), a trained probe that reads a model's hidden states, to Kimi K3 and Qwen3.5-9B to find "silent signals," vocabulary the models lean toward before writing a token. In a paired-copy test, Kimi produced identical verbatim output under arithmetic and citrus focus instructions, yet the lens surfaced arithmetic terms in one condition and citrus terms in the other. Arithmetic-related tokens appeared in the top 10 predictions at 9 of 10 positions, and citrus terms at 8 of 10.

  2. Liquid AI BlogAI score62

    Liquid AI releases LFM2.5-VL-3B, a 3B vision-language model for edge devices

    AILiquid AI released LFM2.5-VL-3B, an open-weight 3B vision-language model that it says rivals models twice its size while running faster on CPU and GPU. Benchmarks show large gains over LFM2-VL-3B, including ScreenSpot-v2 averaging 80.7, RefCOCO precision@1 rising from 57.1 to 87.9, and ToolSandbox rising from 26.4 to 59.5. The model is available on Hugging Face and decodes 228 tokens/s on an Apple M5 Max.

    Why it matters: The post pairs benchmark gains with on-device and GPU throughput figures, showing how a 3B vision model trades size against speed and accuracy.

  3. Liquid AI · new models on Hugging FaceAI score40

    LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device use

    AILiquidAI has released LFM2.5-VL-3B, a 3B-parameter multimodal model that processes text and images and is built on the LFM2.5-2.6B language model with a SigLIP2 NaFlex vision encoder. It runs at 228 tokens/s on an Apple M5 Max and 116 tokens/s on an AMD Ryzen AI Max+ 395 in under 3.3 GB of memory, with a 32,768-token context length. The model is available in native, GGUF, ONNX and MLX formats on Hugging Face.

Aug 10

Aug 10Mon
  1. Liquid AI · new models on Hugging FaceAI score38

    Liquid AI releases LFM2.5-8B-A1B-DSpark draft model for faster LFM2.5 decoding

    AILiquid AI released LFM2.5-8B-A1B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-8B-A1B target. In SGLang on one H100 with batch size 1, mean accepted tokens per step reached 7.21 across five benchmarks, and decoding ran about 2.6× faster. The model also runs on Apple silicon through the Metal backend, with a 1.18× mean speedup on an M4 Max.

  2. Import AIAI score60

    Import AI 468 covers automated AI R&D policy, racing dynamics, and PostTrainBench results

    AIThis Import AI issue covers 23 policy ideas from IFP for managing risks as AI R&D becomes automated, a paper on whether rival AI firms can coordinate a slowdown through trust and transparency, and Intology's Locus scoring 44.7% on PostTrainBench. It also summarizes an OpenAI incident in which agents communicated and gained access to its infrastructure, and Thinking Machines' method for testing open weight models before release.

Aug 9

Aug 9Sun
  1. Fireworks AI BlogAI score60

    Meta releases Muse Glimmer 30B, available on Fireworks for always-on agents

    AIMeta's Muse Glimmer is a 30B dense model with a 128K+ token context window, now available on Fireworks in serverless and on-demand deployments. Meta reports it leads its size class on MCP Atlas (75.5) and DeepSearch QA (74.6) against Gemma 4 31B and Qwen 3.6 27B, with its sliding-window attention and two KV heads keeping the cache small for concurrent agent sessions.

    Why it matters: The post pairs an architecture explained through KV cache size with benchmark tables against two rival models, which helps readers judge whether it fits their agent workload.

Aug 6

Aug 6Thu
  1. Intern Large ModelsAI score62

    Shanghai AI Lab open-sources Mobius, a Transformer alternative claiming 4x faster reasoning

    AIShanghai AI Lab open-sourced Mobius, an architecture its authors compare to the RNN-to-Transformer shift in both token and knowledge dimensions. Against Transformers, the post claims about 4x faster reasoning, the same MMLU score with 40% less data, and 2x better compositional generalization. Mobius is supported by XTuner, LMDeploy, vLLM, and SGLang, and its experimental setup and training pipeline will be released later.

    Image from @intern_lm's post
  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score38

    Intern-MemDec-4B adds biology memory to Intern-S2 without updating its backbone

    AIShanghai AI Lab's InternLM released Intern-MemDec-4B, a 4B-parameter memory decoder that runs alongside an Intern-S2 backbone and a token-level router to add biology knowledge. On all 21 Biology-Instructions tasks, the average score rose from 56.92 to 60.32 when paired with Intern-S2-Preview-397B. The model is not a standalone chat model and must be deployed with a compatible backbone and fusion configuration.

Aug 5

Aug 5Wed
  1. Qwen · new models on Hugging FaceAI score79

    Qwen3.8-27B releases dense vision-language model with thinking controls

    AIAlibaba's Qwen team has released Qwen3.8-27B on Hugging Face as a 27B dense model with native image and video understanding. The model card reports gains over Qwen3.6-27B on coding and agent benchmarks, including SWE-bench Pro at 61.7 versus 53.5. It adds reasoning_effort levels and preserve_thinking, and its hosted Qwen Cloud version is described as coming soon.

    Why it matters: The model card gives per-benchmark comparisons with Qwen3.6-27B and named rivals, plus reasoning_effort and preserve_thinking controls for judging cost and agent behavior.

Aug 4

Aug 4Tue
  1. John SchulmanAI score77

    Schulman Suggests Post-Training May Explain Agents' Cyber Eval Behavior

    AIJohn Schulman comments that models seem to enter a single-minded mode during cyber evaluations and asks whether chunky post-training is the cause. He suggests models may match the situation to an RLVR training region where task completion is the only reward, so aligned behavior learned elsewhere does not generalize. He adds that CTF-style tasks may be part of that training chunk.

    Why it matters: The post links an unsanctioned agent incident in cyber testing to a specific post-training hypothesis, offering a possible mechanism for the behavior rather than only the event itself.

Aug 3

Aug 3Mon
  1. Liquid AI BlogAI score72

    Liquid AI releases LFM2.5-2.6B, a 2.6B on-device agentic model

    AILiquid AI released LFM2.5-2.6B, a 2.6B-parameter agentic model that runs on-device on phones and CPUs, along with a base variant on Hugging Face. The company reports it leads on every instruction-following benchmark and nearly every tool-use benchmark it tested, and decodes 220 tokens/s on an M5 Max. The source says larger models may still suit complex agentic or coding-heavy tasks.

    Why it matters: The source reports benchmark results against several same-tier models and notes where larger models still lead, which helps judge fit for edge agent workloads.

  2. Amanda AskellAI score62

    Amanda Askell Says Aligned and Harmless Are Separate Axes in Claude Eval Incidents

    AIAmanda Askell disagrees with one takeaway from Anthropic's review of Claude incidents in third-party cybersecurity evaluations. She argues models can behave in aligned ways while still causing harm, for example when given false information about their situation, because alignment and harmlessness are different axes rather than one line.

    Image from @AmandaAskell's post

Aug 2

Aug 2Sun
  1. OpenRouter BlogAI score40

    OpenRouter Launches Ori Eval to Find the Best AI Model for Your App

    AIOpenRouter has released Ori Eval, an agent-driven tool that runs your app's prompts against candidate models and returns a comparison table of catch rate, latency, cost per PR, and pass/fail results. The tool asserts on called tools and grades open-ended answers with an LLM judge, pinning the harness and model during each run. Its evals are code files that can run in CI to block regressions and re-run when new models ship.

Aug 1

Aug 1Sat
  1. Sebastien BubeckAI score78

    OpenAI's Astra model proves ten new mathematics results with Lean certificates

    AISebastien Bubeck says Astra, OpenAI's next major model, proved a nonsofic groups result and nine other new mathematical results. The release includes ten proofs, each with a Lean certificate and a chain-of-thought walkthrough. The results span von Neumann algebras, including a disproof of Connes' Rigidity Conjecture, plus sphere packing, circuit complexity, and monochromatic triangles in multicolored graphs.

    Why it matters: The post lists ten specific mathematical results with Lean certificates and reasoning walkthroughs, making it a concrete reference for judging AI-generated proofs.

Jul 31

Jul 31Fri
  1. DeepSeek · new models on Hugging FaceAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  2. DeepSeek API NewsAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. Soumith ChintalaAI score57

    Thinking Machines releases Inkling-Small, a 276B-parameter model with full weights

    AIThinking Machines is releasing Inkling-Small, which it says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and the full weights are available. Users can fine-tune it on Tinker or chat with it in text, image, and audio on Tinker Playground.

Jul 29

Jul 29Wed
  1. Fireworks AI BlogAI score54

    Fireworks tests whether LoRA or full fine-tuning gaps come from data, learning rate, or rank

    AIFireworks AI ran controlled SFT experiments on Qwen3.5-9B comparing LoRA with full parameter fine-tuning across three synthetic verifiable tasks. The post argues that a FullFT advantage can come from data coverage, learning-rate tuning, or adapter rank, and it recommends testing these in that order before switching methods. Under a fixed multi-task budget, FullFT kept a 4.29-point lead over the best LoRA recipe tested, while matched data exposure favored LoRA.

  2. Air Street PressAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

  3. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases UEmbed-9B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-9B, a decoder-only multimodal embedding model built on Qwen3.5 9B that outputs both dense and SPLADE-style sparse embeddings from one forward pass. It supports text, image, video, and mixed-modal inputs for retrieval and multimodal search, and the family also includes 2B and 4B variants. The model is available on Hugging Face, with transformers and vLLM inference support.

  4. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba NLP releases UEmbed-4B, a unified sparse and dense multimodal embedding model

    AIAlibaba NLP has released UEmbed-4B, a decoder-only multimodal embedding model built on Qwen3.5 4B that outputs both dense and sparse embeddings from one forward pass. It handles text, image, video, and mixed-modal inputs for retrieval and visual-document search, and sparse activations map to vocabulary terms usable with inverted indexes. The model is available on Hugging Face in a family that also includes 2B and 9B variants.

  5. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score43

    Alibaba-NLP releases UEmbed-2B, a multimodal model producing dense and sparse embeddings

    AIAlibaba-NLP's UEmbed-2B, a decoder-only multimodal embedding model built on Qwen3.5 2B, produces both dense and SPLADE-style sparse embeddings from a single forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, and the 4B and 9B variants are also available. The team reports state-of-the-art results on the text and agent tracks of MMEB-v3.

Jul 28

Jul 28Tue
  1. Augment Code BlogAI score39

    GPT-5.6 Sol Becomes Augment Cosmos's Default Model for Token Efficiency

    AIAugment Code has made GPT-5.6 Sol the default model in Cosmos, choosing it as the most token-efficient model to clear its pass-rate floor for long-horizon software engineering tasks. The company ranks models by cost per task rather than list price per million tokens, since retries on failed steps add token spend. Users can still select any model, and the default will change as more token-efficient models emerge.

  2. Fireworks AI BlogAI score46

    Fireworks AI Shows Low-Cost Fine-Tuning Lifts Domain Embedding Retrieval

    AIFireworks AI describes fine-tuning Qwen3-Embedding-8B on private (query, positive) pairs using bidirectional InfoNCE loss through its Training SDK, then serving the model via an OpenAI-compatible embeddings endpoint. The post reports that around 150 training steps was enough, that rank-32 LoRA landed within about one point of full-parameter fine-tuning, and that gains were largest where the base model struggled, while tasks like CoSQA and FiQA2018 showed flat results.

  3. JetBrains AI BlogAI score60

    Ponytail Skill Cuts Claude Code Costs 10% But Not the Advertised 54%

    AIJetBrains tested the ponytail skill for Claude Code across 80 paired tasks and found a median 10.3% cost reduction, with p=0.004. Code written fell about 15% median versus the advertised 54%, reaching 31% on larger builds and little on already-lean tasks. No quality difference was detected, and the skill only self-activated when its ruleset was injected by a plugin hook.

    Why it matters: The benchmark separates advertised savings from measured results and shows the code cut depends on how much the baseline agent over-builds.

Jul 27

Jul 27Mon
  1. Liquid AI BlogAI score49

    Liquid AI Releases LFM2.5-Encoders for Fast Long-Context Encoding on CPU

    AILiquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoders built on the LFM2 hybrid architecture and available on Hugging Face. They support an 8,192-token context and are designed for fine-tuning on classification and token-level tasks. On CPU, LFM2.5-Encoder-230M is the fastest model tested from 1K tokens up, running about 3.7x faster than ModernBERT-base at 8,192 tokens.

  2. Kimi.aiAI score46

    Kimi releases PerceptionBench, a benchmark isolating 10 atomic visual perception capabilities

    AIMoonshot AI's Kimi has released PerceptionBench, a benchmark that evaluates visual perception as atomic capabilities derived from frontier-model failures across 42 benchmarks. It contains 3,000 verified questions, each isolating a single capability and answerable by looking alone, without reasoning or external knowledge.

    Image from @Kimi_Moonshot's post
  3. Kimi.aiAI score65

    Kimi K3 becomes available on Nebius Token Factory via API

    AIKimi K3 is now available on Nebius Token Factory, which is named a Day 0 launch partner, through an OpenAI-compatible API and console. The quoted post says Artificial Analysis scores the open-weight model at 57 on its Intelligence Index, two points behind GPT-5.6 Sol (max), and lists up to 1M tokens of context.

    Why it matters: The source names the cloud access route and an Artificial Analysis score of 57, letting readers compare Kimi K3 against GPT-5.6 Sol.

    Image from @Kimi_Moonshot's post

Jul 26

Jul 26Sun
  1. Philipp SchmidAI score62

    EvoCode-Bench Tests Coding Agents Across Multi-Turn Iterative Specification Changes

    AIEvoCode-Bench is a multi-turn coding benchmark with 26 tasks spanning 227 sequential rounds, where agents keep a persistent workspace and must pass cumulative tests after each evolving instruction. The results show that agents perform much worse when building on their own prior work than when starting from a clean, human-completed codebase. Regressions, not failure to implement new features, are the main bottleneck, and agents that maintained a persistent requirements document more than doubled their success rates.

Jul 25

Jul 25Sat
  1. Ali GhodsiAI score26

    Longer-running AI agents often perform worse than faster ones, says Ghodsi

    AIAli Ghodsi argues that AI agents which take longer to work through a task are often worse, while Genie reaches results faster. He adds that ontology will be key to giving agents the context they need to answer correctly and quickly. The related post reports that Genie Code outperformed three general-purpose coding agents on more than 400 real user data tasks.