Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 30

Sep 30Wed
  1. IdeogramAI score23

    Ideogram 4.5 Performs Strongly Across General Image Editing Tasks

    AIIdeogram 4.5 was built for precise, targeted editing but also performs very well across general editing tasks. Design Arena ranks it 15th in Image Editing with an Elo of 1250, placing it in the same performance band as MAI-Image-2.6 and Gemini 3 Pro Image Preview. It is especially strong at typography edits, such as modifying text in infographics.

  2. Ant LingAI score46

    Ant Ling releases Ling-3.1-flash with 1M-token context, plans open-source

    AIAnt Ling introduced Ling-3.1-flash, a model with about 560B total parameters, about 25B active per token, and up to a 1M-token context window. The company plans to open-source the model soon. It reports 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional across work, coding, and healthcare tasks.

    Image from @AntLingAGI's post
  3. Liquid AIAI score42

    LongevityBench: Liquid AI's compact LFMs beat frontier models on aging tasks

    AILiquid AI and InSilicoMeds released LongevityBench, an aging benchmark with 17 tasks spanning clinical records, DNA methylation, transcriptomics, proteomics, and genetics. On several tasks, Liquid AI's compact LFMs outperformed every frontier model the team evaluated. The team plans to present the work to the longevity research community at ARDD this week.

    Video from @liquidai's post
  4. Tencent HyAI score62

    Tencent Hunyuan releases ExplorationBench to test how AI systems discover rules

    AIResearchers from Tencent Hy, Fudan University, and Tsinghua University released ExplorationBench, a benchmark that tests whether AI systems can discover hidden rules in executable Alien World sandboxes. Across 10 frontier systems, getting feedback from experiments outperformed thinking alone, with the best run reaching 89.0% after four rounds. The authors note that rankings barely transfer between the two worlds, and the code is listed as coming soon.

    Image from @TencentHunyuan's post
  5. The SequenceAI score50

    The Sequence Learning Loop: Opus 5.5, DeepSeek Environments, and Claude's DNA Discovery

    AIIssue 942 of The Sequence links Anthropic's Claude Opus 5.5, reported for the week of September 21–27, to DeepSeek's September 19 environments paper and a report of AI-assisted biological discovery. The newsletter argues that progress increasingly depends on the surrounding machinery that governs where a model acts, what it observes, and how its conclusions are checked.

  6. Hamel HusainAI score42

    Hamel Husain Tests Anthropic's Claude Eval Plugin on Leasing Assistant Traces

    AIHamel Husain reviewed Anthropic's new build_eval and hill-climb commands in the claude-api plugin for Claude Code, finding it useful for discovering issues like human handoff, formatting, and voice agent problems. He criticized it for pushing evaluator creation before data review, asking for label validation in Markdown files, and bundling four failure checks into one broad call-transfer evaluator. Husain says he would hold off on using it for now.

  7. ModelScopeAI score62

    InSpatio-World 1.5 turns images and videos into real-time explorable 4D worlds

    AIInSpatio-World 1.5 from InSpatio_AI turns a single image, four images, a panorama, or a video into a navigable scene with wide viewpoint changes. The 1.3B model scores 68.72 on WorldScore-Dynamic, ranking first among evaluated real-time and interactive methods, with speeds up to 24 FPS. The post says the code is released under Apache 2.0 and that dependencies keep their own licenses.

    Video from @ModelScope2022's post
  8. Artificial Analysis ArticlesAI score39

    Upstage Releases Solar Mini 4 Reasoning Model, Scoring 24 on Intelligence Index

    AIKorean AI lab Upstage has released Solar Mini 4, a proprietary reasoning model that scores 24 on the Artificial Analysis Intelligence Index with 35B total and 3B active parameters. It is priced at $0.10/$0.40 per 1M input/output tokens and has a 1M-token context window, but averages 7.1 minutes per task due to heavy output token use. Its weights are not released, and its size cannot be independently verified.

  9. Artificial Analysis ArticlesAI score75

    Gemini 4 Argon matches GPT-6 Astra on intelligence index at lower cost

    AIArtificial Analysis reports that Google's Gemini 4 Argon scores 53 on its Intelligence Index with high reasoning, matching GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max). At the current 50% launch discount, its cost per task is $1.99, about 60% of GPT-6 Astra's $3.26, but the discount's end date is unconfirmed and standard pricing would raise it to $3.98. The model is being rolled out to selected users and is not publicly available.

    Why it matters: The benchmark compares Gemini 4 Argon's cost per task and hallucination rate with GPT-6 Astra, showing where its value depends on a temporary 50% discount.

Sep 29

Sep 29Tue
  1. Jerry LiuAI score20

    Jev, a System One model, tops OSS rivals on document tasks

    AIJerry Liu says Jev, a System One model, outperformed other open-source classifiers and document-specific models on orientation detection, language detection, classification, and splitting. The benchmark measured accuracy, cost, and latency across these fast document decisions, with Jev leading most comparisons. The benchmark code is available in the run-llama/jev_vs_oss repository.

    Video from @jerryjliu0's post
  2. Jerry LiuAI score22

    GPT-6.1 Sol Improves Table Parsing and Reading Order in OCR Benchmarks

    AIJerry Liu benchmarked gpt-6.1 sol on document OCR tasks and found a sizable increase in table parsing and reading order over gpt-6 sol from a week earlier. Its table parsing is similar to gpt-6 astra. He noted frontier models still cost roughly an order of magnitude more than cost-effective document parsing solutions, leaving room to improve the premium end above 1c per page.

    Image from @jerryjliu0's post
  3. Hugging Face BlogAI score46

    Open TTS Leaderboard ranks multilingual and voice cloning models using objective metrics

    AIHugging Face released the Open TTS Leaderboard, which evaluates open-source text-to-speech models using objective metrics instead of arena-style human votes. It measures intelligibility via WER and CER using Qwen3 ASR, speed via RTFx and time-to-first-audio on an H200 GPU, and speaker similarity via WavLM embeddings. The leaderboard covers multilingual results and voice cloning, and it is intended to complement, not replace, human preference rankings.

  4. Apple Machine Learning ResearchAI score38

    LLM Conditioning Study Finds Steering Methods Trade Fluency for Effectiveness

    AIApple researchers systematically tested LLM conditioning methods and found efficient activation steering often degrades fluency. Steering is far less effective on instruction-tuned models than base models, while prompting and full supervised fine-tuning work for concept injection but are weaker at concept removal. Cheap textual metrics correlate highly with costly LLM-as-judge scores.

  5. Liquid AIAI score32

    Liquid AI launches d1, first decision model, beating Jev on HF index

    AILiquid AI announced d1, its first decision model, which it says is the first to outperform Jev on Hugging Face's Decision Index. The company claims d1 wins on multilingual evals, resists prompt injection better, handles longer inputs more effectively, and is built for fast, structured decision-making in software environments. It is available via the Liquid API at console.liquid.ai, with OpenRouter availability coming soon.

    Image from @liquidai's post
  6. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    AIBAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    Why it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

  7. Jerry LiuAI score22

    Jerry Liu and Snorkel's Vincent Sun discuss evals and RL environments

    AIJerry Liu hosted a dinner with Snorkel's Vincent Sun on evals and RL environments, a topic shaped by models rapidly saturating benchmarks. The conversation highlighted that building fair RL environments is hard, since failures are difficult to attribute to input, harness, or reward model, and that long-horizon evals spanning weeks or months remain very difficult. The post also noted that regulated industries still require human-in-the-loop review because 80% accuracy is not sufficient.

    Image from @jerryjliu0's post
  8. Replit BlogAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.