Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Oct 1

Oct 1Thu
  1. Prime IntellectOfficialAI score34

    Qwen3.6 reward rises 2.8x via GRPO on Hosted Training

    AIPrime Intellect reports that after about 100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain. Qwen3.5, trained the same way, reached 0.356, suggesting the method works across model families. Both post-trained models finished well ahead of other open models and narrowed the gap to Claude Opus 4.8, with Qwen3.6 activating only 3B parameters per token.

    Image from @PrimeIntellect's post
  2. Yellowbrick InvestingXAI score18

    Yellowbrick 2.0 launches with leaderboards, API access, and custom feeds

    AIYellowbrick 2.0 is live, tracking 35,000+ stock pitches from 4,000+ authors and adding 300+ new pitches weekly. The rebuilt platform adds author leaderboards, custom feeds and alerts, API access, and paid research partner discounts. Premium subscribers get a 30% discount on Koyfin, which the company says covers the cost of Yellowbrick Premium.

  3. Mustafa SuleymanXAI score40

    Microsoft AI launches MAI-Transcribe-2-Streaming, claiming top real-time transcription accuracy

    AIMicrosoft AI launched MAI-Transcribe-2-Streaming, which Artificial Analysis ranks #1 of 38 models for final transcript accuracy at 2.5% WER, returned 0.13s after end of speech. Artificial Analysis lists its streaming price at $0.54 per hour of audio, at the higher end among leading streaming models. Microsoft's post claims the model is 55% faster and 60% cheaper than ElevenLabs and invites developers to build agents on its platform.

  4. Jerry LiuXAI score42

    LlamaIndex launches Extract v2.5 document extraction agents with improved accuracy

    AILlamaIndex introduced Extract v2.5, a series of agents tuned for document extraction across cost-effective, agentic, and agentic plus tiers. The company reports the agents outperform Opus 5.5 and GPT-6 Sol while costing 30% to 4x less, with accuracy gains on long lists (86.1% to 95.5%), multi-page records (85.5% to 96.5%), and scanned forms (90.9% to 95.7%) on its agentic tier. The release adds advanced citations with bounding boxes and structural reasoning, and the agents are available on LlamaParse.

    Video from @jerryjliu0's post
  5. Lewis Tunstall @ COLM 🌉XAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

    Video from @_lewtun's post
  6. Cloudflare Blog · AIOfficialAI score58

    Cloudflare releases open-source Clef decision models and an RL fine-tuning service

    AICloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and launched a reinforcement learning fine-tuning service. In Cloudflare's tests, Clef classified a domain in 2.2s versus 4.7s for gpt-oss-120b, and the models are Jev-API compatible. The company is offering fine-tuning first through a forward-deployed engineering team, with a self-serve platform planned later.

  7. Hamel HusainXAI score25

    Hamel Husain says not every failure mode needs an automated evaluator

    AIHamel Husain advises against building automated evaluators for every failure mode discovered during AI development. He argues that teams should weigh the costs of different evaluator types, such as code-based checks versus LLM judges, before building an eval.

    Image from @HamelHusain's post
  8. WanOfficialAI score62

    Alibaba's Wan 3.0 ranks first overall on Artificial Analysis video leaderboard

    AIAlibaba's Wan 3.0 ranks #1 overall on the new Artificial Analysis AA-Video-T2V v2.0 text-to-video leaderboard, priced at $12 per minute of video. The benchmark judges models at 1080p using over 68,000 human preference votes across 1,000 prompts, and the author states Wan 3.0 leads 10 of 20 category boards.

  9. AI SupremacyBlogAI score50

    Google announces Gemini 4 Argon, its first frontier model since February

    AIGoogle announced Gemini 4 Argon, a model it says is built to sustain deep reasoning across complex, long-horizon workflows, roughly seven months after its last flagship release in February. The article says cybersecurity testing will be completed after October 1, with no benchmark scores, pricing, or availability details provided.

  10. LangChain BlogOfficialAI score58

    LangChain shows how to build a model router in its Open SWE coding agent

    AILangChain built a model router inside its open source coding agent Open SWE that picks one of three models for each thread. In an A/B test against always using GPT-6 Astra, the median cost per thread fell 64% with no measurable change in merged PR rate. The router runs on the thread's first message, using a base prompt, per-tier criteria, and a classifier model, and the post lists next steps including subagent routing and mid-thread re-routing.

Sep 30

Sep 30Wed
  1. Hamel HusainXAI score38

    Hamel Husain Reviews Claude's New Auto Eval Plugin for Evaluations

    AIHamel Husain has published a longer review of a new Claude Auto Eval plugin after many users asked about it. He invites readers to share their experiences using the plugin and how it went for them. The plugin is part of Claude's ability to help build evaluations and hillclimb on them, as described by @ClaudeDevs.

    Image from @HamelHusain's post
  2. indigoXAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

    Image from @indigox's post
  3. Apple Machine Learning ResearchOfficialAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  4. Google ResearchOfficialAI score40

    Google's science AI tops CDC flu hospital admission forecasts this season

    AIThe CDC announced that Google's science AI model ranked highest among 39 eligible models for forecasting flu-related hospital admissions during the 2025-26 flu season. Google's forecasts were built with Empirical Research Assistance, an AI tool that generates computational solutions across scientific fields.

    Image from @GoogleResearch's post
  5. Yi TayXAI score46

    Gemini 4 Argon launches, reportedly outperforming astra and fable on many tasks

    AIGoogle DeepMind introduced Gemini 4 Argon, a new frontier model built for coding, enterprise knowledge work, and cybersecurity defense, rolling out to trusted testers through its Fairwind Program. Yi Tay says Gemini 4 outperforms astra and fable on many tasks, though the post gives no benchmark figures.

  6. DeedyXAI score22

    Deedy argues AI model pricing signals quality better than benchmarks

    AIDeedy argues that benchmark scores for frontier AI models are increasingly meaningless because labs tune for them before launch. He suggests trusting price instead: a high price indicates a genuinely good model, while a low one suggests the model is benchmaxxed.

  7. whXAI score67

    Gemini 4 Argon previewed with frontier coding and cyber defense claims

    AIThe post quotes Google's Sundar Pichai introducing Gemini 4 Argon as an early look at the next model. It claims frontier performance in complex workflows, cyber defense, and software engineering, and says Google teams are using it for tasks from coding to quantum computing. The author adds that on FrontierSWE the model is very self-critical and often says "Eureka!", a personality they describe as a large improvement over previous Gemini models.

    Image from @nrehiew_'s post
  8. Yuchen JinXAI score12

    Yuchen Jin says Google may be back, citing benchmark gains

    AIYuchen Jin suggests Google is back in the AI race after a model reportedly outperforms Astra and Opus 5.5 across the board on benchmarks. He adds a caveat that the results may reflect benchmark optimization rather than real capability gains, and says he would welcome Google rejoining the race.

    Image from @Yuchenj_UW's post
  9. Google DeepMindOfficialAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  10. Google · Gemini appOfficialAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

  11. PerplexityOfficialAI score20

    Perplexity's embedding preview tops ConTEB benchmark average nDCG@10

    AIPerplexity's embedding preview achieves the highest average nDCG@10 among tested models on ConTEB, though not on every task. It also outperforms voyage-context-4 on chunk retrieval while using 8x less storage per vector, at 1 KB (1024 dims, int8) versus 8 KB (2048 dims, float32).

    Image from @perplexity_ai's post
  12. PerplexityOfficialAI score23

    Perplexity submits turbopuffer's context-bench for blind evaluation

    AIPerplexity says context-bench, a context-aware retrieval benchmark created and privately held by turbopuffer, has 2,099 queries and 38,894 documents. Perplexity submitted it for blind evaluation, with queries and capabilities inspired by turbopuffer customer conversations.

    Image from @perplexity_ai's post
  13. Google Cloud TechOfficialAI score28

    Agent Clinic Ep 3 builds automated eval suite for LangGraph agent

    AITerminal test runs miss multi-turn agent regressions, so Agent Clinic Episode 3 builds an automated eval suite for a LangGraph agent in 60 minutes. The post presents a four-step framework for moving from informal checks to benchmarking AI agents, with a link to the full guide.

    Image from @GoogleCloudTech's post
  14. TypeSafe AIOfficialAI score27

    Jev reranking beats GPT-5 Mini on sales data retrieval

    AIJev reranking retrieves Rox sales data 20x faster, 10x cheaper, and 12% more accurate than GPT-5 Mini. The benchmark compared Jev classification against LLM-based reranking for pulling transcripts, emails, CRM notes, news, and documents.

  15. FireworksOfficialAI score30

    Fireworks' Ember-1 matches Kimi K3 on Vals with fewer reasoning tokens

    AIFireworks' Ember-1 performs close to Kimi K3 on Vals' finance and legal benchmarks while using fewer reasoning tokens per turn. Fewer tokens per turn lower cost and speed up agent loops, and Ember-1 is available on Fireworks Serverless.

  16. Stanford HAIOfficialAI score23

    Foundation models could fill gaps in incomplete biomedical data

    AIBiomedical datasets are often incomplete, such as patients with imaging but no genomic data. Stanford speaker Olivier Gevaert will discuss using foundation models to fill these gaps in multimodal precision medicine modeling on October 7.

  17. IdeogramOfficialAI score23

    Ideogram 4.5 Performs Strongly Across General Image Editing Tasks

    AIIdeogram 4.5 was built for precise, targeted editing but also performs very well across general editing tasks. Design Arena ranks it 15th in Image Editing with an Elo of 1250, placing it in the same performance band as MAI-Image-2.6 and Gemini 3 Pro Image Preview. It is especially strong at typography edits, such as modifying text in infographics.

  18. Ant LingOfficialAI score22

    Ling-3.1-flash scores 65.35 on HealthBench Professional benchmark

    AIAnt Ling reports that Ling-3.1-flash scored 65.35 on HealthBench Professional, a healthcare evaluation rather than clinical certification. The model was trained with feedback from AQ and Haodf.com (Good Doctor Online) healthcare scenarios.

    Image from @AntLingAGI's post
  19. Ant LingOfficialAI score46

    Ant Ling releases Ling-3.1-flash with 1M-token context, plans open-source

    AIAnt Ling introduced Ling-3.1-flash, a model with about 560B total parameters, about 25B active per token, and up to a 1M-token context window. The company plans to open-source the model soon. It reports 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional across work, coding, and healthcare tasks.

    Image from @AntLingAGI's post