Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Oct 6

Oct 6Tue
  1. ARC PrizeOfficialAI score22

    Grok 4.7 uses more reasoning tokens than Grok 4.6 on ARC-AGI-2

    AIGrok 4.7 used more reasoning tokens on average than Grok 4.6 on ARC-AGI-2 semi-private tasks at medium, high, and xhigh reasoning levels, raising its cost per task. Per test-pair attempt, medium used 136% more tokens, high 125% more, and xhigh 173% more, while low used 27% fewer. A chart compares the two models at xhigh on the 20 public tasks where Grok 4.7 increased token use the most.

    Image from @arcprize's post
  2. ARC PrizeOfficialAI score38

    Grok 4.7 scores 1.8% on ARC-AGI-3, trails Grok 4.6 on ARC-AGI-2

    AISpaceXAI's Grok 4.7 scored 90.2% on ARC-AGI-1 at $0.64 per task, higher than Grok 4.6, according to ARC Prize. It reached 61.4% on ARC-AGI-2 at $2.01 per task and 1.8% on ARC-AGI-3 under the standard harness ($2.7k), or 10.0% with the provider adapter harness ($4.8k), both lower than Grok 4.6 on those two benchmarks.

    Image from @arcprize's post
  3. Microsoft ResearchOfficialAI score36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    AIMicrosoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

  4. GoogleOfficialAI score43

    EmbeddingGemma 2 Delivers Best-in-Class Performance at 740M Parameters

    AIGoogle's EmbeddingGemma 2 is a 740M-parameter embedding model that outperforms some models more than twice its size while using about 191MB to 567MB of active RAM. It offers an 8K context window, 4x larger than the first generation, and can process up to 5.5 minutes of audio, 29 images, or 58 video frames in one pass.

    Image from @Google's post
  5. Google DeepMind · The KeywordOfficialAI score72

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, a 740-million-parameter embedding model that maps text, images, audio, and video into a shared space and runs on local hardware under an Apache 2.0 license. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions to 512, 256, or 128, and the model supports an 8K-token context window. The model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform availability coming soon.

    Why it matters: The release shows how a 740M-parameter multimodal embedder runs locally with a 768-to-128 dimension truncation option, useful for judging on-device retrieval designs.

  6. ARC PrizeOfficialAI score14

    ARC Prize hosts Frontier AI benchmarking dinner with Snorkel AI during SF Tech Week

    AIARC Prize is hosting its first Frontier AI Benchmarking Ecosystem Dinner with Snorkel AI during SF Tech Week. The event will bring together 40 experts from nonprofit benchmarking organizations, industry, academia, and government to discuss AI benchmark design and building a more scalable, transparent evaluation ecosystem. Attendees named in the post include representatives from METR, Epoch, Vals, Artificial Analysis, Harvard, and Stanford.

  7. ARC PrizeOfficialAI score14

    DeepSeek to share ARC-AGI-3 results in coming weeks

    AIDeepSeek says its ARC-AGI-3 evaluations are more operationally intensive than expected, so results will follow once testing is complete. The scores will be reported with a new DeepSeek provider adapter harness.

  8. ARC PrizeOfficialAI score25

    ARC Prize finds DeepSeek V4.1 Flash high reasoning gains no clear edge

    AIOn ARC-AGI-1, DeepSeek V4.1 Flash scored 88.5% at high reasoning versus 90.5% at low, with high using 35% more output tokens without consistently better answers. On ARC-AGI-2, the reported per-task cost of max reasoning ($0.129) appears slightly lower than high ($0.133), but after excluding incomplete tasks caused by API issues, max is about 4.5% more expensive per task.

    Image from @arcprize's post
  9. ARC PrizeOfficialAI score46

    DeepSeek V4.1 Flash scores 72.9% on ARC-AGI-2 at $0.13/task

    AIDeepSeek V4.1 Flash reaches 72.9% on ARC-AGI-2 at $0.13 per task and 94.5% on ARC-AGI-1 at $0.07 per task, according to ARC Prize verification. Compared with V4 Flash's best scores, it gains 11.5 points on ARC-AGI-2 and 5.5 points on ARC-AGI-1, but costs about 250% more per task.

    Image from @arcprize's post
  10. Aravind SrinivasXAI score13

    Perplexity Decider is the best decision model, per Tetris test

    AIPerplexity's Decider model was ranked best among eight decision models in a Tetris benchmark, according to a quoted post from @alokbishoyi97. The post says Decider consistently placed at the top of the tests, which were run through the Tetris Royale playground.

  11. SemiAnalysisXAI score18

    ClusterMAX rates FarmGPU underperform on Slurm and Kubernetes testing

    AISemiAnalysis rated FarmGPU as ClusterMAX Underperform after its Slurm layer failed to advertise GPU resources and Kubernetes exposed no RDMA devices for scale-out networking. The post credits FarmGPU's Grafana monitoring, provisioning notes, and trustworthy technical team, while noting the team may be stretched thin across small clusters.

    Image from @SemiAnalysis_'s post
  12. Hamel HusainXAI score12

    Similarity metrics are limited for evaluating LLM outputs

    AIHamel Husain argues that similar wording does not show whether an LLM answer works for a specific application, so teams should check concrete failure modes. He adds that similarity metrics can still help with retrieval evaluation and measuring output diversity.

    Image from @HamelHusain's post
  13. Simon WillisonXAI score36

    Mistral's Pelican SVG Test Passes, Tied to Mistral Large 4 Context

    AISimon Willison reports that Mistral can now generate his pelican SVG test, shared via a Markdown SVG renderer. The post links to a rendered result but gives no benchmark or scoring details. Background from Mistral's own announcement describes Mistral Large 4 as a 1T-parameter, natively multimodal model with 49B active parameters, available via API today and with open weights planned for end of October.

    Image from @simonw's post
  14. Aravind SrinivasXAI score42

    Perplexity Computer plays real-time StarCraft against itself, Blue wins 2-5

    AIPerplexity's Computer ran two agents playing StarCraft against each other in real time, with the game never paused while each agent thought. Blue, playing with 41 Dragoons, lost the final match 2-5 to Red, which used High Templar and Psionic Storm after Blue failed to scout Red's build. Each agent received only its own fog-of-war-limited game state, and video input was not provided.

    Video from @AravSrinivas's post
  15. Nathan LambertXAI score62

    Mistral Large 4 announced as a 1T-parameter multimodal open-weights model

    AIMistral has announced Mistral Large 4, a natively multimodal model with 1T total parameters and 49B active parameters. The company says it is the best open-weights model from the US or Europe on aggregated benchmarks, and that it is available via API today, with open weights due at the end of October.

  16. ARC PrizeOfficialAI score14

    Alexis Fox to speak at ARC Prize Research Summit 2026

    AIDuke University researcher Alexis Fox has been announced as a speaker at the ARC Prize Research Summit 2026. Fox studies how AI can reason over longer horizons and is lead author of PRO-LONG, which uses programmatic memory to support long-horizon reasoning.

    Image from @arcprize's post
  17. ElevenLabsOfficialAI score49

    ElevenLabs' Eleven v4 and v4 Turbo top Artificial Analysis TTS leaderboard

    AIElevenLabs' Eleven v4 and Eleven v4 Turbo rank #1 and #2 on the Artificial Analysis text-to-speech leaderboard. Per Artificial Analysis, Eleven v4 Turbo leads the Provider Voice arena at an Elo of 1,334, costs $40 per 1M characters, and generates 96 characters per second.

  18. Sophia YangXAI score45

    Mistral Large 4 tops benchmarks across cybersecurity, legal, and agentic tasks

    AIMistral Large 4 is a 1T-parameter natively multimodal model with 49B active parameters, which the Mistral account says leads open-weights models from the US or Europe on aggregated benchmarks. The post claims it beats closed frontier models on visual grounding and posts strong results across cybersecurity, legal, and agentic behavior. It is available via API now, with open weights due at the end of October.

    Image from @sophiamyang's post
  19. Arthur MenschXAI score48

    Mistral Large 4 trained on own compute, RL shows no saturation

    AIArthur Mensch says Mistral trained its model on its own compute, and reinforcement learning shows no sign of saturating. The post accompanies Mistral's announcement of Mistral Large 4, a 1T-parameter natively multimodal model with 49B active parameters, available via API today and with open weights planned for end of October.

  20. Guillaume Lample @ NeurIPS 2024XAI score26

    Mistral model beats GLM 5.3 on STEM, CAD, and finance tasks

    AIOn human evaluation, the model outperforms GLM 5.3 on STEM, CAD, and finance tasks and performs on par on agentic coding. The post is part 5 of a thread, so the model's name and other details come from earlier posts not included here.

    Image from @GuillaumeLample's post
  21. Guillaume Lample @ NeurIPS 2024XAI score40

    Mistral's ML4 hits open-model SOTA across capabilities and cyber benchmarks

    AIMistral says its ML4 model reaches state-of-the-art performance among open models across a wide range of capabilities, and outperforms the best models in visual grounding, legal, and spreadsheet manipulation. The post reports ML4 ranks among the best on the AA Cyber Index, scoring 82% on vulnerability reproduction and patching and 93% on Cybench. It argues that self-hosted, auditable open models are the best defense option for enterprises today, and that they do not refuse to help.

    Image from @GuillaumeLample's post
  22. Guillaume Lample @ NeurIPS 2024XAI score42

    Mistral's ML4 matches top open-weight models on coding and agentic benchmarks

    AIMistral's ML4 model matches the best open-weight models on DeepSWE, AutomationBench, and AA-Briefcase, and reaches state-of-the-art results on finance and legal workflows and complex multimodal grounding benchmarks. The post says it can navigate terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks.

    Image from @GuillaumeLample's post
  23. Guillaume Lample @ NeurIPS 2024XAI score78

    Mistral launches Large 4 preview with 1T parameters and open weights due October

    AIMistral has launched a preview of Mistral Large 4 (ML4), a 1T-parameter multimodal model with 49B active parameters. The company says it is the strongest open-weight model from the US or Europe on aggregated benchmarks and is available via API now, with open weights planned for the end of October.

    Why it matters: The post gives parameter counts, a preview timeline, and an open-weights release date, which help readers judge how Mistral's model compares with other open-weight options.

    Image from @GuillaumeLample's post
  24. Mistral AIOfficialAI score80

    Mistral Large 4 launches as a public preview with weights due end of month

    AIMistral AI launched a public preview API for Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters, and says it will release the weights by the end of the month. The company reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4, and 59.9% on AutomationBench. The model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's datacenters in Europe.

    Why it matters: The post gives benchmark figures and a weights timeline for an open-weight model, letting readers compare it with other open models and judge its access terms.

  25. The SequenceBlogAI score62

    Darwin Gödel Machine rewrote its own scaffolding to raise SWE-bench scores

    AIThe Darwin Gödel Machine, a coding agent from Sakana and Jeff Clune's lab, modified its own codebase over roughly eighty iterations without supervision. Its additions included better file viewing, patch validation before submitting fixes, generating and ranking several candidate solutions, and keeping a history of failed attempts. These changes raised its score from 20 to 50 percent on SWE-bench and from 14 to 31 percent on Polyglot.

  26. KhazixXAI score20

    AIHOT adds Pareto frontier view for model price-performance ranking

    AIAIHOT launched a Pareto frontier view on its leaderboard to compare models by price-performance, accessible via the top-right toggle. The cost calculation varies by scenario: general and office use assumes cache-hit:input:output at 7:2:1, programming at 97:2:1, research at input:output 1:3, and knowledge Q&A at 3:1 without caching.

    Image from @Khazix0918's post
  27. Black Forest LabsOfficialAI score38

    FLUX 3 tops Physics-IQ benchmark for video physical understanding

    AIBlack Forest Labs says its FLUX 3 model ranks first on Google DeepMind's Physics-IQ benchmark, which tests whether video models can predict what happens next in real filmed physical experiments. The company says FLUX 3 outperforms Seedance 2.5, MiniMax H3, Gemini Omni 1.1 Flash, Veo 3.1, Sora 2, and Cosmos3 in most cases, and that pairing it with a physics verification layer scores even higher.

    Image from @bfl_ai's post
  28. Latent SpaceBlogAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    AIReflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

  29. Artificial Analysis ArticlesOfficialAI score54

    Mistral Large 4 Preview scores 38 on Artificial Analysis Intelligence Index

    AIMistral has released Mistral Large 4 in Research Public Preview, with open weights for the 1T parameter (49B active) model planned for the end of October. It scores 38 on the Artificial Analysis Intelligence Index, comparable to GPT-6 Luna (max, 38) and DeepSeek V4.1 Flash (max, 39), and 50 on the Cyber Index. The source calls it the most intelligent model from outside the US and China, and notes costs of $1.13 per Intelligence Index task at standard pricing.

  30. Anthropic NewsroomOfficialAI score75

    Anthropic expands Cyber Verification Program into three tiered access levels

    AIAnthropic is launching an expanded Cyber Verification Program with three access tiers for qualifying security professionals, giving each tier different cyber capabilities and reduced blocking classifiers. On CyScenarioBench, Claude Opus 5.5 was blocked on 46 of 50 trials in the Defense Access tier, while the Red Team Access tier had no blocks and completed 34 of 50 tasks. Existing Project Glasswing members will move to the Specialized Access tier, and data retention is required for enrolled organizations.

    Why it matters: The program lays out three verified access tiers with different cyber blocks, and its CyScenarioBench figures show how safeguards change what defenders can do.

Oct 5

Oct 5Mon
  1. IThome · AINewsAI score49

    Reflection AI releases open-weight Beam model to rival DeepSeek and Kimi

    AIReflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.

  2. Goodfire ResearchOfficialAI score62

    Goodfire finds activation probes can detect reward hacking in open-source models

    AIGoodfire Research reports that reward hacking appears in 50–96% of rollouts across three open-source models on three agentic benchmarks. The team found an internal signal tied to cheating and gaming a metric, and simple activation probes catch some hacks that LLM chain-of-thought monitors miss. A probe can screen every transcript cheaply, and in one setup cut LLM monitoring cost by 90% with a roughly 1% precision drop.

    Why it matters: The study links a reward hacking signal in model activations to monitoring cost and detection, showing how probes compare with chain-of-thought monitors on the same runs.

  3. meng shaoXAI score47

    Reflection previews Beam, a 501B-parameter open agentic model

    AIReflection AI previewed Beam, an MoE open model with 501B total and 23B active parameters, claiming 3–4x better inference efficiency than GLM 5.2. The model was pretrained from scratch on 23.8T tokens in four weeks, and its RL run used 10,500 GB300 GPUs over four weeks, which the post describes as possibly the largest publicly recorded. Reflection positions Beam as a workhorse open model for enterprises, governments, and developers, with full weights due this month.

    Image from @shao__meng's post