Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. OpenBMBAI score36

    ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%

    AIReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.

    Image from @OpenBMB's post
  2. Understanding AI (Timothy B. Lee)AI score67

    TypeSafe AI's Jev returns probabilities over fixed answers instead of text

    AITypeSafe AI released Jev, a model that answers yes/no, multiple-choice, or rating questions by outputting the estimated probability of each option. The author notes this design lets the model be served faster and more cheaply than LLMs and fits ordinary if-statement logic, and says he used it to flag spam comments on his blog in place of Gemini 3 Flash.

  3. OpenRouter · New modelsAI score54

    StepFun releases Step 5 Preview, a 600B-parameter agentic model

    AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.

  4. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  5. The DecoderAI score72

    Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna

    AIAnthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.

  6. QbitAIAI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    AIUniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  7. MarkTechPostAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    AINVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

  8. meng shaoAI score39

    Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output

    AIAnthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.

  9. Artificial Analysis ArticlesAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    AIArtificial Analysis is adding trusted-access models to its Cyber Index, starting with GPT-6 Sol (Daybreak Blue, max), which is available only through OpenAI's Daybreak program. The model hits no safety blocks across the Index and scores 32 points higher overall than the publicly available GPT-6 Sol (max), with its largest gains on CyberGym-E2E.

    Why it matters: The source shows how safety refusals shape cyber benchmark scores, with the trusted-access model's gains concentrated on CyberGym-E2E, useful for comparing guarded and unguarded models.

  10. Artificial Analysis ArticlesAI score50

    Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

    AIHarvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

Oct 7

Oct 7Wed
  1. KhazixAI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    AIOpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. François CholletAI score44

    Chollet: Programming and math training don't boost general intelligence

    AIFrançois Chollet compares AI progress to human learning, noting that 1980s research found programming training improves coding but does not transfer to general reasoning. He argues general intelligence is a fundamental brain property rather than a trainable skill, since domain practice improves only that domain. The post is framed as background for his question whether AI's jagged frontier, driven by math and code via RLVR, reflects general capability or continued human-data bottlenecks.

  3. Waymo BlogAI score42

    Sober Drivers Still Face Nearly 4x Nighttime Fatal Crash Risk, Waymo Study Finds

    AIWaymo research found that even fully sober human drivers face nighttime fatal crash risk 3.1 to 3.9 times higher than daytime risk, pointing to systemic hazards beyond impairment. The study used an exposure reconstruction model across the 50 most populous U.S. urban areas, showing removing alcohol-involved drivers lowers the average urban fatal crash rate by 23%, from 1.42 to 1.10 per 100 million miles.

  4. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  5. Google Developers BlogAI score62

    Google's AQuA agent diagnoses production failures in a multi-agent travel concierge

    AIGoogle Developers Blog introduces AQuA, an ambient quality agent that runs in a customer's Google Cloud project and samples production sessions to find recurring agent failures. In a 32-session travel-concierge sweep, it verified six issues and traced two of them to specific prompt lines, and a replay after the fixes raised full-session passes from 5/32 to 13/32. The post notes that verification and diagnosis are model-based, and that the tool proposes edits without applying them.

    Why it matters: The post walks through a concrete production workflow, from sweep and verification to a code-anchored fix and replay, that shows how to diagnose silent agent failures.

  6. Google ResearchAI score23

    Google Research invites COLM visitors to ContinuousBench walkthrough on DP synthetic data

    AIGoogle Research is hosting a walkthrough at its COLM booth #107 today at 5:00 PM of ContinuousBench, a standardized benchmark for measuring knowledge transfer in differentially private synthetic data. The session, led by Alex Bie, asks whether DP synthetic data preserve actual information or only style. A paper is linked on arXiv.

    Image from @GoogleResearch's post
  7. IThome · AIAI score72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    AIAnthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.

  8. TypeSafe AIAI score25

    Jev-killer OpenAI Decisions API benchmarked against Jev for HiringCafe

    AIThe main post is a short reply saying reports of a company's death have been greatly exaggerated, with no details about products or figures. The background post from @h_nilforoshan reports that OpenAI's Decisions API, billed as a "Jev-killer," was benchmarked against Jev for HiringCafe, which serves 2.5 million users. On the task of scoring job-description relevance from 1 to 10, the author reports OpenAI costing 2x more and performing 5-10% worse.

  9. Leandro von WerraAI score36

    Snorkel expands Open Benchmarks Grants to $30M for AI evaluation

    AISnorkel AI is expanding its Open Benchmarks Grants tenfold to a $30M commitment to fund more diverse, robust, and continuously updated open AI benchmarks. The program adds an Open Benchmarks Red Team to test and strengthen those benchmarks, plus a Snorkel Research Fellowship for independent researchers developing new evaluation methods. The source says OBG-funded benchmarks have appeared on model cards from every major frontier lab.

  10. Simon WillisonAI score62

    Anthropic releases Claude Haiku 5.5, priced like GPT-6 Luna up to 100,000 tokens

    AIAnthropic has released Claude Haiku 5.5, priced at $0.10 input and $0.50 output per million tokens up to 100,000 tokens, matching GPT-6 Luna. Beyond 100,000 tokens the price rises to $0.50 and $2.50, and the author found the new tokenizer uses about 1.25x as many tokens as Haiku 4.5 on the same long prompt. The model cannot disable reasoning and defaults to medium effort.

  11. MarkTechPostAI score67

    Anthropic releases Claude Haiku 5.5, a small model with 1M context

    AIAnthropic has released Claude Haiku 5.5, its cheapest and fastest small model, priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. It keeps a 1M token context window, up to 128K output tokens, and is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic reports 72.4% on OSWorld 2.1 (offline subset) versus 15.7% for Haiku 4.5, and the article notes that non-default temperature, top_p or top_k values return a 400 error.

  12. ChatGPTAI score85

    GPT-6 with Intelligent UI rolls out globally in ChatGPT, Free and Go tiers next day

    AIGPT-6 with Intelligent UI begins rolling out globally in the ChatGPT Chat tab for Plus, Pro, Business, and Enterprise users today. The rollout expands to Free and Go tiers starting tomorrow. Plus, Pro, Business, and Enterprise get GPT-6 Sol, while Free and Go get GPT-6 Luna.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”