Skip to content
Read the original: Artificial Ignorance· 46/100AI score46/100

Build Your Own Benchmark: Why Public AI Evals Are Saturating and What Replaces Them

Original titleBYOB: Build Your Own Benchmark

AISummary

Public AI benchmarks such as MMLU, SWE-bench Verified, and GPQA Diamond are saturating or showing contamination, prompting OpenAI to call SWE-bench Verified "no longer suitable" in late February and recommend SWE-bench Pro.

OpenAI's audit found 59.4% of the problems its best model failed had flawed test cases, and GPT-5.2, Claude Opus 4.5, and Gemini 3 Flash could reproduce original fixes from memory.

The article argues that behavioral tests, such as Vending-Bench's simulated vending machine business, may be more useful for everyday model choice.

Read the original ignorance.ai

Source: Artificial Ignorance · ignorance.aiPublished · added here