Skip to content
Read the original: Import AI· Published 44/100AI score44/100

DiG-bench Tests AI Rule Discovery as Opus 5 and Fable 5 Lead

Original titleImport AI 469: Science AI; RSI simulator; and Zuck's technological pessimism

AISummary

DiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.

Read the original importai.substack.com

Source: Import AI · importai.substack.comPublished · added here