DiG-bench Tests AI Rule Discovery as Opus 5 and Fable 5 Lead
Original titleImport AI 469: Science AI; RSI simulator; and Zuck's technological pessimism
AISummary
DiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.
Source: Import AI · importai.substack.comPublished · added here