Skip to content
Epoch AI· Kelly Hong·· YesterdayPickAI score67

Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

Can AI automate Epoch?

AI summary

Epoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

Why it matters

The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

Source: Epoch AI · epoch.ai