Anthropic redesigns its performance engineering take-home as Claude models improve
Original titleDesigning AI-resistant technical evaluations
AISummary
Anthropic's performance engineering lead Tristan Hume describes how a take-home test for hiring performance engineers was repeatedly defeated by successive Claude models.
Claude Opus 4 outperformed most human applicants within the 4-hour limit, and Claude Opus 4.5 matched the best candidates in 2 hours. Anthropic is releasing the original take-home as an open challenge, with the best known Claude result at 1487 cycles.
AIWhy it matters
The post traces how each Claude model defeated the take-home test, showing concrete redesign tradeoffs for evaluating engineers when AI assistance is available.
Source: Anthropic Engineering · anthropic.comPublished · added here