Hamel Husain Tests Anthropic's Claude Eval Plugin on Leasing Assistant Traces
Original titleClaude’s new auto eval tool
AISummary
Hamel Husain reviewed Anthropic's new build_eval and hill-climb commands in the claude-api plugin for Claude Code, finding it useful for discovering issues like human handoff, formatting, and voice agent problems.
He criticized it for pushing evaluator creation before data review, asking for label validation in Markdown files, and bundling four failure checks into one broad call-transfer evaluator. Husain says he would hold off on using it for now.
Source: Hamel Husain · hamel.dev