Skip to content
Read the original: Hugging Face Blog·Published· 5d agoPickAI score67

Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

The Agent Said It Was Done. The Database Disagreed.

AISummary

Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

AIWhy it matters

The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

Read the original huggingface.co

Source: Hugging Face Blog · huggingface.co