Treat agent observability outputs as untrusted inputs to supervision
Original titleTreating agent observability as security-critical infrastructure means handling all AI outputs (transcripts, reasoning, actions, etc.) as...
AISummary
METR argues that agent observability should be treated as security-critical infrastructure, with all AI outputs, including transcripts, reasoning, and actions, handled as untrusted inputs. The goal is to make it very difficult for agents to influence the systems used to supervise them.
Source: METR · x.comPublished · added here