Skip to content
View original post on X: GoodfireOfficial· 36/100AI score36/100

Goodfire launches activation monitors that detect undesired model behaviors

AISummary

Goodfire says its activation monitors use signals from inside a model to detect undesired behaviors, catching more cases, running faster and costing less than text-based monitors. Baseten customers can use them to monitor for prompt injection, actions outside policy, sensitive data exposure and cyber misuse.

Post on XView on X
GoodfireVerified on X
@GoodfireAI

Part of a thread · earlier post

Our activation monitors detect undesired behaviors using signals from inside a model, so they catch more, run faster, and cost less than text-based monitors.

Baseten customers can monitor for:
- Prompt injection
- Actions outside policy
- Sensitive data exposure
- Cyber misuse

Source: Goodfire · x.comPublished · added here