Goodfire's activation-based AI monitors, offered to Baseten customers via Project Beacon
Overview
Goodfire Research has built monitors that read a model's internal activations rather than only its written output, and Baseten is now offering them through Project Beacon, a joint effort to add inline safety controls that can flag unsafe events during generation.
Customers can choose which risks to watch, such as offensive hacking, chemical and biological weapons misuse, reward hacking, prompt injection, or sensitive data exposure, and set automated responses like logging, extra review, refusal, or re-routing.
Goodfire's performance claims are its own and unverified. It says a cheap probe screening sessions before an LLM judge reaches about 93% recall at a 5.5% benign-interruption rate and roughly 50x lower judge cost. On Kimi K3, it says monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, catching 94% of malicious hacking sessions. Goodfire says red-teaming by FAR.AI, the one independent element cited, reduced universal jailbreaks to zero across a static battery of 140 attacks, though that testing used fixed attacks.
Baseten says the capabilities will roll out over several months, starting with selected models and early partners.
Written by AI from the articles below · updated Oct 9, 3:39 PM ET
Check the sources:
Developments
7 developments
- Oct 9, 3:23 PM ET · 1 articleGoodfire post on configurable responses to flagged concerns in AI applicationsGoodfire: Goodfire and Baseten partner on configurable model concern monitoring
- Oct 9, 3:23 PM ET · 1 articleGoodfire launches activation monitors for detecting undesired model behavior on BasetenGoodfire: Goodfire launches activation monitors that detect undesired model behaviors
- Oct 9, 1:01 PM ET · 2 articlesBaseten announces Project Beacon with Goodfire AI for inline safety controls in model inferenceBaseten Blog: Baseten launches Project Beacon with Goodfire AI for inline safety controls on models
- Oct 8, 12:29 PM ET · 1 articleGoodfire reports its monitor cut universal jailbreaks to zero in Far AI red-team testingGoodfire: Goodfire monitors cut universal jailbreaks to zero in red-team test
- Oct 8, 12:29 PM ET · 1 articleGoodfire releases cybersecurity monitors for Kimi K3 and GLM 5.3Goodfire: Goodfire builds cybersecurity monitors for Kimi K3 and GLM 5.3
- Oct 8, 12:14 PM ET · 2 articlesGoodfire deploys probe-based cyber monitors for Kimi K3 and GLM 5.3 on a production inference stackGoodfire Research: Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade
- Oct 8, 12:00 PM ET · 1 articleGoodfire launches inside-out activation monitors for AI agents via BasetenTechCrunch · AI: Goodfire launches inside-out monitors to catch rogue AI agents at lower cost
Article timeline
The articles in this story. Times are ET.
Goodfire@GoodfireAIOfficialGoodfire and Baseten partner on configurable model concern monitoringAIGoodfire says teams can configure how their applications respond when a concern is flagged, including logging, additional review, refusal, and re-routing. The company directs model servers and trainers to a partnership post with Baseten for building monitors into their stack.
Goodfire@GoodfireAIOfficialGoodfire launches activation monitors that detect undesired model behaviorsAIGoodfire says its activation monitors use signals from inside a model to detect undesired behaviors, catching more cases, running faster and costing less than text-based monitors. Baseten customers can use them to monitor for prompt injection, actions outside policy, sensitive data exposure and cyber misuse.
Goodfire@GoodfireAIOfficialGoodfire and Baseten partner on real-time safety monitoring for open modelsAIGoodfire announces a partnership with Baseten to bring real-time safety monitoring to open models. The companies are launching Project Beacon, which combines inference with in-line monitoring so safety signals are caught proactively.
- Baseten BlogOfficialBaseten launches Project Beacon with Goodfire AI for inline safety controls on models
AIBaseten announces Project Beacon with Goodfire AI, adding inline safety controls to model inference. Goodfire's activation-based monitors read a model's internal activations during generation, so policies can flag unsafe events before output reaches a user or tool. Baseten plans to release the capabilities over the next several months with selected models and early partners.
Goodfire@GoodfireAIOfficialGoodfire tests production cyber monitors on Kimi K3AIGoodfire has published research on deploying cyber monitors in production on Kimi K3. The post links to the full research write-up and invites organizations to get in touch about deploying these monitors.
Goodfire@GoodfireAIOfficialGoodfire monitors cut universal jailbreaks to zero in red-team testAIGoodfire says a static battery of attacks from FAR AI red-teamed its monitor, and the monitors reduced successful universal jailbreaks to 0. The same testing cut total jailbroken interactions by 97%.

Goodfire@GoodfireAIOfficialGoodfire builds cybersecurity monitors for Kimi K3 and GLM 5.3AIGoodfire built cybersecurity monitors for Kimi K3 and GLM 5.3 that it says are more accurate, 50x faster, and 50x cheaper than an optimized LLM judge. External red-teaming by Far AI Research reportedly found the monitor greatly reduces universal jailbreaks.

- Goodfire ResearchOfficialGoodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade
AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.
- TechCrunch · AINewsGoodfire launches inside-out monitors to catch rogue AI agents at lower cost
AIGoodfire has launched monitors that read a model's internal signals during agent work instead of reviewing its written output. The monitors are available to Baseten customers, who can choose risks to watch and set automated responses. In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, while catching 94% of malicious hacking sessions.
Heat trend
Not enough continuous observations to show a trend yet.