Skip to content
Trending storyDeveloping

Goodfire's activation-based AI monitors, offered to Baseten customers via Project Beacon

9 articles4 sourcessince Oct 8Last article 3h ago ·

Overview

AISummary of 9 articles

Goodfire Research has built monitors that read a model's internal activations rather than only its written output, and Baseten is now offering them through Project Beacon, a joint effort to add inline safety controls that can flag unsafe events during generation.

Customers can choose which risks to watch, such as offensive hacking, chemical and biological weapons misuse, reward hacking, prompt injection, or sensitive data exposure, and set automated responses like logging, extra review, refusal, or re-routing.

Goodfire's performance claims are its own and unverified. It says a cheap probe screening sessions before an LLM judge reaches about 93% recall at a 5.5% benign-interruption rate and roughly 50x lower judge cost. On Kimi K3, it says monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, catching 94% of malicious hacking sessions. Goodfire says red-teaming by FAR.AI, the one independent element cited, reduced universal jailbreaks to zero across a static battery of 140 attacks, though that testing used fixed attacks.

Baseten says the capabilities will roll out over several months, starting with selected models and early partners.

Written by AI from the articles below · updated Oct 9, 3:39 PM ET

Check the sources:

Developments

7 developments

  1. Oct 9, 3:23 PM ET · 1 article
    Goodfire post on configurable responses to flagged concerns in AI applications
    Goodfire: Goodfire and Baseten partner on configurable model concern monitoring
  2. Oct 9, 3:23 PM ET · 1 article
    Goodfire launches activation monitors for detecting undesired model behavior on Baseten
    Goodfire: Goodfire launches activation monitors that detect undesired model behaviors
  3. Oct 9, 1:01 PM ET · 2 articles
    Baseten announces Project Beacon with Goodfire AI for inline safety controls in model inference
    Baseten Blog: Baseten launches Project Beacon with Goodfire AI for inline safety controls on models
  4. Oct 8, 12:29 PM ET · 1 article
    Goodfire reports its monitor cut universal jailbreaks to zero in Far AI red-team testing
    Goodfire: Goodfire monitors cut universal jailbreaks to zero in red-team test
  5. Oct 8, 12:29 PM ET · 1 article
    Goodfire releases cybersecurity monitors for Kimi K3 and GLM 5.3
    Goodfire: Goodfire builds cybersecurity monitors for Kimi K3 and GLM 5.3
  6. Oct 8, 12:14 PM ET · 2 articles
    Goodfire deploys probe-based cyber monitors for Kimi K3 and GLM 5.3 on a production inference stack
    Goodfire Research: Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade
  7. Oct 8, 12:00 PM ET · 1 article
    Goodfire launches inside-out activation monitors for AI agents via Baseten
    TechCrunch · AI: Goodfire launches inside-out monitors to catch rogue AI agents at lower cost

Article timeline

The articles in this story. Times are ET.

Oct 9
  1. GoodfireOfficial
    Goodfire and Baseten partner on configurable model concern monitoring

    AIGoodfire says teams can configure how their applications respond when a concern is flagged, including logging, additional review, refusal, and re-routing. The company directs model servers and trainers to a partnership post with Baseten for building monitors into their stack.

  2. GoodfireOfficial
    Goodfire launches activation monitors that detect undesired model behaviors

    AIGoodfire says its activation monitors use signals from inside a model to detect undesired behaviors, catching more cases, running faster and costing less than text-based monitors. Baseten customers can use them to monitor for prompt injection, actions outside policy, sensitive data exposure and cyber misuse.

  3. GoodfireOfficial
    Goodfire and Baseten partner on real-time safety monitoring for open models

    AIGoodfire announces a partnership with Baseten to bring real-time safety monitoring to open models. The companies are launching Project Beacon, which combines inference with in-line monitoring so safety signals are caught proactively.

  4. Baseten BlogOfficial
    Baseten launches Project Beacon with Goodfire AI for inline safety controls on models

    AIBaseten announces Project Beacon with Goodfire AI, adding inline safety controls to model inference. Goodfire's activation-based monitors read a model's internal activations during generation, so policies can flag unsafe events before output reaches a user or tool. Baseten plans to release the capabilities over the next several months with selected models and early partners.

Oct 8
  1. GoodfireOfficial
    Goodfire tests production cyber monitors on Kimi K3

    AIGoodfire has published research on deploying cyber monitors in production on Kimi K3. The post links to the full research write-up and invites organizations to get in touch about deploying these monitors.

  2. GoodfireOfficial
    Goodfire monitors cut universal jailbreaks to zero in red-team test

    AIGoodfire says a static battery of attacks from FAR AI red-teamed its monitor, and the monitors reduced successful universal jailbreaks to 0. The same testing cut total jailbroken interactions by 97%.

    Image from @GoodfireAI's post
  3. GoodfireOfficial
    Goodfire builds cybersecurity monitors for Kimi K3 and GLM 5.3

    AIGoodfire built cybersecurity monitors for Kimi K3 and GLM 5.3 that it says are more accurate, 50x faster, and 50x cheaper than an optimized LLM judge. External red-teaming by Far AI Research reportedly found the monitor greatly reduces universal jailbreaks.

    Video from @GoodfireAI's post
  4. Goodfire ResearchOfficial
    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  5. TechCrunch · AINews
    Goodfire launches inside-out monitors to catch rogue AI agents at lower cost

    AIGoodfire has launched monitors that read a model's internal signals during agent work instead of reviewing its written output. The monitors are available to Baseten customers, who can choose risks to watch and set automated responses. In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, while catching 94% of malicious hacking sessions.

Heat trend

Not enough continuous observations to show a trend yet.