Skip to contentSkip to stories
Updated

#Safety/Alignment

Oct 6

  1. Claude BlogAI score62

    Comcast and Booz Allen use Claude Mythos to find exploit chains in codebases

    AIComcast and Booz Allen used Claude Mythos Preview to find vulnerabilities that arise from interactions across code, configuration, and deployment rather than single-file bugs. Comcast identified a critical authentication flaw across 258 systems and about 170 million lines of code before any exploitation was observed. Booz Allen reported that one analyst reviewed eight production systems across 138 repositories in twelve days, a review its team estimated would have taken several months without the model.

    Why it matters: The case studies show how security teams validate and remediate model-found exploit chains, a workflow relevant to anyone managing large codebases.

Sep 14

  1. Google Developers BlogAI score60

    Build zero-trust AI agents that judge intent, not just syntax

    AIPart 2 of the zero-trust agents series moves security checks from agent code to the Gemini Enterprise Agent Platform runtime. Model Armor screens prompts and responses, Semantic Governance Policies judge proposed tool calls against intent and business rules, and Agent Anomaly Detection flags multi-turn drainage that single-turn checks miss. The same Customer Support and Returns Agent from Part 1 is used, with the companion demo open-sourced on GitHub.

    Why it matters: The post walks through a concrete refund agent under four attacks, showing how screening, intent judgment, and anomaly detection each catch what the others miss.

Jun 17

  1. PromptArmor Threat IntelligenceAI score62

    PromptArmor shows Codex auto-review agent approved malware install via prompt injection

    AIPromptArmor demonstrated that OpenAI's Approve-for-me agent approved a malicious NPM install with elevated privileges after a hidden prompt injection in an external GitHub issue influenced the main Codex agent. The malicious package's post-install script then ran unsandboxed with the user's full privileges. The report also gives steps for organizations to disable agentic auto-review in Claude Code and Codex.

    Why it matters: The report shows a prompt-injected GitHub issue leading an approval agent to permit a malicious NPM install, a concrete test of agent-in-the-loop guardrails.

Mar 24

  1. Anthropic EngineeringAI score78

    How Anthropic built Claude Code auto mode to replace skipped permissions

    AIAnthropic describes Claude Code auto mode, which delegates approval of agent actions to model-based classifiers instead of manual prompts or skipped permissions. The classifier reviews tool calls before execution and a separate probe screens tool outputs for prompt injection. Anthropic reports a 0.4% false positive rate on real internal traffic and a 17% false negative rate on real overeager actions.

    Why it matters: The post explains the layered classifier design and its measured tradeoffs, showing how autonomous coding agents can cut approval fatigue without fully removing risk.

That’s everything