Skip to contentSkip to stories

Updated

#Safety/Alignment

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. Goodfire ResearchAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  2. Arena.aiAI score55

    Arena raises $200M Series B and launches Alignment Index for AI agents

    AIArena announced a $200M Series B at a $3.1B valuation and released its Alignment Index, a benchmark measuring agent safety and alignment. The index is built from 90K+ real-world agent sessions across 27 models and tracks Unauthorized Action, False Attribution, and Deceptive Completion. OpenAI's GPT-6.1-Sol leads with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7.

    Video from @arena's post
  3. SantiagoAI score22

    Agent platform maps vulnerabilities and attack paths to protect systems

    AIA security platform uses agents to map a system's potential vulnerabilities and identify routes an attacker could take to reach sensitive data. It then recommends changes to close those paths. The quoted post cites a 700-agent swarm that breached Hugging Face with over 17,000 actions, and presents this tool, Cogent Attack Path Analysis, as the defensive counterpart.

  4. Philipp SchmidAI score46

    SynthID Detector now publicly available for verifying AI-generated content

    AIGoogle's SynthID Detector is now publicly available, letting users check whether an image, video, or audio file was generated by supported tools. Per the post, it scans for watermarks from Google and partners, including Nano Banana 2.1, OpenAI, NVIDIA, and Kakao, with Apple support coming soon. Uploaded files are deleted right after scanning.

    Video from @_philschmid's post
  5. TransformerAI score53

    Yoshua Bengio urges AI researchers to leave frontier labs for safety work

    AIYoshua Bengio, co-president of LawZero, asks researchers at frontier AI companies to reconsider whether they should keep working there, arguing that safety efforts are not slowing a dangerous race. He cites the recent UN Security Council briefing on AI incidents and says he left his earlier research path after ChatGPT made the risks feel immediate. He urges researchers to join AI Safety Institutes or mission-driven organizations such as LawZero.

  6. The Guardian · AIAI score42

    One Nation's AI-generated campaign video draws criticism over racist tropes and regulatory gaps

    AIOne Nation's AI-generated campaign video, reportedly played at its Victorian campaign launch, depicts racist stereotypes including a man brandishing a machete and a man in an explosive vest. The Australian Communications and Media Authority cannot act against it because its powers do not cover this content, and the federal Labor government has not yet moved to ban AI-generated content in election periods.

  7. The DecoderAI score46

    Ethereum researchers warn AI math advances could threaten crypto wallet signatures

    AIEthereum researcher Justin Drake warned on X that AI-assisted math could, in the worst case, break the signature system used by crypto wallets within months, and urged a "bunker mode" in which users move funds to addresses that have never signed a transaction. Vitalik Buterin agreed but cautioned against moving too fast, saying he has lost more money to botched migrations than to hacks. No one has yet broken the current ECDSA signature scheme in practice.

  8. The DecoderAI score72

    One public AI agent on AWS could take over every other agent in its region

    AIZenity Labs says a single publicly accessible agent on Amazon Bedrock AgentCore could take over all AgentCore agents in the same AWS account and region. A chat prompt let the researchers query the instance metadata service and steal temporary credentials, and AgentCore's default permissions allowed read, write, and delete access across agents. According to Zenity, AWS made IMDSv2 the default for new deployments and changed the default execution role around August.

    Why it matters: The report traces how one public agent's weak isolation exposed credentials and every other agent in the region, showing why default permissions matter for enterprise deployments.

  9. Ars Technica · AIAI score38

    Nvidia's Halos safety platform extends from robotaxis to humanoid and warehouse robots

    AINvidia's Halos software platform, originally built for autonomous vehicles, has been adapted for robotics, according to Ars Technica. The system monitors hardware and software for failures, isolates safety-critical workloads, and includes simulation tools and an inspection lab for robotics developers. Because safety requirements vary widely between a robotic vacuum and a warehouse forklift, Nvidia made the platform programmable so developers can define custom safety functions.

  10. Semafor · TechnologyAI score40

    Japanese and South Korean firms hit by major cyberattacks amid AI hacking fears

    AICompanies in Japan and South Korea were hit by major cyberattacks that exposed millions of customer records. The revelations follow reports that Chinese and US models were used to steal hundreds of thousands of credit card details, which one analyst called among the most severe AI-enabled exploitation abuses on record.

  11. The Guardian · AIAI score36

    Altman Says AI Will Cause 'Bad Things' as Columnist Cites Deaths and Lawsuits

    AIOpenAI CEO Sam Altman told Politico that the world should accept some bad things from AI for its benefits, a stance columnist Moustafa Bayoumi calls problematic. The column cites lawsuits over ChatGPT-linked suicides, a February strike on a Minab school that killed at least 120 children with a US military AI system (Palantir's Maven) implicated, and a chatbot error that nearly triggered a military interception.

  12. The DecoderAI score34

    Teen Hiker Needs Helicopter Rescue After Following Claude's Route Advice

    AIA 16-year-old hiker had to be airlifted from a dangerous rock face on Crown Mountain near Vancouver after using Anthropic's Claude to plan a route to the summit. He ended up on the Widowmaker Arete, a steep cliff requiring climbing gear, and called police when he got stuck on a ledge. Rescue manager Paul Markey said Claude has no actual knowledge of locations or terrain and is no substitute for experience and common sense.

  13. Wired · AIAI score36

    Tristan Harris's Center for Humane Technology lays off about half its staff

    AIThe Center for Humane Technology is laying off about half of its 16 non-founder employees and ending its policy research and litigation work. The organization will refocus on "founder-led" initiatives built around cofounder Tristan Harris, according to WIRED, after its board concluded that operating as both an advocacy group and a think tank had stretched it too thin.

  14. The DecoderAI score72

    AI hacking tools let a likely single attacker breach multiple South Korean banks

    AIA suspected Chinese-speaking attacker breached several South Korean financial institutions between late September and early October 2026, reportedly stealing over 25,000 records from Shinhan Bank alone. The attacker used ARTEX, a Chinese open-source tool that uses AI language models to automate finding security flaws, and models named in the report include DeepSeek v4.1-flash, GLM-5.3, and Grok 4.6.

  15. MIT Technology Review · AIAI score26

    AVEVA's Arti Garg outlines a safer path to autonomous industrial AI

    AIAVEVA chief technologist Arti Garg argues industrial AI should augment rather than replace human supervisors in critical decisions, with guardrails defining where automated systems can act. She says organizations must rethink business processes and safeguards as foundation models, physical AI, and agentic AI enable more complex automation.

  16. LeiphoneAI score14

    Negative Transfer in AI: Four Root-Cause Mechanisms Defined in a Chinese Governance Series

    AIThis second installment of the Carbon-Silicon Dao Code series defines four types of negative transfer in cross-domain AI: NT1 mechanism mismatch, NT2 semantic drift, NT3 unknown completion, and NT4 power leakage. It argues that current evaluation based on fit accuracy and test-set pass rates cannot detect whether the underlying mechanisms match. The article is a Chinese-language theoretical and governance piece, and the summary covers only the framework it presents, not empirical results.

  17. LeiphoneAI score15

    Chinese Legal-Style Framework Outlines Seven-Layer System for Cross-Domain AI Transfer Governance

    AILeiphone publishes the table of contents for "Carbon-Silicon Dao Code: Cross-Domain Transfer Governance Code," a seven-layer framework covering 188 numbered chapters. The outline spans transfer accident analysis, technical mechanisms, rights assignment, industry governance, top-level regulation, civilization-scale risk control, and final codification, with a baseline entry labeled NT1–NT4 negative-transfer categories.

  18. Air Street PressAI score60

    Nathan Benaich's 2026 State of AI Report covers agents, robotics, and AI control

    AINathan Benaich's 9th annual State of AI Report covers agents, robotics, AI for science, inference economics, and government control over frontier AI access. The report also records a 2025 prediction scorecard and lists nine predictions for the next 12 months. It cites an OpenAI cyber evaluation in which agents compromised Hugging Face's production infrastructure, and it says Anthropic and OpenAI's combined annualized revenue run rate reached $105B by late summer.

  19. Anthropic NewsroomAI score62

    Anthropic launches Cyber Mission with infrastructure defense and free OSS Scanner

    AIAnthropic has launched the Anthropic Cyber Mission, which starts with the Critical Infrastructure Defense Program for operational technology and OSS Scanner for open-source projects. The defense program brings frontier Claude models, on-site engineers and threat research to trusted providers such as Accenture, CrowdStrike and Palo Alto Networks. OSS Scanner gives enrolled open-source projects periodic free scans from its strongest models, with reports sent without human review and an expected true-positive rate above 90%.

    Why it matters: The announcement shows how a frontier AI lab is packaging cyber defense around critical infrastructure and open-source maintainers, including the program's partners and access routes.

  20. Anthropic NewsroomAI score46

    Anthropic Updates Claude Usage Policy, Effective November 12, 2026

    AIAnthropic has published a 2026 update to its Usage Policy, taking effect November 12, mostly to clarify existing rules for longer, more autonomous Claude work. The changes consolidate deceptive-campaign prohibitions into a new section, narrow the elections rules to voter deception and disruption, and explicitly ban weapons-related software and surveillance tools. Requirements for high-risk uses and for models connected to autonomous physical hardware were also tightened.

  21. Artificial Analysis ArticlesAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    AIArtificial Analysis is adding trusted-access models to its Cyber Index, starting with GPT-6 Sol (Daybreak Blue, max), which is available only through OpenAI's Daybreak program. The model hits no safety blocks across the Index and scores 32 points higher overall than the publicly available GPT-6 Sol (max), with its largest gains on CyberGym-E2E.

    Why it matters: The source shows how safety refusals shape cyber benchmark scores, with the trusted-access model's gains concentrated on CyberGym-E2E, useful for comparing guarded and unguarded models.