Skip to content

#Safety/Alignment

Oct 8

TodayOct 8Thu26 items
  1. Andrew CurranAI score28

    To become a member of the Association for Human Mathematics you must make three vows: 1. Members commit to not working with commercial AI companies by providing technical labor, mathematical knowledge, consultation, or publicity. 2. Members commit to not publishing mathematical texts (including papers, reviews, referee reports, lecture notes, teaching materials, and books) generated by AI models. 3. Members of the AI-free caucus do not use AI models in a research capacity.

    To become a member of the Association for Human Mathematics you must make three vows: 1. Members commit to not working with commercial AI companies by providing technical labor, mathematical knowledge, consultation, or publicity. 2. Members commit to not publishing mathematical texts (including papers, reviews, referee reports, lecture notes, teaching materials, and books) generated by AI models. 3. Members of the AI-free caucus do not use AI models in a research capacity.

  2. Tessl BlogAI score42

    Agent Skills Should Be Treated as Supply Chain Components

    Tessl's talk at AI Native DevCon London argues that agent skills, which can be markdown files with instructions and bundled material, act as supply chain components that can shape agent behavior. The author says reading SKILL.md once is insufficient because risks can sit in supporting files, updates, and workspace trust settings. He identifies the danger as the combination of private context, untrusted content, and external communication, and cites research scanning roughly 4,000 public skills for issues including malware-like behavior.

  3. SemiAnalysisAI score72

    SemiAnalysis Finds China's AI Safety Rules Target Applications, Not Frontier Models

    SemiAnalysis argues China's AI safety regime is speed-first, with rules covering content and public-facing services but no frontier-risk duties tied to training compute or capability. Its dataset of 857 releases from nine leading Chinese developers found only 31 (3.6%) ever had a published safety result, and just 9 at launch. The analysis also finds that technical experts favor binding frontier rules while the top leadership's development-first preference settled the policy debate.

  4. Miles BrundageAI score22

    I bet they're just doing that to help with the IPO and curry favor with the administration (which will love it) as part of a regulatory capture play https://x.com/AndrewCurran_/status/2108244808494154089?s=20

    I bet they're just doing that to help with the IPO and curry favor with the administration (which will love it) as part of a regulatory capture play https://x.com/AndrewCurran_/status/2108244808494154089?s=20

  5. Ethan MollickAI score9

    At the very start of my grad school in 2005, I wrote paper about the original computer hacking/phreaking/BBS scene: hackers were often driven by curiosity but the tools they made were widely exploited by "Script Kiddies" who caused most damage & chaos Anyhow, about AI hacking...

    At the very start of my grad school in 2005, I wrote paper about the original computer hacking/phreaking/BBS scene: hackers were often driven by curiosity but the tools they made were widely exploited by "Script Kiddies" who caused most damage & chaos Anyhow, about AI hacking...

  6. Microsoft ResearchAI score16

    The most important AI failures may not be the obvious ones. Microsoft Partner Research Manager Jennifer Neville explains to host Chad Atalla why “surprising failures” can reveal where human expectations about intelligence diverge from how AI systems actually work. https://msft.it/6017aU8r7

    The most important AI failures may not be the obvious ones. Microsoft Partner Research Manager Jennifer Neville explains to host Chad Atalla why “surprising failures” can reveal where human expectations about intelligence diverge from how AI systems actually work. https://msft.it/6017aU8r7

  7. TransformerAI score67

    Bengio urges safety-minded AI researchers to leave frontier labs

    Yoshua Bengio, co-president of LawZero, writes to researchers urging those who prioritize safety to leave frontier AI companies for safety institutes or mission-driven organizations. He argues that safety efforts at the labs are not sufficiently slowing a dangerous race toward recursive self-improvement, and cites LawZero's recent C$200 million-plus funding from Canada and Germany as an alternative path.

  8. The DecoderAI score46

    Ethereum researchers warn AI math advances could threaten crypto wallet signatures

    Ethereum researcher Justin Drake warned on X that AI-assisted math could, in the worst case, break the signature system used by crypto wallets within months, and urged a "bunker mode" in which users move funds to addresses that have never signed a transaction. Vitalik Buterin agreed but cautioned against moving too fast, saying he has lost more money to botched migrations than to hacks. No one has yet broken the current ECDSA signature scheme in practice.

  9. Clément DelangueAI score22

    We urgently need more public traces of AI agents attacking and defending systems. Defenders can’t learn from what they can’t see. If you have traces and are being pressured to keep them private, my DMs are open. Let’s level the playing field and fight the asymmetry and lack of transparency in AI!

    We urgently need more public traces of AI agents attacking and defending systems. Defenders can’t learn from what they can’t see. If you have traces and are being pressured to keep them private, my DMs are open. Let’s level the playing field and fight the asymmetry and lack of transparency in AI!

  10. Nathan LambertAI score18

    The South Korean bank hack fits with this. Yes, open models are being fitted to be used by some attackers, but so long as Claude Code can be the orchestrator of a hack we desperately need open models to diffuse defense as well. A hard balance societally to stomach, but needed.

    The South Korean bank hack fits with this. Yes, open models are being fitted to be used by some attackers, but so long as Claude Code can be the orchestrator of a hack we desperately need open models to diffuse defense as well. A hard balance societally to stomach, but needed.

  11. The Guardian · AIAI score36

    Altman Says AI Will Cause 'Bad Things' as Columnist Cites Deaths and Lawsuits

    OpenAI CEO Sam Altman told Politico that the world should accept some bad things from AI for its benefits, a stance columnist Moustafa Bayoumi calls problematic. The column cites lawsuits over ChatGPT-linked suicides, a February strike on a Minab school that killed at least 120 children with a US military AI system (Palantir's Maven) implicated, and a chatbot error that nearly triggered a military interception.

  12. Gergely OroszAI score48

    You can either hold crypto and keep being stressed out if a math breakthrough would drain your wallet; or someone stealing your keys would drain your wallet; or someone kidnapping you and forcing you to hand over your keys would drain your wallet Or you can just not hold crypto

    You can either hold crypto and keep being stressed out if a math breakthrough would drain your wallet; or someone stealing your keys would drain your wallet; or someone kidnapping you and forcing you to hand over your keys would drain your wallet Or you can just not hold crypto

  13. MIT Technology Review · AIAI score26

    AVEVA's Arti Garg outlines a safer path to autonomous industrial AI

    AVEVA chief technologist Arti Garg argues industrial AI should augment rather than replace human supervisors in critical decisions, with guardrails defining where automated systems can act. She says organizations must rethink business processes and safeguards as foundation models, physical AI, and agentic AI enable more complex automation.

  14. Leiphone (雷峰网)AI score14

    Negative Transfer in AI: Four Root-Cause Mechanisms Defined in a Chinese Governance Series

    This second installment of the Carbon-Silicon Dao Code series defines four types of negative transfer in cross-domain AI: NT1 mechanism mismatch, NT2 semantic drift, NT3 unknown completion, and NT4 power leakage. It argues that current evaluation based on fit accuracy and test-set pass rates cannot detect whether the underlying mechanisms match. The article is a Chinese-language theoretical and governance piece, and the summary covers only the framework it presents, not empirical results.

  15. Leiphone (雷峰网)AI score15

    Chinese Legal-Style Framework Outlines Seven-Layer System for Cross-Domain AI Transfer Governance

    Leiphone publishes the table of contents for "Carbon-Silicon Dao Code: Cross-Domain Transfer Governance Code," a seven-layer framework covering 188 numbered chapters. The outline spans transfer accident analysis, technical mechanisms, rights assignment, industry governance, top-level regulation, civilization-scale risk control, and final codification, with a baseline entry labeled NT1–NT4 negative-transfer categories.

  16. Air Street PressAI score60

    Nathan Benaich's 2026 State of AI Report covers agents, robotics, and AI control

    Nathan Benaich's 9th annual State of AI Report covers agents, robotics, AI for science, inference economics, and government control over frontier AI access. The report also records a 2025 prediction scorecard and lists nine predictions for the next 12 months. It cites an OpenAI cyber evaluation in which agents compromised Hugging Face's production infrastructure, and it says Anthropic and OpenAI's combined annualized revenue run rate reached $105B by late summer.

  17. DeedyAI score24

    “Cryptography is a subfield that’s extremely conspicuous by its absence from OpenAI’s list of 376 papers! But my sources tell me that the AI cos have now started, gingerly and discreetly, investigating whether their latest internal models can break important cryptographic protocols and primitives.” - Scott Aaronson, CS chair at UT Austin

    “Cryptography is a subfield that’s extremely conspicuous by its absence from OpenAI’s list of 376 papers! But my sources tell me that the AI cos have now started, gingerly and discreetly, investigating whether their latest internal models can break important cryptographic protocols and primitives.” - Scott Aaronson, CS chair at UT Austin

Oct 7

Oct 7Wed
  1. Andrew CurranAI score52

    AI Labs Reportedly Test Internal Models Against Cryptographic Protocols

    Scott Aaronson reports, based on his sources, that some AI companies have begun discreetly investigating whether their latest internal models can break important cryptographic protocols and primitives. He notes that cryptography is conspicuously absent from OpenAI's list of 376 papers, and the quoted post adds that the US government has censored academic quantum cryptanalysis results.

  2. Miles BrundageAI score26

    IMO it is possible to have an impact from the inside but also people generally overestimate inside impacts, and there's a bit of a "streetlight effect" - you see/focus on all the internal opportunities for impact but not the zillion ones outside https://x.com/KatjaGrace/status/2107672181622878665?s=20

    IMO it is possible to have an impact from the inside but also people generally overestimate inside impacts, and there's a bit of a "streetlight effect" - you see/focus on all the internal opportunities for impact but not the zillion ones outside https://x.com/KatjaGrace/status/2107672181622878665?s=20

  3. Andrew CurranAI score32

    Anyone who believes hallucination is an intractable problem for current models needs to update on both reality and directionality. So much has changed in the last six months. https://deploymentsafety.openai.com/gpt-6-october

    Anyone who believes hallucination is an intractable problem for current models needs to update on both reality and directionality. So much has changed in the last six months. https://deploymentsafety.openai.com/gpt-6-october

  4. Amazon ScienceAI score10

    Amazon Scholar and @UTAustin professor @mattlease explores how to harness AI for scientific discovery while assessing potential risks, drawing on work from @UTGoodSystems and @CosmicAI_Inst. Catch his Expo Talk at @COLM_conf Thursday at 1pm PT. #COLM2026

    Amazon Scholar and @UTAustin professor @mattlease explores how to harness AI for scientific discovery while assessing potential risks, drawing on work from @UTGoodSystems and @CosmicAI_Inst. Catch his Expo Talk at @COLM_conf Thursday at 1pm PT. #COLM2026

  5. Andrew CurranAI score42

    Andrew Curran says OpenAI's 722 math results omit cryptography breakthroughs

    Andrew Curran notes OpenAI's 722 published mathematical results show striking under-representation of cryptographic breakthroughs, and he says he has personally witnessed US government censorship of academic quantum cryptanalysis results. He calls backroom government interventionism his base case and says rumors suggest yesterday's OpenAI math release was only the first of three batches.

  6. Gergely OroszAI score17

    one more reason AI agents will increasingly run on the cloud, not locally but even so, that doesn't deal with how AI agents can do stupid stuff with sensitive data as well these things are nondeterministic. To gurantee something does not happen, you need a deterministic sys

    one more reason AI agents will increasingly run on the cloud, not locally but even so, that doesn't deal with how AI agents can do stupid stuff with sensitive data as well these things are nondeterministic. To gurantee something does not happen, you need a deterministic sys

  7. Semafor · TechnologyAI score34

    Alex Stamos Criticizes Silicon Valley's "Nihilism" and Separates Real AI Risks From Imagined Ones

    Cognition CISO and former Facebook security chief Alex Stamos criticized "nihilism" in Silicon Valley and argued that some AI risks are real while others are shaped by "almost religious beliefs" held by people at AI companies. He said AI systems "are not conscious, they do not have souls," and that he plans to "work the problem" to help shorten the expected "dark age" of cybersecurity.

Oct 6

Oct 6Tue
  1. PlatformerAI score49

    Anthropic and OpenAI Leaders Weigh Hard Caps on AI Intelligence

    Speakers at The Curve, a Berkeley AI conference, discussed limiting how intelligent large language models can become, amid concerns over recursive self-improvement. Proposed approaches include Anthropic's responsible scaling policy, limits on compute and model copies, and restrictions on using frontier models for AI research. The column notes such enforcement tools do not yet exist and that the Trump administration opposes such restrictions.

  2. METRAI score22

    Treating agent observability as security-critical infrastructure means handling all AI outputs (transcripts, reasoning, actions, etc.) as untrusted inputs: ideally we can make it very difficult for agents to influence the systems we use to supervise them.

    Treating agent observability as security-critical infrastructure means handling all AI outputs (transcripts, reasoning, actions, etc.) as untrusted inputs: ideally we can make it very difficult for agents to influence the systems we use to supervise them.