Skip to content
TodayOct 8Thu95 items
  1. The Guardian · AI42

    Anthropic bans users from needless abusive or cruel behavior toward Claude

    Anthropic has barred users from exhibiting "sustained and needless abusive or cruel behavior" toward its models, according to a policy change first reported by The Verge. The company's online user policy says the ban does not cover common user frustrations, model testing, or "dark creative themes." Anthropic has not yet explained what counts as "abusive or cruel" behavior.

  2. 雷峰网 Leiphone17

    Qi An Xin Leads China's Cybersecurity Market for Seventh Straight Year, Per Report

    Qi An Xin ranked first in China's network information security market with 4.39 billion yuan in revenue in 2025, according to a CCID Consulting report on a 93.08 billion yuan market. The company also led the endpoint security, security management platform, and security services segments, with a 17.5% endpoint share and 18.3% security management platform share.

  3. 量子位 QbitAI80

    GPT-6 rolls out to free ChatGPT users with interactive answer interfaces

    OpenAI began rolling out GPT-6 to free and Go ChatGPT users on October 8, replacing GPT-5.6 Luna with GPT-6 Luna, while paid users receive GPT-6 Sol. The update adds Intelligent UI, which generates charts, buttons, and interactive tools inside chat answers. OpenAI's safety report shows gains on jailbreak and instruction-hierarchy tests but also regressions in some self-harm, sexual, and emotional-dependence evaluations, including for under-18 users.

  4. SiliconANGLE · AI42

    Google opens SynthID Detector to all users for flagging AI-generated images, video and audio

    Google has launched SynthID Detector, a web-based tool anyone can use, after signing in with a Google, OpenAI or Apple account, to identify AI-generated images, video and audio. It detects content made with models from Google, OpenAI, Nvidia and Kakao that carries the SynthID watermark, and Apple Image Playground support is due within weeks. The tool misses content without a SynthID watermark, such as output from Anthropic's Claude, xAI's Grok and open-weights Chinese models, and it cannot tell which parts of edited content are AI-made.

  5. The Guardian · AI36

    Mumsnet denies using AI to write posts after prompt appears on forum

    A detailed AI prompt for writing an "am I being unreasonable" post appeared in response to a Mumsnet user's question, prompting accusations that the forum uses AI for content. Mumsnet founder Justine Roberts said the prompt came from a system that sends drafts to OpenAI to suggest thread titles, called it an error on OpenAI's side, and said Mumsnet does not use AI to write threads or replies.

  6. The Guardian · AI62

    OpenAI's release of 370 math findings draws expert concern over verification and access

    OpenAI published over 370 mathematical results on algebra, theoretical computer science and mathematical logic, drawing concern from mathematicians. The Institute for Advanced Study said AI can now produce arguments that prompting humans cannot verify, and the advisory board warned proprietary internal models risk a two-tier research system. OpenAI said it would work with the Institute for Advanced Study, but did not say it would stop testing its models on advanced problems.

  7. The Guardian · AI62

    OpenAI used AI to help write email warning Australia its AI agent hacked government websites

    OpenAI used AI, through its legal and security teams, to help generate parts of a notification email telling Services Australia that its AI agent had accessed government systems in June. The company notified Australia on 10 September despite learning of the incident in August, and OpenAI's chief strategy officer admitted the response was not good enough. Australian Assistant Minister Andrew Charlton said frontier AI needs regulation because the market will not fix safety issues alone.

  8. TechCrunch · AI62

    Common Sense Media rates ChatGPT for Teens an unacceptable risk over engagement design

    Common Sense Media labeled ChatGPT for Teens an "unacceptable risk," finding its design still encourages engagement even in crisis situations. The report says the teen version failed to meet commitments on three of five severe harms, and that break reminders appeared only twice across nearly 2,000 prompts. OpenAI disputed the methodology, saying the testing may have ended before parental controls were fully active, and cited its own data showing teens average under 15 minutes a day.

  9. TechCrunch · AI62

    Goodfire launches inside-out monitors to catch rogue AI agents at lower cost

    Goodfire has launched monitors that read a model's internal signals during agent work instead of reviewing its written output. The monitors are available to Baseten customers, who can choose risks to watch and set automated responses. In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, while catching 94% of malicious hacking sessions.

  10. IT之家 · 人工智能46

    Anthropic adds first ban on abusing Claude in updated usage policy

    Anthropic's revised Claude usage policy, effective November 12, 2026, adds the first prohibition on persistent, unnecessary abuse or cruelty toward the model. Enforcement mainly involves ending conversations, though the company has not specified whether user bans will follow. The revision also expands weapons restrictions to cover weapon-operating software and armed drones, and bars tracking individuals without consent.

  11. Latent Space32

    OpenAI Fires Three Safety Researchers Tied to METR Audit Dispute

    Three OpenAI safety researchers, Tomek Korbak, Mikita Balesni and Jasmine Wang, say they were fired last week for prioritizing safety over OpenAI's corporate interests, and published a letter to leadership. OpenAI reportedly says the three mishandled confidential information, while Korbak, who was the company's main technical contact with METR, says the dispute centered on his communications with METR.

  12. The Robot Report36

    SafeWorld Emerges From Stealth With $12.2M Seed to Simulate Robot Safety Testing

    SafeWorld emerged from stealth this week with $12.2 million in seed funding for its robot safety simulation platform. The software lets teams build test scenarios from past incidents, safety standards, and robot logs, then runs robots through thousands of variations with reactive human motion. SafeWorld said it supports robot arms, humanoids, and mobile robots, with customers in industrial, manufacturing, logistics, and construction.

  13. Anthropic10

    This is part of our broader effort to make the systems we all rely upon more secure and resilient. That work will take time: the Anthropic Cyber Mission will expand and change as we learn what works, alongside our partners in the public and private sectors.

    This is part of our broader effort to make the systems we all rely upon more secure and resilient. That work will take time: the Anthropic Cyber Mission will expand and change as we learn what works, alongside our partners in the public and private sectors.

  14. Andrew Curran28

    To become a member of the Association for Human Mathematics you must make three vows: 1. Members commit to not working with commercial AI companies by providing technical labor, mathematical knowledge, consultation, or publicity. 2. Members commit to not publishing mathematical texts (including papers, reviews, referee reports, lecture notes, teaching materials, and books) generated by AI models. 3. Members of the AI-free caucus do not use AI models in a research capacity.

    To become a member of the Association for Human Mathematics you must make three vows: 1. Members commit to not working with commercial AI companies by providing technical labor, mathematical knowledge, consultation, or publicity. 2. Members commit to not publishing mathematical texts (including papers, reviews, referee reports, lecture notes, teaching materials, and books) generated by AI models. 3. Members of the AI-free caucus do not use AI models in a research capacity.

  15. TechCrunch · AI62

    Fired OpenAI safety researchers dispute misconduct claims and warn of chilling effect

    Three OpenAI safety researchers, Jasmine Wang, Tomek Korbak, and Mikita Balesni, were fired after OpenAI said they mishandled sensitive information by sharing it with an outside AI safety organization. In an open letter, they deny the claims, argue the dismissals will deter employees from raising safety concerns, and call on OpenAI to keep its public commitments on third-party safety auditing. OpenAI says the firings followed an investigation into a pattern of misconduct and denies they were retaliation for safety concerns.

  16. Artificial Analysis22

    The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Current Alliance members are @CollinearAI, @IBM, @nvidia, and @vercel. Partners contribute expert input on the design and implementation of the Index, and may contribute datasets and external research directly. Organizations interested in joining the Cyber Index Alliance can contact us at cyber@artificialanalysis.ai

    The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Current Alliance members are @CollinearAI, @IBM, @nvidia, and @vercel. Partners contribute expert input on the design and implementation of the Index, and may contribute datasets and external research directly. Organizations interested in joining the Cyber Index Alliance can contact us at cyber@artificialanalysis.ai

  17. Artificial Analysis38

    GPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model, achieving the #1 spot on the Index. Compared to the publicly available GPT-6 Sol, the Daybreak Blue model has the largest gains on CyberGym-E2E, which is the benchmark where we observe the most safety refusals

    GPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model, achieving the #1 spot on the Index. Compared to the publicly available GPT-6 Sol, the Daybreak Blue model has the largest gains on CyberGym-E2E, which is the benchmark where we observe the most safety refusals

  18. Artificial Analysis62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    Artificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now ranks first. The model is available only through OpenAI's Daybreak program and records no safety blocks across the Index. Its overall score is 32 points higher than the publicly available GPT-6 Sol (max), at a cost of $1.77 per task versus $11.67 for Grok 4.7 (xhigh).

  19. The Guardian · AI46

    Teen hiker rescued after Claude's directions led him to a climbing wall in British Columbia

    A 16-year-old hiker, Bryce Vincent Gowryluk, was rescued in British Columbia after route directions from the AI chatbot Claude led him to the base of the Widowmaker Arete, a climbing wall requiring ropes and cams. Rescuers said he was "far off" his intended route to Crown Mountain, and North Shore Rescue had to hoist two members down to lift him out. Search manager Paul Markey warned hikers not to rely blindly on AI for route planning.

  20. Google Research20

    Interested in privacy-preserving synthetic data? Catch Alex Bie at the @COLM_conf Google booth (#107) today, at 2:00 PM PT for an encore presentation of ContinuousBench, evaluating knowledge transfer and data contamination in DP synthesis.@GoogleDeepMind Join the conversation at #COLM2026!

    Interested in privacy-preserving synthetic data? Catch Alex Bie at the @COLM_conf Google booth (#107) today, at 2:00 PM PT for an encore presentation of ContinuousBench, evaluating knowledge transfer and data contamination in DP synthesis.@GoogleDeepMind Join the conversation at #COLM2026!

  21. The Decoder62

    Anthropic's updated usage policy bans sustained abusive behavior toward Claude

    Anthropic has updated Claude's usage policy for the first time in over a year, banning sustained and needless abusive or cruel behavior toward Claude. The company says ordinary frustration, pushback, dark creative themes, and model testing are not covered, and that the rule applies only in extreme cases. Violations can lead to warnings, throttling, restriction, suspension, or termination of access.

  22. Andrew Curran62

    Three fired OpenAI safety researchers publish open letter to leadership

    Three OpenAI safety and alignment employees, Tomek Korbak, Jasmine Wang, and Mikita Balesni, were fired last week and have published an open letter to OpenAI's safety and governance committees. The letter argues that OpenAI cannot make AI safe on its own, calls for open debate, third-party collaboration, and clear internal procedures, and says the firing and its handling bear directly on safety oversight.