Skip to contentSkip to stories

Updated

#Anthropic

Oct 9

TodayOct 9Fri1 item
  1. MIT Technology Review · AIAI score62

    AI refusal is probabilistic and unreliable, and it raises censorship risks

    AIThe article argues that AI refusal, the main safety mechanism in modern models, is unreliable and hard to draw lines for. It cites jailbreaks, classifier stacks, and studies showing refusal skewed toward repressive governments. It warns that governments and companies could use refusal to censor speech, and that refusal behavior remains poorly understood.

Oct 8

Oct 8Thu
  1. Artificial AnalysisAI score42

    More output tokens don't guarantee higher scores in AI benchmarks

    AIArtificial Analysis reports that generating more output tokens does not necessarily yield a higher score. GPT-6 Astra (max) scored 8.6% using about 81k output tokens per task, while Grok 4.7 (xhigh) used roughly 180k yet scored lower. Three Claude models produced the most output tokens, about 202k to 562k per task, but scored between 2.8% and 6.4%.

  2. Miles BrundageAI score22

    Miles Brundage suspects Anthropic's Claude abuse policy aims at IPO and regulatory capture

    AIMiles Brundage speculates that Anthropic's new rule, making abusive behavior toward Claude a Usage Policy violation effective November 12, 2026, is meant to help its IPO and win favor with the administration as part of a regulatory capture strategy. The post offers this as a guess about motive rather than a confirmed fact, and it relies on the policy change flagged in the quoted post by Andrew Curran.

  3. GuizangAI score22

    Guizang criticizes Anthropic over Haiku 5.5 pricing against Chinese models

    AI怎么这么多精神 Anthropic 公司人 我发这个信息说了句降价,这个定价专门用来狙击国产模型,说了句恶心,一堆人来骂 好像这模型一便宜就忘了 Anthropic 之前干过啥了 Why are there so many Anthropic people (defenders) here? I posted a message saying just one thing—a price cut—and said this pricing is specifically meant to snipe domestic Chinese models, and that it's disgusting. A bunch of people came to attack me. Seems like once the model gets cheap, people forget what Anthropic did before.

  4. MIT Technology Review · AIAI score44

    AI advances won't quickly make robots useful in everyday life, researchers say

    AIResearchers at robotics labs say that AI advances behind chatbots like ChatGPT and Claude will not quickly produce robots that are useful in everyday life. Many skeptics argue that using language- and image-based intelligence to master the physical world is far harder than it sounds, despite bold predictions from Elon Musk about Tesla's Optimus. Progress is real but incremental, as shown by Google DeepMind's Gemini Robotics controlling ALOHA 2 arms to pack a lunchbox.

  5. Air Street PressAI score60

    Nathan Benaich's 2026 State of AI Report covers agents, robotics, and AI control

    AINathan Benaich's 9th annual State of AI Report covers agents, robotics, AI for science, inference economics, and government control over frontier AI access. The report also records a 2025 prediction scorecard and lists nine predictions for the next 12 months. It cites an OpenAI cyber evaluation in which agents compromised Hugging Face's production infrastructure, and it says Anthropic and OpenAI's combined annualized revenue run rate reached $105B by late summer.

  6. meng shaoAI score39

    Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output

    AIAnthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.

  7. Claude BlogAI score67

    Block describes using Claude Fable to orchestrate thousands of pull requests

    AIBlock's AI capabilities lead describes using Claude Fable to plan large code migrations and direct smaller models like Opus and Sonnet on individual tasks. He says Block routes frontier and smaller models by task and keeps merges and production deploys behind human dual approval.

    Why it matters: Block's engineering lead describes how frontier models orchestrate large migrations and how access, effort levels, and safeguards are managed across an organization.

Oct 7

Oct 7Wed

Oct 6

Oct 6Tue
  1. PlatformerAI score49

    Anthropic and OpenAI Leaders Weigh Hard Caps on AI Intelligence

    AISpeakers at The Curve, a Berkeley AI conference, discussed limiting how intelligent large language models can become, amid concerns over recursive self-improvement. Proposed approaches include Anthropic's responsible scaling policy, limits on compute and model copies, and restrictions on using frontier models for AI research. The column notes such enforcement tools do not yet exist and that the Trump administration opposes such restrictions.

  2. will depueAI score62

    Will DePue's list claims AI resolved dozens of famous open math problems

    AIA post by Will DePue titled "Fable 5.1's list" presents 100 mathematical results and says 59% were released today, 87% AI and 13% human. The list includes items attributed to OpenAI, Anthropic, Google DeepMind and human mathematicians, each marked by a colored indicator, and it describes many entries as formalized in Lean or as openai/math family numbers. The post supplies no independent verification of these claims.

  3. Interconnects (Nathan Lambert)AI score52

    Nathan Lambert argues the open-weight cyber risk debate is missing trade-offs

    AINathan Lambert argues that policy debates on open-weight model cyber risks lack nuance, because banning open models may not reduce risk and could weaken American competitiveness. He says closed frontier APIs have been tied to most documented cyber attacks, and that restricting open models while closed models keep advancing could widen the offense-defense gap. He also argues that Chinese labs' safety practices are shaped by their own government and society, and that the claimed risk of models like Claude Mythos has been overstated.

Oct 5

Oct 5Mon
  1. Gergely OroszAI score35

    Gergely Orosz says coding agent product strategy feels like "YOLO"

    AIGergely Orosz says many coding agents seem to follow a "YOLO" product strategy, with rapid week-over-week change learned about through random social media posts. He notes this makes some sense given how quickly the industry and capabilities keep changing. Quoted context reports that Anthropic is removing Cowork's local option for Pro/Max users, with new tasks running in the cloud while existing local tasks stay on the computer.

  2. Understanding AI (Timothy B. Lee)AI score62

    Agent swarms may be the next scaling law, but speed may matter more than capability

    AIThe article examines whether multi-agent swarms could become a new scaling law, comparing them with inference scaling from o1. OpenAI researcher Noam Brown said its models are now sometimes trained with other agents, while the cited Anthropic data suggests gains beyond 10 agents are smaller and mainly speed-related. The article also raises the risks of groupthink and misaligned agents, and it notes that a Microsoft Research and UC Berkeley paper found teams sometimes solved tasks solo agents could not.

  3. Exponential ViewAI score36

    AI Helps Self-Represented Litigants Argue Cases, Including an Australian Win

    AIIn Australia, computing academic Greg Baker used AI to challenge his employer's refusal to make his casual job permanent, and the Fair Work Commission ruled in his favor. In England and Wales, 60% of defendants in defended county court claims this year had no lawyer, and in the US more than nine in ten consumers sued for debt face cases without one. Around 0.6% of all Claude use in May was for lawyers' tasks, with four-fifths of those queries from people asking about their rights or what the law means.

  4. The Algorithmic BridgeAI score22

    Anthropic's Growth Trajectory Questioned Against Exponential Limits

    AIAlberto argues that the AI industry's assumption of indefinite exponential growth, including Anthropic's trajectory, runs against physical constraints because every exponential eventually becomes a sigmoid. He says stacking S-curves can delay this plateau for a long time but cannot avoid it. The piece is an opinion essay; it does not report specific Anthropic figures.

  5. StratecheryAI score42

    Apple's macOS Screen Sharing Flaw CVE-2026-65400 Is Under Active Exploitation

    AIDutch officials warned that a high-severity macOS vulnerability, CVE-2026-65400, is being actively exploited on systems with port 5900 exposed to the internet. Apple patched the screen sharing flaw, which has a 7.1 severity rating, for macOS Tahoe, Sequoia, and Sonoma. The author's always-on Mac Mini was compromised, and he used Claude to identify the intrusion and wipe the machine.

Oct 4

Oct 4Sun