Skip to contentSkip to stories

Updated

#Anthropic

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 7

Oct 7Wed
  1. Wired · AIAI score46

    Pentagon's Tradewinds Program Uses Five-Minute Videos to Speed AI Purchases

    AIThe Department of Defense's Chief Digital and Artificial Intelligence Office runs the Tradewinds program, which grants "post-competitive" status to AI vendors based on videos of five minutes or less. That status can let government buyers use other transaction agreements and sometimes make awards in less than a week. OpenAI, Anthropic, and Google are listed as participants, though none commented.

  2. 🚨 AI News | TestingCatalogAI score47

    Daily AI brief covers Mistral Large 4, Google, OpenAI, and Anthropic updates

    AIMistral released Mistral Large 4 "Le Chonk", a 1T-parameter (49B active) multimodal model, with open weights planned in about three weeks. Google rolled out Nano Banana 2.1 across Gemini, AI Studio, and the Gemini API, and released EmbeddingGemma 2, a 740M-parameter open multimodal embedding model under Apache 2.0. OpenAI launched the Decisions API in beta with gpt-6-luna, returning typed answers 10x faster than the Responses API.

  3. Latent SpaceAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    AIOpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

  4. Claude BlogAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    AIAnthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

  5. Artificial Analysis ArticlesAI score60

    Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index

    AIAnthropic released Claude Haiku 5.5, which scores 43 on the Artificial Analysis Intelligence Index, up 26 points from the last Haiku release. Pricing is $0.10/$0.50 per 1M input/output tokens up to 100k tokens, rising to $0.50/$2.50 above that, but at max effort it uses about 162k output tokens per Intelligence Index task, roughly 3x GPT-6 Luna.

    Why it matters: The benchmark shows Haiku 5.5 scores well but uses far more output tokens than GPT-6 Luna, so cost per task matters beyond list price.

  6. Claude BlogAI score66

    Claude skill commands build evals and hillclimb them against overfitting

    AIAnthropic added build-eval and hillclimb commands to its claude-api skill for designing evaluations and iteratively improving applications against them. The article covers eval design principles, including production-representative tasks, headroom and low variance, and guards against overfitting through train/test splits. Two examples report results: a customer support benchmark where cost fell to under half while accuracy rose, and a claude-api skill eval that rose from 66% to 88%.

    Why it matters: The article gives a concrete workflow for designing evals and hillclimbing without overfitting, with two worked cost and performance examples that show the tradeoffs.

Oct 6

Oct 6Tue
  1. PlatformerAI score49

    Anthropic and OpenAI Leaders Weigh Hard Caps on AI Intelligence

    AISpeakers at The Curve, a Berkeley AI conference, discussed limiting how intelligent large language models can become, amid concerns over recursive self-improvement. Proposed approaches include Anthropic's responsible scaling policy, limits on compute and model copies, and restrictions on using frontier models for AI research. The column notes such enforcement tools do not yet exist and that the Trump administration opposes such restrictions.

  2. Claude Apps Release NotesAI score60

    Claude Haiku 5.5 launches as a fast, low-cost small model, and Max and Team plans gain monthly API credits

    AIAnthropic launched Claude Haiku 5.5, which it describes as the cheapest, fastest, and most capable small model it has released, aimed at high-volume, cost-sensitive tasks. Max and Team plans now include monthly API credits for running their own apps and agents on the Claude Platform, rolling out over a few days. Users claim the credits by linking a Claude Console organization in Settings > Billing for Max or Organization settings > Billing for Team.

    Why it matters: The notes name a new small model and a credit change for Max and Team plans, with the claim path, which matters for teams budgeting API use.

  3. Epoch AIAI score60

    Epoch AI finds frontier models fall short of an end-to-end AI research task

    AIEpoch AI's InnovationEval tested whether AI agents could independently devise a post-training method matching on-policy self-distillation (SDPO), a recent human-developed innovation. GPT-5.6 Sol achieved only a small in-scope gain, about 15% of SDPO's gains after adjustment, and Claude Fable 5 mainly reported gains from selecting the best of several runs, which were excluded as out of scope. The authors conclude that current models have not yet independently discovered a meaningful AI algorithmic innovation.

    Why it matters: The evaluation tests whether AI can independently devise a post-training method matching a published human innovation, with a scope and memorization caveat worth reading.

  4. will depueAI score62

    Will DePue's list claims AI resolved dozens of famous open math problems

    AIA post by Will DePue titled "Fable 5.1's list" presents 100 mathematical results and says 59% were released today, 87% AI and 13% human. The list includes items attributed to OpenAI, Anthropic, Google DeepMind and human mathematicians, each marked by a colored indicator, and it describes many entries as formalized in Lean or as openai/math family numbers. The post supplies no independent verification of these claims.

    Image from @willdepue's post
  5. Boris ChernyAI score38

    Boris Cherny shares prompts for formally verifying Claude Agent SDK

    AIBoris Cherny says he used Opus 5.5 with Lean to formally verify the Claude Agent SDK, with a couple of short prompts producing 16 PRs fixing bugs and race conditions. He also reports that TLA+ works well, sometimes combined with Lean to find data flow, concurrency, and state management issues. The post links to his actual prompts as another example.

  6. AnthropicAI score49

    Anthropic expands Cyber Verification Program for verified security professionals

    AIAnthropic is expanding its Cyber Verification Program to give verified security professionals broader access to its most capable models. Through the program, they can use Claude Mythos 5.1, Opus 5.5, and Sonnet 5.5 with safeguards designed for defensive work. New tiers will also allow authorized offensive work such as penetration testing and red-teaming.

  7. Claude Code · GitHub ReleasesAI score40

    Claude Code v2.1.292 adds plugin marketplace flag and fixes security issues

    AIClaude Code v2.1.292 adds a --marketplace option to claude plugin install, which adds the marketplace if needed and then installs the plugin from it. The release also adds an effort parameter to the Agent tool and fixes several security issues, including permission prompts bypassed for network (UNC) file reads and a sandboxed read path that could return files outside approved access.

  8. Interconnects (Nathan Lambert)AI score52

    Nathan Lambert argues the open-weight cyber risk debate is missing trade-offs

    AINathan Lambert argues that policy debates on open-weight model cyber risks lack nuance, because banning open models may not reduce risk and could weaken American competitiveness. He says closed frontier APIs have been tied to most documented cyber attacks, and that restricting open models while closed models keep advancing could widen the offense-defense gap. He also argues that Chinese labs' safety practices are shaped by their own government and society, and that the claimed risk of models like Claude Mythos has been overstated.

  9. indigoAI score40

    SemiAnalysis: Anthropic's subscriptions yield about 5x OpenAI's API-equivalent value

    AIAccording to SemiAnalysis, Anthropic's subscriptions may be worth about 5 times more than OpenAI's when measured by API-equivalent value, even though subscriptions make up roughly 10% of Anthropic's revenue. The post says Anthropic's API revenue accounts for 75–85% of its total, while OpenAI's subscriptions reached 65% of revenue in 2026 Q1. It also claims Claude's average U.S. paid user pays about $45 per month, with a free-to-paid conversion of about 9%, versus roughly 6% for OpenAI.

    Image from @indigox's post
  10. Claude BlogAI score62

    Claude now works inside Google Docs, Sheets, and Slides in public beta

    AIClaude for Google Workspace is in public beta on all paid Claude plans, adding a sidebar to Google Docs, Sheets, and Slides. It can read the open file, edit text, build formulas, pivot tables, charts, and slides, and it asks for approval before changes unless the user chooses "Accept all edits." New Docs, Sheets, and Slides connectors in beta let Claude create and edit Google files from the chat, with access matching existing Google sharing permissions.

    Why it matters: The source specifies how Claude edits Docs, Sheets, and Slides in place and where users keep control, which clarifies the practical workflow change.

  11. Claude BlogAI score62

    Comcast and Booz Allen use Claude Mythos to find exploit chains in codebases

    AIComcast and Booz Allen used Claude Mythos Preview to find vulnerabilities that arise from interactions across code, configuration, and deployment rather than single-file bugs. Comcast identified a critical authentication flaw across 258 systems and about 170 million lines of code before any exploitation was observed. Booz Allen reported that one analyst reviewed eight production systems across 138 repositories in twelve days, a review its team estimated would have taken several months without the model.

    Why it matters: The case studies show how security teams validate and remediate model-found exploit chains, a workflow relevant to anyone managing large codebases.