Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. Microsoft CopilotOfficialAI score30

    Microsoft unveils new Copilot Home combining Chat and Cowork

    AIMicrosoft says the new Copilot Home brings Chat and Cowork into one experience, so users can move from thinking to doing without losing context. Microsoft Copilot EVP Jacob Andreou explains why Home is the new starting point for work.

    Video from @MSFTCopilot's post
  2. LangChain BlogOfficialAI score40

    LangChain adds emoji reactions to Managed Deep Agents Slack channels

    AILangChain's Managed Deep Agents v0.9 adds a reactions attribute for Slack channels that accepts either an emoji string or a callable returning one. The article shows a function that returns a bug emoji when a message contains "broken" and eyes otherwise. It also shows a TypeSafe Classifier that picks from a seven-emoji vocabulary and falls back to eyes below 25% confidence.

  3. 🚨 AI News | TestingCatalogXAI score62

    Anthropic moves dynamic workflows in Claude Managed Agents into public beta

    AIAnthropic has expanded dynamic workflows in Claude Managed Agents into a public beta, according to Testing Catalog. Users can configure their agents for multiagent orchestration, with Claude planning and operating a fleet of agents to achieve a goal. The post also links a video from Anthropic's ClaudeDevs account, which the author describes as a new SWE norm.

    Video from @testingcatalog's post
  4. merveXAI score28

    Hugging Face lets agents train Qwen3.8-27B on Nebius GPUs

    AIHugging Face launches an arena where users bring their own agent, which gets Nebius GPUs to build RL environments that improve Qwen3.8-27B across eight domains. The arena runs on PostTrainArena from BenchFlow, with compute from Nebius. Setup requires only a few steps through the linked OpenEnv Arena space.

    Video from @mervenoyann's post
  5. ClaudeDevsOfficialAI score60

    Claude Managed Agents adds dynamic workflows in public beta

    AIAnthropic's ClaudeDevs account announces that dynamic workflows for Claude Managed Agents are now available in public beta. The feature is a new type of multiagent orchestration in which a lead agent writes a plan that runs across many agents in phases, then combines their results at the end.

    Why it matters: The post describes how a lead agent plans work across many agents in phases and merges their results, a structure useful for understanding complex agent orchestration.

    Video from @ClaudeDevs's post
  6. Perplexity DevelopersOfficialAI score34

    Perplexity releases cookbook for a browser agent using the Decisions API

    AIPerplexity Developers says its new cookbook builds a browser agent that sends a screenshot and questions to pplx-decider-v1.1-27b through the Decisions API, which accepts text and image inputs. The developer's code converts the returned probabilities into clicks, scrolls, and stops.

  7. Replit ⠕OfficialAI score22

    TikTok Ads MCP lets users run TikTok Ads from Replit

    AIReplit's X account shares a showcase of TikTok Ads MCP, which runs TikTok Ads from Replit. The post is a broadcast link with no further details on features, pricing, or availability.

  8. elvisXAI score60

    Meta researchers propose agent plasticity to measure self-improvement efficiency

    AIResearchers from UC Berkeley, Meta Superintelligence Labs, and other institutions introduce agent plasticity, the gain on held-out tasks per dollar of learning cost, with model weights frozen. The paper reports that in chess, Go, and Hex, Claude Fable 5 reaches the highest final score while GPT-5.6 Sol gains the most per dollar, and in NetHack only Claude Opus 5.5 improves significantly.

    Image from @omarsar0's post
  9. Ai2OfficialAI score22

    Ai2 replaces its GPU scheduler after idle jobs hoarded capacity

    AIAi2 says its old scheduler made every scheduled workload eventually run at HIGH priority. Researchers kept idle jobs running to reserve GPUs for experiments, because the incentives rewarded holding capacity even with no active work.

  10. Ai2OfficialAI score25

    Ai2's new scheduler delivers 98% of owed GPU hours in 30-day test

    AIAi2 reports that over a 30-day test of its new scheduler, teams received 98% of the GPU hours they were owed, based on actual demand. Cluster occupancy stayed at 98%, and spare capacity went to interruptible work without drawing down team budgets.

  11. The Algorithmic BridgeBlogAI score40

    Meta's AI comeback follows heavy Anthropic Claude spending and a new Muse Spark model

    AIMeta spent heavily on Anthropic's Claude models, with internal use reaching up to 60,000 employees and a projected $10 billion yearly spend, according to The Algorithmic Bridge. The author says Meta then released Muse Spark, which scored 52 on the Artificial Analysis intelligence benchmark, on par with Claude Opus 4.6.

  12. Ai2 (Allen Institute for AI)OfficialAI score46

    Ai2 describes GPU time budgets that replaced its priority-based cluster scheduler

    AIAi2's AI Infrastructure team replaced its priority-based scheduler for GPU clusters with GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change moved debates over how much GPU time each research project deserves from case-by-case operational decisions into a transparent budgeting process. The clusters range from 88 to 1024 GPUs across NVIDIA H100, B200, and B300 hardware, and serve about 150 internal researchers.

  13. AWS Machine Learning BlogOfficialAI score67

    How Postman runs Agent Mode for 40 million developers on Amazon Bedrock

    AIPostman describes the architecture behind Agent Mode, its AI agent for API testing, documentation, discovery, and implementation. The post covers limiting tools per task, using schema-based queries, building purpose-shaped context handlers, and running on Amazon Bedrock with cross-Region inference and prompt caching. Postman reports that tool-selection errors rose once the visible toolset exceeded about 40 tools.

    Why it matters: The post shows concrete patterns for tool scoping, context handling, and Bedrock routing and caching, which apply to any team moving an agent past a prototype.

  14. AWS Machine Learning BlogOfficialAI score36

    AWS recaps September 2026 Bedrock, AgentCore, and Strands updates for AI builders

    AIAmazon Bedrock Managed Agents, powered by OpenAI, entered public preview, and OpenAI's GPT-6 Astra, GPT-6.1 Sol, and GPT-6.1 Luna became generally available on Amazon Bedrock. AWS also released Strands Decider 2B, a 2B-parameter open source decision model that answers in about 115ms locally, and said the Strands harness uses 28 percent fewer tokens than popular harnesses while matching their accuracy.

  15. Hugging Face BlogOfficialAI score38

    Ai2 replaces priority scheduler with GPU time budgets for cluster allocation

    AIAi2's AI Infrastructure team replaced its priority-based GPU cluster scheduler with a system using GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change turns decisions about how much GPU time each research project receives into a transparent administrative budgeting process. Its clusters, which range from 88 to 1024 GPUs including H100, B200, and B300 units, serve about 150 researchers facing demand two to three times available capacity.

  16. ElevenLabsOfficialAI score32

    ElevenLabs partners with Banner Health on AI voice agents for patient calls

    AIElevenLabs says it is partnering with Banner Health to answer patient calls with AI voice agents, starting with primary care scheduling. The ElevenAgents system books, reschedules, or cancels appointments directly in Banner's electronic medical record at any hour, and transfers calls to a Banner team member with context when a patient asks for a person.

    Image from @ElevenLabs's post
  17. Andrew CurranXAI score62

    OpenAI responds to three fired employees' letter on safety and trust

    AIOpenAI's research leaders say they parted ways with Jasmine, Mikita, and Tomek after an investigation found they violated policies on handling sensitive information. The company says the decision was not about raising safety concerns and that it is finalizing contracts with third-party safety assessors, with details to follow in the coming weeks.

  18. OpenAI DevelopersOfficialAI score46

    Codex on Windows gets new MXC-based sandbox mode

    AIOpenAI says Codex on Windows now has a new sandbox mode built on Microsoft's Execution Containers (MXC), offering faster setup, stronger network enforcement, and granular file access controls. The mode requires a compatible Windows 11 device. Background from Microsoft's announcement says MXC is now generally available on Windows 11, keeping agents within boundaries the operating system enforces.

  19. Baseten BlogOfficialAI score61

    How to choose which layers to run at NVFP4 quantization precision

    AIBaseten explains how to decide which layers of a model can run in 4-bit NVFP4 without losing needed information. The post compares architecture-based heuristics, isolated-layer sensitivity scoring, and SaturationQuant, which accounts for other quantized layers. It also covers calibration with representative data and block-level scales of 16 values.

    Why it matters: The post explains how to choose which layers run at NVFP4 precision using heuristics, sensitivity scoring, and saturation-aware scoring, with clear calibration steps.

  20. GuizangXAI score22

    Guizang releases a one-click Grok bot for daily AI news videos

    AIGuizang says he turned his workflow into a Grok bot that users can install with one click. The bot runs on Grok's cloud virtual machine to collect content, write code, and render a daily morning AI news video without using a local computer.

  21. Arena.aiOfficialAI score24

    Arena weekly update: Nano Banana 2.1, Mistral Large 4, Claude Haiku 5.5 rankings

    AIArena's weekly update says Nano Banana 2.1 ranked in the top six across three Image Arena modes, with #4 in Multi-Image Edit at 1431 points. Mistral Large 4 placed #43 overall in Agent Arena, 11 spots above Mistral Medium 3.5, and Claude Haiku 5.5 (High) landed #30 in Code Arena WebDev at 1587 points, priced at $0.10/$0.50 per 1M input/output tokens. The post also introduces Arena's Alignment Index and announces a $200M Series B at a $3.1B valuation.

  22. Boris PowerXAI score28

    Boris Power calls OpenAI integer multiplication progress "Wow!"

    AIBoris Power, who owns the OpenAI account, posted only the word "Wow!" with no details. Background from a separate post says the integer multiplication problem #109 witness value κ rose to 2⁻¹⁰·⁵⁴⁷ (about 6.6857 × 10⁻⁴), past the 2⁻¹¹ threshold. The author notes gains are now fractional and a major breakthrough is still needed.

  23. DatabricksOfficialAI score25

    Databricks pairs Temporal and Lakebase for durable cloud agents

    AIDatabricks has published a reference implementation pairing Temporal with Lakebase Postgres so cloud agents can survive worker, container, or deployment replacement. The design keeps recorded work and evidence and review state queryable, and lets human decisions arrive days later. Unity Catalog remains the governed policy source through synced tables.

    Image from @databricks's post
  24. Kilo (acq. by Anaconda)OfficialAI score60

    StepFun's Step 5 Preview is free in Kilo for one week

    AIKilo says StepFun has announced Step 5 Preview, which is free to use in Kilo for one week. The post lists 600B total parameters with 27B active per token, a 1M-token context window with vision, and highlights strong coding and finance performance at lower cost.

    Image from @kilocode's post
  25. South China Morning Post · TechNewsAI score52

    Anthropic alleges Chinese AI firms covertly used its Claude model

    AIAnthropic claims Chinese AI developers used fraudulent accounts and proxy networks to extract reasoning data from its flagship model, Claude. The company says some firms used Claude as a covert back end for their own apps. A joint advisory from the NSA, FBI and CISA last month, and US Treasury Secretary Scott Bessent's July warning about large-scale distillation, add to the allegations.