Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 9

Oct 9Fri
  1. The Algorithmic BridgeBlogAI score40

    Meta's AI comeback follows heavy Anthropic Claude spending and a new Muse Spark model

    AIMeta spent heavily on Anthropic's Claude models, with internal use reaching up to 60,000 employees and a projected $10 billion yearly spend, according to The Algorithmic Bridge. The author says Meta then released Muse Spark, which scored 52 on the Artificial Analysis intelligence benchmark, on par with Claude Opus 4.6.

  2. Ai2 (Allen Institute for AI)OfficialAI score46

    Ai2 describes GPU time budgets that replaced its priority-based cluster scheduler

    AIAi2's AI Infrastructure team replaced its priority-based scheduler for GPU clusters with GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change moved debates over how much GPU time each research project deserves from case-by-case operational decisions into a transparent budgeting process. The clusters range from 88 to 1024 GPUs across NVIDIA H100, B200, and B300 hardware, and serve about 150 internal researchers.

  3. AWS Machine Learning BlogOfficialAI score36

    AWS recaps September 2026 Bedrock, AgentCore, and Strands updates for AI builders

    AIAmazon Bedrock Managed Agents, powered by OpenAI, entered public preview, and OpenAI's GPT-6 Astra, GPT-6.1 Sol, and GPT-6.1 Luna became generally available on Amazon Bedrock. AWS also released Strands Decider 2B, a 2B-parameter open source decision model that answers in about 115ms locally, and said the Strands harness uses 28 percent fewer tokens than popular harnesses while matching their accuracy.

  4. Hugging Face BlogOfficialAI score38

    Ai2 replaces priority scheduler with GPU time budgets for cluster allocation

    AIAi2's AI Infrastructure team replaced its priority-based GPU cluster scheduler with a system using GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change turns decisions about how much GPU time each research project receives into a transparent administrative budgeting process. Its clusters, which range from 88 to 1024 GPUs including H100, B200, and B300 units, serve about 150 researchers facing demand two to three times available capacity.

  5. ElevenLabsOfficialAI score32

    ElevenLabs partners with Banner Health on AI voice agents for patient calls

    AIElevenLabs says it is partnering with Banner Health to answer patient calls with AI voice agents, starting with primary care scheduling. The ElevenAgents system books, reschedules, or cancels appointments directly in Banner's electronic medical record at any hour, and transfers calls to a Banner team member with context when a patient asks for a person.

    Image from @ElevenLabs's post
  6. Andrew CurranXAI score62

    OpenAI responds to three fired employees' letter on safety and trust

    AIOpenAI's research leaders say they parted ways with Jasmine, Mikita, and Tomek after an investigation found they violated policies on handling sensitive information. The company says the decision was not about raising safety concerns and that it is finalizing contracts with third-party safety assessors, with details to follow in the coming weeks.

  7. PixVerseOfficialAI score14

    PixVerse hosts sessions demoing ChatGPT plugin video creation

    AIPixVerse says each session includes a platform walkthrough, a live OpenAI demo showing ChatGPT generating a creative brief and finished video via the PixVerse Plugin, a creator sharing their workflow, and live Q&A. The post presents these as recurring sessions rather than a new product launch.

  8. OpenAI DevelopersOfficialAI score46

    Codex on Windows gets new MXC-based sandbox mode

    AIOpenAI says Codex on Windows now has a new sandbox mode built on Microsoft's Execution Containers (MXC), offering faster setup, stronger network enforcement, and granular file access controls. The mode requires a compatible Windows 11 device. Background from Microsoft's announcement says MXC is now generally available on Windows 11, keeping agents within boundaries the operating system enforces.

  9. Baseten BlogOfficialAI score61

    How to choose which layers to run at NVFP4 quantization precision

    AIBaseten explains how to decide which layers of a model can run in 4-bit NVFP4 without losing needed information. The post compares architecture-based heuristics, isolated-layer sensitivity scoring, and SaturationQuant, which accounts for other quantized layers. It also covers calibration with representative data and block-level scales of 16 values.

    Why it matters: The post explains how to choose which layers run at NVFP4 precision using heuristics, sensitivity scoring, and saturation-aware scoring, with clear calibration steps.

  10. The Verge · AINewsAI score47

    Instinct AI agent holds its own against Muse and Dots in testing

    AIInstinct, a startup AI agent that works through text messages, handled online tasks such as swim lesson searches, an eye doctor appointment email, and an Ikea return about as well as Muse and Dots, according to The Verge's testing. The startup was valued at $10 billion in late September, and it has no app or subscription fee for now, with access by invitation or waitlist. Its founder, Noah Shinn, says its focus on a personal assistant sets it apart from OpenAI and Meta.

  11. GuizangXAI score22

    Guizang releases a one-click Grok bot for daily AI news videos

    AIGuizang says he turned his workflow into a Grok bot that users can install with one click. The bot runs on Grok's cloud virtual machine to collect content, write code, and render a daily morning AI news video without using a local computer.

  12. Arena.aiOfficialAI score24

    Arena weekly update: Nano Banana 2.1, Mistral Large 4, Claude Haiku 5.5 rankings

    AIArena's weekly update says Nano Banana 2.1 ranked in the top six across three Image Arena modes, with #4 in Multi-Image Edit at 1431 points. Mistral Large 4 placed #43 overall in Agent Arena, 11 spots above Mistral Medium 3.5, and Claude Haiku 5.5 (High) landed #30 in Code Arena WebDev at 1587 points, priced at $0.10/$0.50 per 1M input/output tokens. The post also introduces Arena's Alignment Index and announces a $200M Series B at a $3.1B valuation.

  13. Boris PowerXAI score28

    Boris Power calls OpenAI integer multiplication progress "Wow!"

    AIBoris Power, who owns the OpenAI account, posted only the word "Wow!" with no details. Background from a separate post says the integer multiplication problem #109 witness value κ rose to 2⁻¹⁰·⁵⁴⁷ (about 6.6857 × 10⁻⁴), past the 2⁻¹¹ threshold. The author notes gains are now fractional and a major breakthrough is still needed.

  14. DatabricksOfficialAI score25

    Databricks pairs Temporal and Lakebase for durable cloud agents

    AIDatabricks has published a reference implementation pairing Temporal with Lakebase Postgres so cloud agents can survive worker, container, or deployment replacement. The design keeps recorded work and evidence and review state queryable, and lets human decisions arrive days later. Unity Catalog remains the governed policy source through synced tables.

    Image from @databricks's post
  15. Kilo (acq. by Anaconda)OfficialAI score60

    StepFun's Step 5 Preview is free in Kilo for one week

    AIKilo says StepFun has announced Step 5 Preview, which is free to use in Kilo for one week. The post lists 600B total parameters with 27B active per token, a 1M-token context window with vision, and highlights strong coding and finance performance at lower cost.

    Image from @kilocode's post
  16. South China Morning Post · TechNewsAI score52

    Anthropic alleges Chinese AI firms covertly used its Claude model

    AIAnthropic claims Chinese AI developers used fraudulent accounts and proxy networks to extract reasoning data from its flagship model, Claude. The company says some firms used Claude as a covert back end for their own apps. A joint advisory from the NSA, FBI and CISA last month, and US Treasury Secretary Scott Bessent's July warning about large-scale distillation, add to the allegations.

  17. SiliconANGLE · AINewsAI score35

    SailPoint's Navigate event highlights a push to secure AI agent identities in real time

    AISailPoint's Navigate conference in Austin, Texas, featured executives arguing that enterprises must secure AI agent identities at machine speed through just-in-time access and enforcement outside the agent. Mark McClain, SailPoint's founder and chief executive, said real-time decision-making is needed because manual administration cannot keep up. The event also covered the Entro Security acquisition and a partnership with AWS on Amazon Bedrock AgentCore, which grew 15-fold in the first six months of the year.

  18. The Verge · AINewsAI score40

    Alexa Plus excels at running a smart home but falls short as a personal assistant

    AIAmazon's Alexa Plus, powered by generative AI, now responds in three to five seconds and handles multistep smart home commands, cooking questions, and calendar imports more reliably than the original Alexa, according to a year-long test by The Verge. The reviewer says its personal assistant features remain underbaked and frustrating, and that ads on Echo Show displays are excessive. Alexa Plus costs $19.99 a month in the U.S. unless users have an Amazon Prime membership, and the Echo Dot Max is recommended as the ad-free option.

  19. Qwen · new models on Hugging FaceOfficialAI score49

    Qwen releases Qwen-Image-2.1-Turbo, an 8-step accelerated image generation checkpoint

    AIQwen has published Qwen-Image-2.1-Turbo on Hugging Face, an accelerated checkpoint of Qwen-Image-2.1 for text-to-image generation and image editing with 8 denoising steps. The checkpoint uses the same 7B visual generation architecture, loads directly with QwenImage21Pipeline in Diffusers, and includes its recommended sampling schedule. It defaults to CFG=1 and uses prefix KV caching to reuse text and reference-image context across steps.

  20. Claude BlogOfficialAI score54

    Claude Managed Agents guide shows how to build scheduled agent automations

    AIThe Claude Blog published a guide to building scheduled agent automations with Claude Managed Agents (beta) that reads custom sources such as Slack and GitHub and posts a daily brief. The guide covers scoped vault credentials, per-source bookmarks so no window is lost or repeated, and confirming each Slack post before updating records. It also covers read-only access, a per-run spending cap, and a reference implementation with a Claude Code setup command.

  21. ModelScopeOfficialAI score60

    Qwen-Image-2.1-Turbo cuts image generation and editing to 8 denoising steps

    AIModelScope announces Qwen-Image-2.1-Turbo, an accelerated checkpoint that keeps the 7B visual architecture and runs image generation and editing in 8 denoising steps. The source says it uses CFG=1 and prefix KV caching to reuse text and reference-image context across steps, supports 2048 resolution with square, portrait, landscape, and widescreen presets, and loads through QwenImage21Pipeline in Diffusers. It is released under the Qwen Research License Agreement.

    Why it matters: The source names a concrete speedup path, 8 sampling steps and CFG=1 with prefix KV caching, which matters to anyone weighing image generation latency.

    Image from @ModelScope2022's post
  22. Don't Worry About the Vase (Zvi Mowshowitz)BlogAI score73

    OpenAI releases 719 AI-generated math manuscripts, splitting the mathematics community

    AIZvi Mowshowitz reports that OpenAI released 722 math manuscripts from an internal frontier model on GitHub, later reduced to 719 after three withdrawals, covering 90 of the top 500 open problems. He says the work came mostly from a single prompt, with an average of three hours of compute per solution. Mathematicians reacted with mixed feelings, and the post highlights concerns about unread papers, cryptography implications, and the role of Lean verification.

    Why it matters: The post traces how OpenAI's release of 719 math manuscripts divided mathematicians and reshaped verification, credit, and publication norms in the field.

  23. Simon WillisonBlogAI score27

    Simon Willison builds a new blog feature largely by voice with Codex

    AISimon Willison says he built a Newsletters index for his blog almost entirely by voice, using the ChatGPT desktop app's Codex voice mode while cooking dinner. The feature imports weekly Substack posts via RSS and undocumented API, monthly newsletters from a GitHub archive repository, and a private sponsors-only newsletter. He says he switched back to typing for review and fixes before deploying the pull request.

  24. Rohan PaulXAI score46

    Microsoft's TeleTune evolves agent skills from raw usage logs

    AIMicrosoft researchers present TeleTune, which lets agents learn software skills from raw usage logs by keeping only skill edits that better predict users' next actions. The method needs no live test environment, because next-action accuracy on held-out logs tracked live success. Unlike earlier methods such as Agent Workflow Memory, which need goal-labeled examples or a live environment, TeleTune guesses each session's goal and uses wrong guesses to suggest edits to a text skill library.

    Image from @rohanpaul_ai's post
  25. Sakana AIOfficialAI score37

    Sakana AI paper uses LLMs to catch errors in research papers

    AISakana AI researchers introduce a benchmark that plants contradictions in papers to test whether LLM reviewers can detect errors, and propose Multi-Layered Review, modeled on the Three-Pass Approach to reading. Their system detected more errors than the other review systems tested, including in papers withdrawn for real mistakes, while its paper-quality assessments stayed broadly consistent with human judgments. The work, accepted at TMLR, is framed as support for human reviewers rather than a replacement.

    Video from @SakanaAILabs's post
  26. Google WorkspaceOfficialAI score12

    RigStrips uses Google Workspace with Gemini to draft customer replies faster

    AIRigStrips co-founder Zhach Pham says he uses Google Workspace with Gemini to draft customer support replies instantly, saving hours that would otherwise go to outdoor gear testing. The post highlights the company's more than 150,000 units shipped, though it gives no specific figures for time saved.

    Video from @GoogleWorkspace's post