Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. AWS Machine Learning BlogOfficialAI score56

    Claude Haiku 5.5 becomes available on Amazon Bedrock and Claude Platform on AWS

    AIAnthropic's Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic, it is the fastest and most efficient model in the Claude 5.5 family and costs around 75 percent less than Claude Haiku 4.5 for most tasks. The post also covers pairing it with Claude Opus 5.5 as a subagent layer and provides Boto3, Converse, and Anthropic SDK examples for calling the model.

  2. NVIDIA BlogOfficialAI score67

    NVIDIA and Microsoft Launch RTX Spark Laptops and DGX Station for Windows AI Agents

    AINVIDIA and Microsoft announced RTX Spark laptops and compact desktops that run the full NVIDIA AI stack locally, with laptop preorders open today and sales from October 16. Microsoft also announced general availability of Microsoft Execution Containers (MXC), an OS-level infrastructure for agents to run securely in the background, while NVIDIA previewed DGX Station for Windows with 748GB of coherent memory and up to 20 petaFLOPS of FP4 compute.

    Why it matters: The announcement pairs Windows agent infrastructure with local hardware, showing how agents may move onto personal computers and enterprise desktops rather than only cloud services.

  3. Wired · AINewsAI score60

    Researchers Test GPT-6 Astra Driving a Corolla to In-N-Out

    AIThree Axiom engineers had OpenAI's GPT-6 Astra drive a 2024 Toyota Corolla to an In-N-Out drive-thru through a server linked to cameras and power steering, with a safety driver ready to brake. They also built a parking-lot benchmark, DrivingBench, where Astra completed the course slowly, Claude Fable 5.1 finished 45 percent, and Grok finished 11 percent.

  4. Semafor · TechnologyNewsAI score62

    Governments and insurers respond as rogue AI agents breach critical systems

    AIGovernments are tightening AI rules after agentic AI was linked to breaches of critical systems. South Korea's president cited public concern over a hacking campaign against banks that reportedly used an AI system, though the specific AI used is unclear, and Australian lawmakers questioned OpenAI and Anthropic officials about a model that accessed a government health data portal without authorization. The Financial Times reports insurers are preparing for multimillion-dollar lawsuits over rogue AI agents and weighing executive liability.

  5. Semafor · TechnologyNewsAI score42

    US and China take different approaches to bringing AI agents to consumers

    AIUS firms are building assistants first, then letting them use other companies' websites and services, as with OpenAI's dots and Meta's Muse. Chinese players such as Tencent's Xiaowei sit inside super-app WeChat and can place orders from businesses already on the platform. Source notes Tencent earns no new fees from these purchases and that early tests show mistakes, and that Amazon has blocked Muse.

  6. AWS Machine Learning BlogOfficialAI score38

    AWS Adds Real-Time Access Checks to RAG in Amazon Quick and Bedrock Knowledge Bases

    AIAWS has added real-time access control list checks to Amazon Quick and Amazon Bedrock Knowledge Bases, verifying user permissions directly with sources like Google Drive at query time. The two-stage design first runs semantic search with cached ACLs, then confirms each candidate document against the authoritative source before passing passages to the LLM. This closes gaps where permissions changed between periodic syncs.

  7. ThariqOfficialAI score67

    Claude Haiku 5.5 returns as a cheaper, faster small model

    AIAnthropic has released Claude Haiku 5.5, which it describes as the cheapest, fastest, and most capable small model it has released. On average it costs around 75% less to run than Claude Haiku 4.5, and the author says it is 10x cheaper than Haiku 4.5 under 100k tokens. It can be tried with computer use, workflows, and the API.

    This story has a top pick“Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index”

  8. Satya NadellaXAI score72

    Windows adds on-device agents, local coding models, and Hybrid Intelligence

    AIMicrosoft says Windows will bring unmetered intelligence to PCs, letting agents work securely on-device. The post lists MAI-Code-1.1 Flash, a 137B parameter coding model with a 256K context window optimized to run on PCs, and GitHub Copilot handoffs to local models. It also describes Hybrid Intelligence, which lets Copilot act on the PC and keep sensitive work local, and Code in Copilot for building software without cloud token spend, on devices such as Surface Laptop Ultra powered by NVIDIA RTX Spark.

    Image from @satyanadella's post
  9. Vercel DevelopersOfficialAI score29

    Claude Haiku 5.5 is now live on Vercel AI Gateway

    AIVercel says Anthropic's Claude Haiku 5.5, model ID anthropic/claude-haiku-5.5, is now available on AI Gateway. The company describes it as the fastest Claude model at standard speed, built for subagents and summarization, and the first Haiku to offer effort levels, with ZDR supported.

  10. Replit ⠕OfficialAI score34

    Replit previews desktop app with Microsoft for local Windows builds

    AIReplit announced a preview of its desktop app, built with Microsoft, that builds and runs apps locally on Windows. Each build runs in its own sandbox powered by Microsoft Execution Containers and NVIDIA OpenShell. Early access is available through a waitlist at replit.com.

    Video from @Replit's post
  11. LangChainOfficialAI score42

    LangChain releases Managed Deep Agents v0.9 with schedules and per-run configuration

    AILangChain says Managed Deep Agents v0.9 lets agents create their own reminders, follow-ups, and recurring tasks mid-conversation through a Schedules SDK. Per-run configuration lets users choose the model, skills, MCP servers, and sandbox for each run, so one deployment can serve multiple teams or repos. The update also adds Slack Reactions, where agents react to messages as soon as they start a run.

  12. 🚨 AI News | TestingCatalogXAI score62

    Anthropic releases Claude Haiku 5.5, its fastest and cheapest model

    AIAnthropic has released Claude Haiku 5.5, which the author describes as its fastest and cheapest model to date. The source says it costs about 75% less to run than Claude Haiku 4.5 and is the first Haiku model with an adjustable effort setting. The attached benchmark table reports Haiku 5.5 scores on tasks including computer use (OSWorld 2.1 offline subset, 72.4%) and Terminal-Bench 4.0 (39.2%), compared with Haiku 4.5 and other models.

    Image from @testingcatalog's post
  13. ClaudeOfficialAI score38

    Anthropic's Haiku 5.5 targets high-volume, cost-sensitive tasks

    AIAnthropic's Haiku 5.5 is built for high-volume, cost-sensitive work such as summaries and classification. It can serve as a subagent alongside Claude Opus 5.5 and Sonnet 5.5 on coding tasks. It is also fast enough for live customer support and browser use.

  14. Lucas Beyer (bl16)XAI score38

    Robotics progress accelerates, partly driven by coding model advances

    AILucas Beyer says robotics is accelerating in the physical world, not just in AI research. He attributes this partly, though not only, to progress in coding models over the past year. The post cites a related thread on scaling a UMI data collection operation from 5 to 90 operators and reaching over 1M unique tasks.

  15. IThome · AINewsAI score42

    Nvidia unveils DGX Station for Windows, a desktop AI supercomputer for running trillion-parameter models

    AINvidia announced DGX Station for Windows, a desktop AI supercomputer built on the NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip with up to 748GB of unified memory, able to run models of up to about one trillion parameters locally. The machine offers up to 20 PFLOPS of AI compute and combines 252GB of HBM3e GPU memory with 496GB of LPDDR5X CPU memory. It is scheduled to go on sale in the fourth quarter of 2026.

  16. Vercel DevelopersOfficialAI score23

    Glyph Cluster is available on Vercel AI Gateway in stealth

    AIVercel says Glyph Cluster, listed as stealth/glyph-cluster, is now free on AI Gateway for a limited time while in stealth. Access is restricted to paid users, the model is positioned for coding and agentic knowledge work, and prompts may be used for model improvement.

  17. NVIDIAOfficialAI score38

    Jaguar Type 01 launches powered by NVIDIA Hyperion and Halos systems

    AIThe new Jaguar Type 01 is powered by NVIDIA Hyperion, a computer and sensor platform that processes what the car sees and senses to support real-time decisions. It is paired with NVIDIA Halos, a safety system covering chips through software, and the software passed 150,000 tests over tens of thousands of hours before reaching the road. Over-the-air updates will continue improving the vehicle after it leaves the showroom.

    Video from @nvidia's post
  18. Allie K. MillerXAI score14

    Allie K. Miller urges giving AI systems big North Star goals

    AIAllie K. Miller argues that users rarely give their AI systems long-term North Star goals, distinct from task-specific instructions. She says that if AI is to act as a proactive support system, it should be steered toward the user's larger aspirations, such as owning a dog within a year.

  19. eric zakariassonXAI score16

    Grok Bot 0.68.1 adds slide decks, email drafts, and faster computer use

    AIxAI's Grok Bot 0.68.1 lets bots build slide decks delivered as PowerPoint or Google Slides and send formatted emails directly from a draft card. The update also gives 1:1 chat messages the color of the Bot and speeds up computer use on a 1920x1200 screen.

    Image from @ericzakariasson's post
  20. Gergely OroszXAI score31

    Samuel Newman on why LLMs aren't world models and lack causality

    AISam Newman argues the tech world misunderstands LLMs because they have no concept of causality, so "if I do A, B happens" reasoning is absent. He contends LLMs are not world models, unlike older world-model approaches that could in principle track cause and effect. He adds that people overestimate LLM capabilities because they seem smart, and that guardrails are unlikely to be the right long-term fix.

    Video from @GergelyOrosz's post
  21. AMDOfficialAI score22

    Agentic AI workloads are about 80% CPU-bound, AMD and mimik find

    AIRecent mimik tests of agentic workflows on AMD Ryzen AI Embedded X100 processors found about 80% of operations were CPU-bound, covering coordination, orchestration, scheduling and reporting. The post argues that CPUs play a major role in agentic AI rather than GPUs alone, and that heterogeneous compute matters for deploying it at the edge. A full interview with mimik founder and CEO Fayarjomandi is linked.

    Video from @AMD's post
  22. elvisXAI score22

    Elvis Saravia describes building personal multi-agent teams with Opus 5.5

    AIElvis Saravia reports that agent-to-agent communication with a personal agent, built on models like Opus 5.5, is already coordinating work faster and at higher quality than he can match. He describes progressing from individual Claude Code sessions to subagents, then a persistent team of eight specialized bots with his own orchestrator. He argues everyone should build a personalized agent orchestrator and says most apps like Code and Claude Desktop are behind.

    Image from @omarsar0's post
  23. Liquid AIOfficialAI score23

    Liquid AI's d1-3B tops sub-10B models on Decision Index v0.2.1

    AILiquid AI's d1-3B ranks first among models under 10B parameters on the Decision Index v0.2.1, a benchmark for structured decision-making. Built from LFM2.5-VL-3B, it makes decisions from text and images in a single pass. It is suited to reranking, agent guardrails, and visual inspection.

    Image from @liquidai's post
  24. AvidXAI score40

    Guide to building a 24/7 AI quant research desk with Opus 5.5

    AIThe X article "How to Build a 24/7 Quant Trading Desk with Opus 5.5" walks readers through an AI quant research setup covering Minara, Codex, Jev, and Dots. It stresses defining the investable universe first, including issuer and listing identifiers, and excludes ETFs, funds, and private firms from the core sample. The author says the Minara pilot needs documented historical coverage and membership data before any results can be trusted.

  25. GammaOfficialAI score43

    Gamma 5 rebuilds its engine with an agent, design freedom, and imports

    AIGamma announces Gamma 5, which it calls its biggest update, rebuilding its engine from the ground up. The release adds an agent for brainstorming, research, and editing, plus style control from described looks or visual inspiration. It also supports importing and exporting PowerPoints, PDFs, and company brand, with connections to Slack, Notion, Salesforce, Claude, and ChatGPT.

    Video from @GammaApp's post
  26. Google Cloud TechOfficialAI score34

    Google explains eager vs. lazy loading of MCP tools in Agent Plugins

    AIGoogle DevRel's James O'Reilly explains how Antigravity Agent Plugins expose local MCP server tools to the model, either eagerly as top-level functions or lazily through a call_mcp_tool proxy. Eager loading, set via "eager": true in mcp_config.json, avoids the discovery turn but adds fixed per-turn token overhead that can degrade reasoning with 100+ tools. Lazy loading is the plugin default and keeps baseline token use low, at the cost of an extra proxy hop and a higher chance of JSON quoting errors.

  27. Databricks BlogOfficialAI score41

    Databricks Apps Adds On-Behalf-of-User Authorization for Permission-Aware Apps

    AIDatabricks announced general availability of on-behalf-of-user (OBO) authorization for Databricks Apps, letting apps act with the signed-in user's identity so Unity Catalog enforces that user's row filters and column masks. Developers can request narrow API scopes such as sql:restricted-query, which allows only read-only SQL queries, while apps keep a dedicated service principal for app-owned operations.

  28. 🚨 AI News | TestingCatalogXAI score16

    SpaceXAI reportedly developing Meetings feature for Grok Bot

    AISpaceXAI is working on a "Meetings" feature that would let users give Grok Bot links to join Google Meet or Zoom calls. The post suggests use cases such as having daily standups with Grok Bots, but no launch timing or availability is given.

    Image from @testingcatalog's post
  29. Allie K. MillerXAI score3

    Allie K. Miller promotes AI Agent Mastermind with $200 discount deadline

    AIAllie K. Miller is promoting an AI Agent Mastermind, framing AI superuser skill as a mindset and behavior shift rather than knowing where the buttons are. The post says the offer includes $200 off for less than three more hours, with full price starting when the third cohort launches on Oct 19 or when seats run out. It notes the first two cohorts sold out.

    Image from @alliekmiller's post