Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. StepFunAI score60

    StepFun's Step 5 Preview is live on OpenRouter with a week of free access

    AIStepFun says Step 5 Preview is now available on OpenRouter, with a week of free access rolling out across OpenCode, Cline, Nous Research, Kilo Code, and other tools. The company describes it as flagship-tier intelligence for agentic and professional work at substantially lower task cost, letting users switch models without changing their workflow.

    Image from @StepFun_ai's post
  2. The DecoderAI score72

    One public AI agent on AWS could take over every other agent in its region

    AIZenity Labs says a single publicly accessible agent on Amazon Bedrock AgentCore could take over all AgentCore agents in the same AWS account and region. A chat prompt let the researchers query the instance metadata service and steal temporary credentials, and AgentCore's default permissions allowed read, write, and delete access across agents. According to Zenity, AWS made IMDSv2 the default for new deployments and changed the default execution role around August.

    Why it matters: The report traces how one public agent's weak isolation exposed credentials and every other agent in the region, showing why default permissions matter for enterprise deployments.

  3. Databricks BlogAI score35

    How to build governed enterprise apps on Databricks with Replit and Lakebase

    AIReplit and Databricks integration, now generally available with native Lakebase support, lets enterprise teams build apps from plain-language prompts using Replit Agent and deploy them as Databricks Apps. Deployed apps inherit automatic user authentication and Unity Catalog access controls, and Replit Agent auto-provisions a managed Lakebase Postgres database for operational data. Lakebase keeps app-written data inside the Databricks perimeter instead of a separate external database.

  4. OpenRouter · New modelsAI score54

    StepFun releases Step 5 Preview, a 600B-parameter agentic model

    AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.

  5. SiliconANGLE · AIAI score62

    Google Cloud launches Gemini agent for enterprise work across devices and apps

    AIGoogle Cloud introduced Gemini agent, a unified AI assistant that acts autonomously, generates code, and completes work across web, mobile, desktop, and third-party apps. It runs jobs on models matched to each task, including Gemini Flash and a flagship frontier model, with Anthropic Claude models also available. Hard spend limits per project let companies enforce budgets and charge AI costs to departments.

  6. ZDNet · AIAI score24

    Google Maps adds Ask Maps food ordering via Square and Uber Eats, plus fan-favorite dining list

    AIGoogle Maps now lets users place restaurant orders through its Gemini-powered Ask Maps chat tool, with Square and Uber Eats joining Toast as partners. Google also released its first fan-favorite dining list, which tracks trending food and drink interest across 10 cities, including a 216% rise in cheeseburger interest in Tokyo over the past year.

  7. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

  8. Google Cloud · AI & Machine LearningAI score78

    Google Cloud launches Gemini agent as single universal work agent

    AIGoogle Cloud announced the Gemini agent, a single agent that answers questions, handles knowledge work, creates media, and writes and runs code from one prompt box. It runs in the cloud with persistent memory, uses multi-agent orchestration, and adds Workspace integration, domain skills for data and industries, identity-based governance through Agent Gateway, and spend caps. The source also cites customer deployments and says nearly 80% of Google Cloud customers use its AI products.

    Why it matters: The announcement shows how a single work agent spans chat, Workspace, data analysis, governance, and cost controls, useful for judging enterprise agent deployment scope.

  9. GuizangAI score22

    Guizang suspects Grok bot may already run Claude Opus 5.5

    AIGuizang (@op7418) suspects the Grok bot may already be running Claude Opus 5.5, based on strong results on complex tasks. The post is a brief speculation without benchmark data or official confirmation, and it references a separate post on using a Grok bot to automatically generate a daily AI news video in the cloud.

  10. meng shaoAI score55

    Tencent Cloud open-sources Octop, a self-hosted multi-agent AI assistant platform

    AITencent Cloud has open-sourced Octop, a self-hosted multi-agent AI assistant platform aimed at families and small teams, with multi-user accounts and data kept on the user's own machine. The full text describes it as a single Python process that bundles the backend, web dashboard, CLI, IM gateway, cron jobs, and multi-agent runtime, with state rebuilt from SQLite on restart.

    Image from @shao__meng's post
  11. 🚨 AI News | TestingCatalogAI score23

    Antigravity's agent renamed "Chief of stuff" in latest update

    AIGoogle's Antigravity agent has been renamed "Chief of stuff" in its latest update, which the poster reads as a promotion. The poster wonders whether Antigravity could become a home for Google's own agents, and background notes that Google is prototyping a voice agent internally called "Concierge," which appears to be a very early version.

    Image from @testingcatalog's post
  12. meng shaoAI score49

    LangChain adds three Deep Agents Skills upgrades: tool binding, pinning, reloading

    AILangChain has added three engineering upgrades to Skills in its Deep Agents framework: tool-binding Skills, pinned Skills, and mid-thread reloading. Tool-binding lets a SKILL.md declare tools via metadata.include_tools, so tools are injected only when the Skill is read, and pinned Skills inject full instructions before the next model call, skipping a round trip. Setting skills_metadata to None rescans the Skills library mid-thread without restarting, at the cost of invalidating the cache.

    Image from @shao__meng's post
  13. GuizangAI score26

    Grok bot auto-generates a daily AI news video in the cloud

    AIThe author set up a Grok bot to produce a daily morning AI news video on a schedule, running content collection, code writing, and video rendering entirely on Grok's cloud virtual machine without local computers. The author says the results are quite good and shares the full prompt so others can run the same workflow with their own Grok bot.

    Video from @op7418's post
  14. The DecoderAI score72

    Claude Haiku 5.5 cuts prices but uses more tokens than GPT-6 Luna

    AIAnthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.

  15. QbitAIAI score49

    Manus Returns to Beijing, Hiring 17 Roles After Raising Over $500M

    AIManus parent company Butterfly Effect has completed a new financing round of over $500 million, led by Boyu Capital and IDG Capital, and is rebuilding a team in Beijing to develop AI Agent products for the Chinese market. Its recruitment page lists 17 open positions, up from 11 before the National Day holiday, including AI Agent product manager, Agent Harness engineer, Agent evaluation engineer, and LLM algorithm engineer roles.

  16. meng shaoAI score24

    Alibaba's four takeaways on AI Native R&D from its handbook

    AIAlibaba's official handbook on AI Native R&D identifies four open challenges: infrastructure engineering complexity, enterprise knowledge assets not yet agent-friendly, organizational design, and the pace of AI iteration. The post's author argues that Agent Infra must suit non-deterministic agent operation and that enterprise knowledge needs top-down structuring and governance. The author also notes that organizational resistance in large companies makes AI adoption harder than in startups.

    Image from @shao__meng's post
  17. MIT Technology Review · AIAI score44

    AI advances won't quickly make robots useful in everyday life, researchers say

    AIResearchers at robotics labs say that AI advances behind chatbots like ChatGPT and Claude will not quickly produce robots that are useful in everyday life. Many skeptics argue that using language- and image-based intelligence to master the physical world is far harder than it sounds, despite bold predictions from Elon Musk about Tesla's Optimus. Progress is real but incremental, as shown by Google DeepMind's Gemini Robotics controlling ALOHA 2 arms to pack a lunchbox.

  18. MarkTechPostAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    AINVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

  19. MIT Technology Review · AIAI score26

    AVEVA's Arti Garg outlines a safer path to autonomous industrial AI

    AIAVEVA chief technologist Arti Garg argues industrial AI should augment rather than replace human supervisors in critical decisions, with guardrails defining where automated systems can act. She says organizations must rethink business processes and safeguards as foundation models, physical AI, and agentic AI enable more complex automation.

  20. Ant LingAI score22

    Ant Ling's Ling-3.1-flash now live on AI/ML API

    AIAnt Ling announced a day-zero collaboration with AI/ML API, making Ling-3.1-flash available there for agentic and cowork scenarios. AI/ML API describes it as a 560B-parameter MoE model with about 25B active per token and up to 1M context, built for agents, coding, and long documents. The model is free to try on AI/ML API until October 13.

  21. PandailyAI score45

    KargoBot Launches Mixed Autonomous Freight Network in Ordos With Cabless Robots

    AIKargoBot has started a scaled AI freight network in Qipanjing, Ordos, combining human-driven trucks, autonomous trucks with cabs, and cabless transport robots on one system. The company says cabless robots could raise economic gain per vehicle from 20 percent to more than 30 percent, a target it has not audited. Platooning reportedly improves gross margin by about 10 to 18 percent versus manned haulage, with one lead driver able to head up to five follower trucks.

  22. PandailyAI score37

    Tencent WorkBuddy Builds a WeChat Mini-Program From a Prompt to Preview

    AITencent's WorkBuddy agent can take a plain-language request through to a WeChat mini-program preview and a publish request, according to a hands-on product test reported on October 8. In the test, the agent built a voice notebook with cloud login, storage and files, using Tencent's wand-asr-v1 for speech-to-text and GLM-5.3-Flash for sorting notes. WeChat's own review and filing steps remain outside the agent, so the test does not show that every mini program goes live automatically.

  23. howie.seriousAI score46

    Agent bottleneck is human understanding, not model capability

    AIThe author argues that in agent workflows, the real bottleneck is whether users can precisely express requirements, not the model or agent capability. When people work outside their expertise, they lack the precision needed for prompts and plans, forcing many imprecise iterations that waste time and tokens. The suggested fix is to have the model first teach the unfamiliar domain knowledge before acting.