Skip to contentSkip to stories

Updated

#Agent

Sep 17

Sep 17Thu
  1. OpenBMBAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

  2. Baidu Inc.AI score22

    Apollo Go plans to seek commercial autonomous driving approval in Hong Kong

    AIBaidu's Apollo Go plans to apply for commercial operation of its autonomous driving service in Hong Kong, citing the HKSAR Government's support in its first Five-Year Plan and the 2026 Policy Address. The company says it builds on fully driverless trials already conducted in the city. Baidu hopes Hong Kong can become a global benchmark for commercial autonomous driving in right-hand-drive markets.

  3. KrASIA · Big TechAI score50

    SenseTime's Lin Dahua Says Multimodal AI Breakthrough Could Come Within Two Years

    AISenseTime chief scientist Lin Dahua argues that native multimodal AI, which processes language, vision and other information in one shared model, is essential for AI to move beyond coding into industries and the physical world. SenseTime released the open-source SenseNova U1 in April and U1.5 Lite nearly four months later, and reported first-half 2026 revenue of RMB 2.91 billion, up 23.4% year-on-year. Lin's claim that a breakthrough could come within two years is the source's prediction, not a confirmed result.

  4. Gemini API ChangelogAI score38

    Antigravity Agent 09-2026 replaces 05-2026 with new built-in file and search tools

    AIGoogle released the antigravity-preview-09-2026 agent, which replaces and deprecates antigravity-preview-05-2026. Remote sandbox users reading only output_text or model_output steps need only update the agent string, while local-environment users or those parsing function_call steps must adapt to renamed tools, PascalCase parameters, and line-range file edits. The 05-2026 preview shuts down on October 5, 2026.

Sep 16

Sep 16Wed
  1. hardmaruAI score38

    Schmidhuber traces four decades of recursive self-improvement research to 1987

    AIJürgen Schmidhuber's new post surveys his recursive self-improvement (RSI) work since 1987, from self-modifying policies and the Gödel Machine to modern LLM agents. His background note says he published the first concrete RSI algorithms in 1987, when compute was about 100,000,000 times more expensive, and argues software RSI is now practical while full RSI will also require self-improving hardware in the physical world.

  2. Google Developers BlogAI score38

    Google and Speakeasy open-source OpenAPI SDK generator suite under AGPLv3 license

    AISpeakeasy is open-sourcing its full OpenAPI client suite under the AGPLv3 license, including generators for seven languages (Python, TypeScript, Go, Java, C#, PHP, Ruby), an agent-native CLI generator, and a documentation MCP server generator. Google said the move followed the May 2026 shutdown of the SDK generation provider it had been using, which it cited as evidence that closed-source generators pose platform risk. Google's new Google GenAI SDKs for the Interactions, Agents, and Webhooks APIs were built with this pipeline across six targets.

  3. Greg BrockmanAI score62

    Databricks rolls out Astra to all engineers, reports 60% higher coding spend

    AIDatabricks rolled out Astra to every engineer, about 3,500 people, after a pilot with around 200 users. Engineers given Astra increased coding spend by roughly 60% compared to baseline. The company reports Astra outperforms Opus 5 and Sol 5.6 on highly complex system design tasks, but sees no clear gain on medium or low complexity coding. Astra gets a separate sub-budget in Unity Gateway to encourage selective use.

  4. Google for DevelopersAI score38

    Three companies cut video AI token costs using Gemini agentic understanding

    AIMosaic, Ponder Studio, and Revyl tested Google's agentic video understanding with early-access Gemini Flash models for production video editing, B-roll selection, and mobile app testing. Mosaic reported a 97% cut in median token usage for its video agent Motion, and Ponder Studio reported a 0.967 F1 score with roughly 72% lower token costs when selecting B-roll. Revyl reported a 65% accuracy improvement in catching UI bugs and animation stutters, and the capability is available now via the Gemini API for video uploads and YouTube videos.

  5. Matei ZahariaAI score44

    Agent harness choice strongly affects coding cost, not task success rate

    AIMatei Zaharia says agent harnesses make a large difference in cost, even on open-source coding benchmarks, and Melissa Pan's research examines why. Her quoted evaluation of seven models across Claude Code, Codex, and Pi found harness choice had little effect on task success but significantly affected cost. A simple harness can be competitive, and the native harness is not always the best.

  6. Baseten BlogAI score54

    Baseten launches Hosted Tools with web search for open-source models

    AIBaseten has launched Hosted Tools, starting with Baseten Grounded Inference, a server-side web search capability for models hosted on Baseten. Developers enable it by adding a hosted search tool to a Messages, Chat Completions, or Responses request, and the platform runs the search loop with partners Exa, Keenable, Parallel, and You.com. In Baseten's benchmarks, agents using the hosted tools saw a 15% reduction in end-to-end latency compared with client-side tools, and the feature is in playground preview with 25 RPM rate limits and $2 of free credits.

  7. Cat WuAI score60

    Claude merges Cowork and chat into one product with automatic routing

    AIAnthropic is merging Claude Cowork and chat into one Claude, and Claude Design is integrated so users can ask for slides, designs, or docs without switching apps. Claude decides from the prompt whether to give a quick answer or do deeper agentic work, and users can still stop, redirect, or adjust its effort. The change rolls out to Pro and Max over the next few weeks.

  8. Mike KriegerAI score46

    Claude Cowork and Chat Merge into One Unified Claude

    AIAnthropic is merging Claude Cowork and Chat into a single Claude starting today, which Mike Krieger says removes the friction of choosing which product to start with. Per the @claudeai announcement, Claude will carry tasks forward even after the laptop is closed, asking for clarification when needed while users keep final say. The rollout to Pro and Max plans will take place over the coming weeks.

  9. X.PINAI score49

    Shengyu Liu warns AI could turn programming into a hobby, not a profession

    AIFormer DeepSeek kernel engineer Shengyu Liu argues that AI industrializing software production could reduce programming to a recreational craft and erode students' engineering skills. His central concern is less whether AI can outthink humans than whether access to it stays widespread or gets concentrated in a few corporations. The post, cited by X.PIN, contrasts this with Western warnings about AI escaping human control.

Sep 15

Sep 15Tue
  1. Noah ZwebenAI score17

    Anthropic offers Claude Tag office hours for on-call triage feedback

    AIAnthropic is hosting office hours for teams interested in using Claude Tag for on-call work, and it is asking Team or Enterprise plan users to share triage feedback. Claude Tag can start investigating when a Slack alert fires by pulling metrics, diffing deploys, and checking flags to propose a likely cause and fix. Sign-up is through a Google Calendar booking link.

  2. Zed BlogAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  3. Claude Apps Release NotesAI score72

    Claude Cowork moves into every conversation, adding designs, slides, and docs

    AIClaude now makes Cowork capabilities available from any conversation without choosing a mode first, with chats, tasks, projects, connectors, and skills carrying over. Users can also create designs, decks, and docs in any conversation, including Claude Code and the Artifacts tab, and edit them with Claude.

    Why it matters: The release merges Cowork tasks into ordinary chats and adds design, slide, and doc creation, changing how Claude users start larger work.

  4. Jazzyear · InsightsAI score67

    HiDream's vivago R1 agent targets five-minute AI video delivery

    AIHiDream.ai launched vivago R1, a content creation agent, globally, with a domestic version upgrade. The company says R1 can output five-minute high-quality videos through agent planning, with a claimed 85% usable-output rate and support for multi-round extensions. It also released HiDream-O1-Video-1.0, a native omni-modal video model supporting single shots of 5 to 20 seconds at 1080p.

  5. Google Developers BlogAI score46

    Google Launches Agent Anomaly Detection in Private Preview on Gemini Enterprise Agent Platform

    AIGoogle has put Agent Anomaly Detection into Private Preview on the Gemini Enterprise Agent Platform, a reasoning-based audit layer that reviews agent reasoning traces, tool calls, and execution flow to flag behavioral anomalies and policy violations. It runs asynchronously without adding runtime latency and publishes findings to Security Command Center. The preview requires ADK 1.2 or later.

  6. Mark ZuckerbergAI score30

    Zuckerberg says labs should prioritize alignment and safety as core capabilities.

    AIMark Zuckerberg argues that every AI lab has both the incentive and responsibility to train models safely, since users will reject misaligned agents and labs face liability for harm. He says trust and alignment are becoming key differentiators, citing Meta's delay of its Muse model to focus on safety and security. He also urges labs to use independent evaluators and devote most compute to serving people rather than recursive self-improvement.

  7. Google AI StudioAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    AIGoogle AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    Why it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

  8. Microsoft Foundry BlogAI score32

    Microsoft Launches Foundry Dev Pack to Install Foundry Development Tools in One Command

    AIMicrosoft has launched Foundry Dev Pack, an all-in-one installer that sets up tools for Microsoft Foundry development across the terminal, IDE, and coding agents. Depending on the environment, it installs Azure CLI (az), Azure Developer CLI (azd) with the Microsoft Foundry Extension for azd, the Microsoft Foundry Skill, the Microsoft Foundry Toolkit for Visual Studio Code, and Foundry Canvas (preview), with the last two conditional on VS Code or GitHub Copilot App being present.

  9. Greg BrockmanAI score46

    ChatGPT Work adds Data agent for dashboards and actions on company data

    AIOpenAI's Greg Brockman says ChatGPT Work can operate over and act on a company's data, including building dashboards, by connecting existing tools such as PowerBI, Tableau, Clickhouse, Oracle BI, and AWS Redshift. The linked ChatGPT announcement describes a Data agent with a Data Plugin that turns company data into answers, interactive dashboards, and actions through conversation.

  10. Google AIAI score72

    Google rolls out Gemini 3.8 Live and Extended Thinking across consumer, developer, and enterprise channels

    AIGoogle is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking across several channels. Consumers get them in Search Live and Gemini Live, developers get public preview access through the Gemini API, and enterprises get private preview through Gemini Enterprise, with Customer Experience support coming soon.

    Why it matters: The post lays out where each Gemini 3.8 Live variant reaches consumers, developers, and enterprises, which clarifies access paths for a voice model release.

  11. Google DeepMindAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    AIGoogle DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    Why it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  12. Google AI StudioAI score72

    Google launches Gemini 3.8 Live and Extended Thinking voice models

    AIGoogle introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents that reason and speak simultaneously. The Extended Thinking version scores 82.6 on Artificial Analysis' Speech to Speech Quality Index and 97.7% on Big Bench Audio, while 3.8 Live targets scale and cost efficiency. Developers can access both through the Gemini API in Google AI Studio, and enterprise and consumer rollouts vary by product.

    Why it matters: The source names the two models, their access paths, and specific benchmark results, showing how the voice agent capabilities differ between the two tiers.

  13. Cognition Blog (Devin, Windsurf)AI score60

    Cognition and AWS sign multi-year deal to deploy Devin for enterprise modernization

    AICognition and AWS have entered a multi-year Strategic Collaboration Agreement to help enterprises deploy the Devin autonomous engineer in production. Devin can be purchased through AWS Marketplace, and the companies are exploring deeper engineering integrations within customers' AWS environments. Mercedes-Benz reportedly used Devin to analyze more than 200,000 lines of COBOL, reducing an estimated eight-month modernization project to eight days.

    Why it matters: The collaboration shows how an autonomous coding agent is being packaged for enterprise legacy modernization inside existing AWS environments, with concrete customer migration figures.

  14. Baseten BlogAI score40

    LangChain uses Baseten Loops to train custom models for LangSmith Engine

    AILangChain is using Baseten Loops, a managed fine-tuning service, to train custom models for LangSmith Engine, its in-platform agent that debugs and improves AI agents. The article says LangChain fine-tunes large open-weight models on agent traces and trains smaller open-weight models such as Qwen for tasks like failure-mode categorization. Baseten Loops supports supervised fine-tuning, reinforcement learning, and long-context workloads, and lets checkpoints be evaluated and deployed directly to inference.