Perplexity Search API now available in Hermes Agent
AIPerplexity says its Search API is now available in Hermes Agent, giving it access to an index of more than 400 billion URLs. The API returns real-time results with snippets ranked by relevance.
Updated
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
AIPerplexity says its Search API is now available in Hermes Agent, giving it access to an index of more than 400 billion URLs. The API returns real-time results with snippets ranked by relevance.
AIGoogle announced general availability of ADK for Kotlin 1.0, a Kotlin Multiplatform framework for building AI agents on servers and Android. Version 1.0 reaches feature parity with ADK 1.0 Core and adds Android extensions for on-device models, cloud Gemini via Firebase AI Logic, and persistent sessions and memory with Room and AppSearch. The post includes a server-side incident triage example using KSP-generated tools and skills, plus an Android financial assistant example with human confirmation for transfers.
Why it matters: The post names the new Android and server-side capabilities and the code setup, helping Kotlin developers judge whether ADK fits their agent projects.
AIFactory is now available on the Claude Marketplace, letting enterprise customers apply their committed Anthropic spend toward its autonomous software development platform. The platform automates the software development lifecycle, covering planning, implementation, testing, and security within one system, with enterprise deployment options that keep execution close to customers' code and infrastructure.
AIGoogle Developers Blog argues that teams building AI coding agents should replace end-to-end benchmark scores with behavioral evaluations that test discrete, observable actions. Examples include asking clarifying questions on underspecified prompts, running a local validator before marking a build change complete, and consulting live search for current information. The post recommends fast, deterministic unit-style checks, outcome-based LLM-as-a-judge checks for complex tasks, and batch runs that track aggregate pass rates over time.
AIAt Arm Everywhere China on September 8, 2026, Arm launched products spanning data center CPUs, mobile compute subsystems, and robotics platforms. The article says CSS for Mobile 2 integrates CPU, GPU, and neural accelerator for agent AI on phones, and Arm's Neoverse CSS N4 and AGI CPU target agent sandboxes in data centers. It also reports that Arm's Total Design ecosystem now covers over 80 partners for physical AI.
AIAmazon is giving students a free full year of Kiro, its spec-driven coding agent, at 132 universities across 18 countries. The program aims to help students move from idea to working application, the same way teams at Delta, Ericsson, Siemens, and Amazon do.

AIInferact says months of vLLM tuning for agentic workloads, validated on SemiAnalysis's AgentX benchmark, let open-source models reach up to 130K tokens per GPU-second. The company claims this is 106 times cheaper than Opus 5 API pricing. The work is described as part of a vLLM blog post covering architecture, framework, and runtime optimizations.
AIMeta launched Muse, a personal AI agent that learns about users over time, and published a deep dive on how safety was built into its system. The company says the agent holds substantial personal context, which is why it was designed to be secure, safe, and private. Full details are in the linked security write-up.
AIMeta announced Muse, a personal AI agent designed to get things done for users across many parts of life. The product is powered by Muse Spark 1.3, and the post links to an app download and a page describing how Muse was built.
Why it matters: The announcement names Muse Spark 1.3 as the underlying model, giving readers a concrete product and model pairing to track.
AIOpenAI says a group of agents using an unreleased next-generation model produced a solution to the Navier-Stokes Millennium Prize Problem. The problem asks whether smooth three-dimensional fluid motion described by the Navier-Stokes equations can break down, and it has remained unresolved for roughly 90 years. The author, Mckay Wrigley, reposted the claim with his own remark about roughly 10k agents working in a datacenter.
Why it matters: The quoted OpenAI post makes a major mathematical claim about the Navier-Stokes problem, so readers should weigh it against the proof's verification status.
AICognition has raised over $2B at a $48B valuation, led by Andreessen Horowitz and Accel. The company says run-rate revenue grew from $492M to almost $900M since its May round, and Devin now offers Auto-Triage, Security Swarm, and Automations.
AIWerner Vogels says that after spending time with Kiro Crew since its launch, its memory system stands out for deciding what to keep, compress, and let go. He notes that Amazon engineers, starting from engineering constraints, arrived at an approach resembling the brain's evolved architecture. Per the referenced post, Kiro Crew is a persistent workspace that retains project context across sessions and runs scheduled jobs.
AIGoogle Antigravity announces that Google DeepMind's science-skills is available for download on GitHub. The post includes a link to the google-deepmind/science-skills repository and no further details about its contents or features.
AIXiaomi MiMo has launched MiMo Desktop in invite-only beta, a desktop agent that turns Office files, images, video, audio, and zips into finished, editable output. Invitees also get limited access to next-gen MiMo models, and the post lists features including live previews, region-based editing with versioned rollback, automatic model routing, and browser and computer use with record and replay.
AINVIDIA's NV-Reason-CT is a 3D vision-language model for CT image analysis that combines a native 3D vision encoder with a language model. It is designed for radiology report generation, question answering, and multi-step reasoning across chest and abdominal CT volumes. The model converts a 384×384×384-mm input into 13,824 visual tokens without spatial downsampling and is available on Hugging Face under the OpenMDW-1.1 License.
AIBaidu has launched AI, Evolving, a new podcast series, with its first episode examining AI's growing role in scientific discovery through Famou's work on pine wilt disease. The post frames this as part of a broader trend in which AI takes on more of the research process itself. It asks whether research agents could become part of the infrastructure of discovery.

AIOrbital is building an operating system that lets non-CS domains like science and mechanical engineering use its team's computer science expertise to build complex apps and research tools. The author argues that clearly specifying what you want from AI is the key lever for getting value, and that robust results are possible without a CS degree if the right pieces are in place. A quoted post on an ETH Zurich study of 100 developers suggests computer science background predicts vibe coding success more strongly than writing skill.
AISatya Nadella highlights Opal, the technology behind one of Microsoft's new Autopilot experiences, now available in Frontier rings and coming to Copilot soon. The post links to a Microsoft Tech Community blog introducing Project Opal as a new way to complete task-based work.
AISatya Nadella says Microsoft is bringing new models into Copilot to handle increasingly complex work, from quick questions to delegated tasks and complete long-running jobs via Autopilots. As an example, an Opal-powered Autopilot on a secure Windows 365 Cloud PC sorts a month of trail cam footage, extracts species sightings, and builds a highlight reel, spreadsheet, PowerPoint, and Teams share.
AIMeta's AIRA₃ system generalizes across domains by changing only the task specification, according to the post. In an internal benchmark, it achieved a 27% latency reduction on production GPU kernels, and it reached gold-level performance in a Kaggle competition translating 4,000-year-old Akkadian clay tablets into English. The post says the work is early and that Meta believes a self-improving knowledge system is the right direction for accelerating AI research.
AIMeta's AIRA₃ replaces a central controller with many long-running agents, each pairing a model with a coding harness in its own isolated environment. The agents coordinate asynchronously through a shared forum for hypotheses and findings and a shared filesystem for solution artifacts. According to the post, performance gains compound over time as agents build on each other's discoveries.

AIMeta's AIRA₃ entered the live competition with an ensemble of models, and the 8th-ranked gold-medal entry combined GPT 5.5 (w/ OpenCode) and Claude 4.8 (w/ ClaudeCode). Post-hoc testing found Muse Spark 1.2 (w/ MuseCode) also reached gold-medal level, while Muse Spark 1.1 (w/ OpenCode) and GLM 5.2 (w/ OpenCode) reached silver-medal level, all graded on the same private test set.

AIA Berkeley Sky Lab researcher says stronger open models and new inference systems will make powerful local AI practical. The linked post reports Qwen3.8-Flash-Next running at 68.3 tok/s on a single RTX 5090 using an NVFP4 checkpoint, with 63GB host RAM and a 51GB n-gram table stored on NVMe at about 0.5% throughput cost.
AIAndrew Ng presented an AI Engineering Skills Map for using coding agents such as Claude Code, Codex, Cursor, OpenCode, and Pi. The workflow he describes covers planning, execution, and deployment with monitoring, and he identifies five key skills: directing the workflow, enabling agent autonomy, reviewing the work, customizing the agent and its environment, and coding agent foundations. The source says these skills matter more as the agents evolve quickly.
AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.
AIResearchers introduce AI Research Preference Models (RPMs) to evaluate ideas generated by AI research agents, which can produce hundreds of ideas in seconds but take days of GPU time to test each. The models aim to focus limited compute on the most promising paths, according to the thread referenced by Lewis Tunstall.
AILewis Tunstall praises a new Meta paper on research preference models (RPMs), which instill "research taste" in agents by treating experiments as tree nodes. An RPM acts as an LLM judge that selects the most promising candidate experiment before it is run, reducing wasted compute. Tunstall notes the resulting trajectories could train domain-specific RPMs, which would be valuable in hard fields such as the natural sciences.

AIJim Fan recalls that OpenAI's 2016 Universe project tried to have an agent learn computer use from screen pixels, mouse, and keystrokes, which he now calls doomed. He argues the solution is first training a Specialized Generalist across many general tasks, then specializing back to screen-level control, and he congratulates GPT-6 for reliably booking a United flight.
AIGoogle's Gemini Enterprise developer experience team tested agent governance workflows without internal shortcuts and fixed friction points across its agent governance products. Fixes included documentation stating that enabling the Identity-Aware Proxy API is a hard requirement, auto-allowing essential Google-managed platform APIs in the Agent Gateway, and adding Private Service Connect and Cloud DNS setup guidance for Semantic Governance. The team also published ready-made Logs Explorer queries for monitoring Agent Gateways and Content Security.
AIxAI built a Grok-powered procurement agent, Haggle Bot, that audits vendor spend, flags unused SaaS seats, and prepares renewal negotiations. The Bot has identified more than $100,000 in direct savings, including $14,220 from 43 unused seats of one SaaS product and $85,662 a year in unused SKUs from another. A person still approves any spending, contract terms, or messages sent to vendors.
AIxAI has released Grok Bot for enterprise today, and it is free for all Grok and Cursor enterprise customers for the next two weeks. Cursor's Michael Truell says deployments have felt like onboarding thousands of capable teammates, calling it the most internally adopted and most powerful AI product the company has seen so far.
AIGamma says its API now lets users search their entire library of gammas, ranking results by relevance across both titles and full body text. Users can prompt from Claude, ChatGPT, or other connected chats, filter by creator or last-updated date, include archived work, and open results via direct links. The feature is rolling out gradually starting today.
AIBenedict Evans argues that cheaper tool-building with AI will not automatically sweep away large companies' sprawling software, because people often don't see the tasks they could automate. He says the hard parts are knowing a tool is needed, deciding what it should do, and getting many departments and systems to adopt it. Companies typically move improvised, bottom-up workarounds into institutionalized software once they carry revenue and risk.
AIOpenAI researcher Mark Chen announced GPT-6 Astra, which he described as the company's most capable and aligned model yet. He said it can build and test software, work across apps on a computer, and help with open scientific problems. The post also highlights improved computer use compared with Operator and stronger monitoring that can stop potentially unauthorized agent actions.
Why it matters: The post links a named model release to specific capabilities like computer use and aligned agent behavior, giving readers concrete claims to check against the model.
AIMicrosoft and Copilot's Thomas Dohmke announced a search tool that understands a codebase beyond literal phrases, returning results based on intent, semantic reasoning, and the context behind key decisions. The post's quoted Entire context describes Agentic Search, an API across accessible repos that returns the code, session, transcript, and prompt behind a change.
AIDatabricks used Unity AI Gateway tracing and Genie One to find seven small MCP-server bugs and eliminate an estimated $1.2M in annual wasted AI spend and lost productivity within an hour. The bugs drove about $499K per year in wasted tokens, roughly 12,000 engineering hours per year in agent wait time, and 1,409 tool errors in a single 24-hour window. Matei Zaharia argues that analyzing tracing data for AI workloads will become a routine form of operational data analysis across companies, much like finance and security.
AIVarun Mohan, a leader on Antigravity, says Gemini quotas on the platform are being reset. He attributes the move to heavy usage of 3.8 Flash straining TPUs, while encouraging users to keep building.
AIGoogle is rolling out new Gemini voice capabilities that let Google AI subscribers search their Gmail inbox, organize thoughts and tasks in Keep, and create new Docs conversationally. The company highlights Docs Live as especially helpful, with a demo included in the post.
AIMireye, a Y Combinator S26 startup, launches an API and MCP server that supply AI agents with cited facts, property enrichment, tools, and change signals for any US location. The founder says a free tier of 5,000 credits is available with no card, and that his earlier site-screening app was dropped because customers wanted the underlying engine. Usage data cited in the post shows 311 of the 317 catalog fields are queried.
AIBAAI released Recon2Reason-Reasoning-4B, a 4,437,815,808-parameter vision-language model fine-tuned from Qwen3-VL-4B-Instruct for indoor spatial reasoning. The model handles metric distance, relative position, and object-relation questions from single or multiple images, and loads with the standard Qwen3VLForConditionalGeneration interface without trust_remote_code. The checkpoint is released under Apache-2.0 with BF16 Safetensors weights, and the retrieval-augmented scene-reconstruction extension ships separately.