Skip to content

#Agent

Oct 8

TodayOct 8Thu226 items
  1. NVIDIA BlogAI score49

    How Developers Use Frontier AI Agents to Build Omniverse Simulations

    Developers are pairing frontier AI models with NVIDIA Omniverse libraries to turn simulation ideas into working applications, from humanoid warehouse simulators to autonomous-driving test environments. In the examples, developers direct AI agents through natural-language instructions and review results, while Omniverse provides GPU-accelerated physics, rendering and sensor simulation. One experiment reported a simulated Unitree G1 humanoid clearing a hurdle in 64 of 100 trials.

  2. ThariqAI score22

    you now get Claude API credits with your MAX plans ($100, or 200 matching your plan) every month, use this to build more personal AI for yourself! I made an AI homepage that's generated everyday based on sites I read and replaces my 'new tab' page in Chrome

    you now get Claude API credits with your MAX plans ($100, or 200 matching your plan) every month, use this to build more personal AI for yourself! I made an AI homepage that's generated everyday based on sites I read and replaces my 'new tab' page in Chrome

  3. NVIDIA Technical BlogAI score26

    How to create SimReady robotics assets from CAD with frontier AI models

    NVIDIA's Omniverse libraries, guided by SimReady Foundation specifications and agentic NVIDIA skills, provide a structured workflow for converting CAD assets to OpenUSD for robotics simulation. The workflow covers configuring and validating materials, collision geometry, joints, and other physics properties before testing robot behavior.

  4. Sara HookerAI score17

    @adaption_ai discovery agent reviews past AutoScientist training runs and identify critiques. These end up being at the edge of current capabilities and domain specific. We use that checklist to filter training data automatically.

    @adaption_ai discovery agent reviews past AutoScientist training runs and identify critiques. These end up being at the edge of current capabilities and domain specific. We use that checklist to filter training data automatically.

  5. Higgsfield AIAI score34

    Introducing Higgsfield Katana, powered by Claude Motion. Our most powerful AI video editing tool, inside Claude. Upload a reference and create editable motion graphics, product launch videos, or aura-farming edits. Available now in Claude via Higgsfield MCP.

    Introducing Higgsfield Katana, powered by Claude Motion. Our most powerful AI video editing tool, inside Claude. Upload a reference and create editable motion graphics, product launch videos, or aura-farming edits. Available now in Claude via Higgsfield MCP.

  6. SiliconANGLE · AIAI score62

    Manus raises over $500M at reported $4B valuation after Meta deal collapsed

    Manus, the developer of the Manus AI agent, has raised more than $500 million led by Boyu Capital, with Tencent also investing. Bloomberg had reported a $4 billion valuation, roughly double Meta's reported offer last December, which Chinese regulators blocked in April. The funding comes days after Manus 2.0 added a new harness, Cloud Computer, and Cue, which lets agents use email accounts and digital wallets.

  7. SiliconANGLE · AIAI score24

    CoreWeave Pitches Open Full-Stack AI Cloud With Forge Development Platform

    CoreWeave is positioning its AI cloud around an open development loop, connecting training, inference and evaluation through its newly announced CoreWeave Forge platform. Chief marketing officer Jean English said the company wants production learnings to improve models and agents and that the loop should work across different models, frameworks and clouds. She argued that competitive differentiation extends beyond GPUs to partner tooling, infrastructure and APIs.

  8. ClaudeDevsAI score42

    ICYMI we've cut the price of Sonnet 5.5 cache reads in half. In the Claude Platform, it's now $0.10 per million tokens (input is $2, output is $10). This means Sonnet 5.5 now runs ~20% cheaper on most agentic work. (API only, no change to Claude Code usage limits.)

    ICYMI we've cut the price of Sonnet 5.5 cache reads in half. In the Claude Platform, it's now $0.10 per million tokens (input is $2, output is $10). This means Sonnet 5.5 now runs ~20% cheaper on most agentic work. (API only, no change to Claude Code usage limits.)

  9. Claude Code · GitHub ReleasesAI score56

    Claude Code v2.1.295 adds hook failure blocking and gateway controls

    Claude Code v2.1.295 adds onFailure: "block" for command and HTTP hooks, so a hook that cannot start, times out, or exits unexpectedly blocks the action. The release also adds an optional models list for Claude apps gateway upstreams, plus upstream_request_id in the inference audit event, and fixes a range of MCP, plugin, and terminal issues.

  10. Artificial AnalysisAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    Artificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now ranks first. The model is available only through OpenAI's Daybreak program and records no safety blocks across the Index. Its overall score is 32 points higher than the publicly available GPT-6 Sol (max), at a cost of $1.77 per task versus $11.67 for Grok 4.7 (xhigh).

  11. The DecoderAI score62

    Anthropic's updated usage policy bans sustained abusive behavior toward Claude

    Anthropic has updated Claude's usage policy for the first time in over a year, banning sustained and needless abusive or cruel behavior toward Claude. The company says ordinary frustration, pushback, dark creative themes, and model testing are not covered, and that the rule applies only in extreme cases. Violations can lead to warnings, throttling, restriction, suspension, or termination of access.

  12. ClaudeAI score13

    When a question needs deeper analysis, send the dashboard to your analytics tool. When an animation needs finishing touches, open it in your video editor. See the full list of supported tools: https://claude.com/resources/articles/dashboards-and-motion

    When a question needs deeper analysis, send the dashboard to your analytics tool. When an animation needs finishing touches, open it in your video editor. See the full list of supported tools: https://claude.com/resources/articles/dashboards-and-motion

  13. Codex · GitHub ReleasesAI score36

    Codex 0.162.0 adds managed worktree tools and clickable URLs in the TUI

    OpenAI's Codex 0.162.0 release adds tools for creating and listing managed Git worktrees from trusted local projects when the worktrees feature is enabled. The update also lets users pin tasks in the agent Command Center, copy transcript blocks with /copy, and make URLs clickable in approval headers, questions, and warnings, along with several Linux and Windows sandbox fixes.

  14. Testing CatalogAI score36

    Gemini Agent for Business may add Claude Opus 5 and Sonnet 5.5

    Google's recently announced Gemini Agent for Gemini Business is reportedly set to offer Gemini Argon 4, Gemini Flash 3.8, Claude Opus 5, and Claude Sonnet 5.5. If accurate, it would mark the first time Claude models appear on Google's platform alongside Google's own models, which the post frames as a way for Google to compete for enterprise customers.

  15. Tessl BlogAI score29

    One Brain Means Owning Your Organizational Memory

    Leapfrog, a small team doing high-volume AI visual and production work for fashion and brand clients, is building a "one brain" system that makes company knowledge and client context searchable through natural-language agents. The starter stack described is OpenClaw in a sandbox, a GitHub repository, Obsidian on the local machine, and Telegram as the access point. The system's research structure had roughly 1,200 files at the time of the talk.

  16. Tessl BlogAI score42

    Agent Skills Should Be Treated as Supply Chain Components

    Tessl's talk at AI Native DevCon London argues that agent skills, which can be markdown files with instructions and bundled material, act as supply chain components that can shape agent behavior. The author says reading SKILL.md once is insufficient because risks can sit in supporting files, updates, and workspace trust settings. He identifies the danger as the combination of private context, untrusted content, and external communication, and cites research scanning roughly 4,000 public skills for issues including malware-like behavior.

  17. Elvis SaraviaAI score42

    Don't sleep on domain-specific harnesses. Coding agents are great because of their harnesses, but they aren't built for creative work. Creative work needs its own harness. Voyager looks great. It's an open harness for video, graphics, and games. The agent works with your files and drives apps like Blender, DaVinci Resolve, and Unity right on your desktop. Bring Opus, Astra, or DeepSeek. Excited to try this one.

    Don't sleep on domain-specific harnesses. Coding agents are great because of their harnesses, but they aren't built for creative work. Creative work needs its own harness. Voyager looks great. It's an open harness for video, graphics, and games. The agent works with your files and drives apps like Blender, DaVinci Resolve, and Unity right on your desktop. Bring Opus, Astra, or DeepSeek. Excited to try this one.

  18. Tessl BlogAI score52

    Cisco engineer argues agent skills need a context pipeline with evals

    John Groetzinger, writing in a personal capacity rather than for Cisco, argues that enterprise skills need packaging, evaluation, syncing, and distribution rather than scattered markdown files. He describes using skills to make cheaper models viable, converting curated TAC knowledge-base articles into maintained skills, and rolling out an eval framework across teams. He also describes syncing a repository README to Confluence with a deterministic script.

  19. LiveKitAI score22

    Know exactly what your agent can do, catch gaps before customers do, and test any model against your own scenarios before you switch. Your agent has a QA team now. Try LiveKit Simulations free through October: https://livekit.com/products/agent-simulations

    Know exactly what your agent can do, catch gaps before customers do, and test any model against your own scenarios before you switch. Your agent has a QA team now. Try LiveKit Simulations free through October: https://livekit.com/products/agent-simulations

  20. Artificial AnalysisAI score34

    Harvey LAB-AA uses @harvey's LAB dataset and was built in collaboration with Harvey. Explore the full results: https://artificialanalysis.ai/evaluations/harvey-lab-aa See Harvey's commentary on the evaluation and human expert preferences: https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark https://www.harvey.ai/blog/augmenting-human-preference-in-complex-domains

    Harvey LAB-AA uses @harvey's LAB dataset and was built in collaboration with Harvey. Explore the full results: https://artificialanalysis.ai/evaluations/harvey-lab-aa See Harvey's commentary on the evaluation and human expert preferences: https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark https://www.harvey.ai/blog/augmenting-human-preference-in-complex-domains

  21. AWS Machine Learning BlogAI score46

    AWS Pays Per Inference for AI Agents with BlockRun and Incarna

    Amazon Bedrock AgentCore payments lets AI agents pay for model inference one request at a time, using x402 with USDC on the Base network. Incarna used the service to connect its agents to BlockRun, a pay-as-you-go router serving more than 90 models from more than 15 providers. Spending limits are enforced at the infrastructure layer, outside the model.

  22. Testing CatalogAI score49

    Voyager has launched a desktop app that lets AI agents work inside creative tools like After Effects, DaVinci Resolve, Blender, and Unity on a Mac. It reads project files, operates creative apps, and outputs editable results. > Video edits, motion graphics, and color grading in the apps creators already use. > 3D scenes in Blender and game prototypes in Unity. > Built-in skills, custom skills, and a memory that learns how each user works.

    Voyager has launched a desktop app that lets AI agents work inside creative tools like After Effects, DaVinci Resolve, Blender, and Unity on a Mac. It reads project files, operates creative apps, and outputs editable results. > Video edits, motion graphics, and color grading in the apps creators already use. > 3D scenes in Blender and game prototypes in Unity. > Built-in skills, custom skills, and a memory that learns how each user works.

  23. Lauren TanAI score29

    if you use Grok Bot on Omarchy, or are building plugins for it, please let me know if you have any feedback or feature requests! Would be cool to see what interesting integrations we could support https://plugins.omarchy.org/?q=grok+bot#catalog

    if you use Grok Bot on Omarchy, or are building plugins for it, please let me know if you have any feedback or feature requests! Would be cool to see what interesting integrations we could support https://plugins.omarchy.org/?q=grok+bot#catalog

  24. TechCrunch · AIAI score72

    Google launches unified Gemini agent for businesses, with consumer rollout later

    Google announced a unified Gemini agent that can plan and complete tasks on a user's behalf from a single interface, starting with businesses. The agent can connect to systems such as Google Workspace, Microsoft 365, Slack, Jira, and any MCP server, and it writes audit trails attributed to itself. Google said consumers will get the agent later.

  25. PyTorch BlogAI score62

    NVIDIA Dynamo adds session-level IDs to route and cache agentic inference

    NVIDIA Dynamo uses a unified session-level identifier to make its inference stack aware of agent sessions, subagents, and their KV cache across turns and tool calls. On SWE-bench, two TP4 MiniMax-M2 replicas on one 8xH100 node gained roughly 12-16% throughput from program-aware scheduling over KV-aware routing alone. The post also describes experimental shared-pool indexing and a proposed KvHint interface for session-aware cache policies in vLLM and SGLang.

    AIWhy it matters: The post explains how session identifiers let an inference stack track agent working sets, with measured throughput gains on SWE-bench and agentic RL rollouts.

  26. Google Cloud TechAI score38

    2️⃣ Gemini connects to the apps you already run like Workspace, Microsoft 365, Slack, Jira, Confluence, Git, Salesforce, ServiceNow, and any MCP server. Every tool call passes through our Agent Gateway to enforce policy and block data leaks.

    2️⃣ Gemini connects to the apps you already run like Workspace, Microsoft 365, Slack, Jira, Confluence, Git, Salesforce, ServiceNow, and any MCP server. Every tool call passes through our Agent Gateway to enforce policy and block data leaks.

  27. SantiagoAI score42

    This is a harness for creative work. You can use it to make videos, graphics, and even games. It works with Blender, DaVinci Resolve, After Effects, Ableton, and Unity. It works like Codex or Claude Code: it connects models with creative applications to build your project.

    This is a harness for creative work. You can use it to make videos, graphics, and even games. It works with Blender, DaVinci Resolve, After Effects, Ableton, and Unity. It works like Codex or Claude Code: it connects models with creative applications to build your project.

  28. Google Cloud TechAI score16

    1️⃣ Turn your best employee’s workflow into the team default. Package proven procedures once as Skills and share them in one central registry for your whole team to access, raising accuracy and cutting token spend.

    1️⃣ Turn your best employee’s workflow into the team default. Package proven procedures once as Skills and share them in one central registry for your whole team to access, raising accuracy and cutting token spend.