Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. The Verge · AINewsAI score30

    SpaceXAI Backs Omarchy Linux Distro With $1.5 Million in Grok Tokens

    AISpaceXAI is joining the Omacom Foundation, which oversees the Omarchy Linux distribution, as a Founding Corporate Patron and donating $1.5 million worth of Grok tokens to the project. According to David Heinemeier Hansson's blog post, the tokens will primarily accelerate development, review code, and patch bugs. The partnership follows earlier controversy over Hansson's anti-immigration posts, which have drawn criticism of Omarchy's corporate contributors, including 1Password and Cloudflare.

  2. Artificial IgnoranceBlogAI score52

    Charlie Guo maps the core primitives that make AI agents work over time

    AIThe author argues that agent systems are converging on shared primitives grouped into doing the work, continuing the work, and delegating the work. These include instructions and skills, tools and connectors, sandboxes, sessions, compaction, schedules, and subagents. He also flags memory, proactivity, and agent identity as emerging areas still lacking settled standards.

  3. Dhravya ShahXAI score42

    MemoryRepo: open-source implementation of Cognition's dreaming agent memory

    AISupermemory introduces MemoryRepo.dev, an open-source implementation of Cognition's dreaming memory system built on Cloudflare Artifacts, Durable Objects, Alchemy, and Effect. The project follows Cognition's Devin memory design, which builds a memory graph across sessions and prunes stale records overnight. Supermemory says it will incorporate learnings from this research into its own product.

    Video from @DhravyaShah's post
  4. ClineOfficialAI score46

    Cline makes Step 5 Preview free, citing strong DeepSWE coding scores

    AICline says Step 5 Preview is now free in its coding tool and scores ahead of Kimi K3 and GLM-5.3 on DeepSWE. The company describes it as one of the strongest open-weights coding models available. StepFun's background announcement describes Step 5 Preview as a 600B total / 27B active MoE model with 1M context and vision, and says open weights arrive on Oct 15.

    Image from @cline's post
  5. MuseOfficialAI score20

    Muse builds a personalized movie feed and books tickets for users

    AIMuse, created by @JJEnglert, is trained on the user's movie and TV preferences and weekly curates a feed of films in theaters and on owned streaming services. Users can upvote or downvote picks to refine the feed, and when they want to see a movie, Muse books the ticket through a wallet link integration.

    Image from @Muse's post
  6. Luke EdwardsXAI score38

    Pocketty brings SSH and herdr agent alerts to iPhone and iPad

    AIPocketty launches as an SSH app for iPhone and iPad, built for herdr, that notifies users when an agent on any host is blocked. Tapping a notification opens the exact pane, and the app supports Tailscale and Bonjour natively, shows diffs for every agent turn, and requires no account or subscription.

    Video from @lukeed05's post
  7. SiliconANGLE · AINewsAI score30

    Liquid AI Builds On-Device Personal AI Around Device-Level Context

    AILiquid AI is building personal AI that runs on devices such as phones, wearables, PCs, and cars, using its Liquid Context layer, which is optimized for Snapdragon processors, to sit between models, agents, and hardware. The company's agent harness uses its own models to decide which user context to retain and how to compress it within fixed compute limits. Liquid AI is also collaborating with Mercedes-Benz Group AG to bring on-device AI to its cars and plans observability and continuous improvement loops for self-improving agents.

  8. The Verge · AINewsAI score62

    Anthropic updates Claude usage policy to ban abusive treatment and expand misuse rules

    AIAnthropic is revising its usage policy for the first time in over a year, adding bans on sustained abusive or cruel behavior toward Claude and on deceptive election and propaganda campaigns. The update also expands weapons restrictions, tightens surveillance bans, and requires a qualified operator able to stop equipment when Claude controls autonomous physical hardware. Terminating conversations remains the primary enforcement mechanism, and the company did not say whether user bans would follow.

  9. Tessl BlogOfficialAI score44

    Continuous AI Brings Agentic Automation to Repository Workflows

    AITessl's blog post argues that repository automation needs Continuous AI, a third pillar alongside CI and CD for scheduled, auditable AI workflows that improve repositories over time. The article describes GitHub Agentic Workflows, which harden agentic workflow specifications into GitHub Actions that can run coding agents such as Claude Code, Copilot CLI, Gemini CLI, or Codex-style agents. It emphasizes read-only agent steps, restricted outputs, and human review of pull requests.

  10. elvisXAI score46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    AIRSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  11. Daniel HanXAI score38

    Unsloth adds OS-level sandboxing for Linux, Mac, and Windows

    AIUnsloth now supports OS-level sandboxing on Linux via bwrap, on Mac via seatbelt, and on Windows via Microsoft's MXC. Per-tool-call latency is under 100ms across all three, and its software-style sandboxing with regex AST checks adds about 3ms. The Windows integration was built in collaboration with Microsoft.

  12. Google ResearchOfficialAI score14

    Google Research demos EnvHarness for co-evolving LLM agents and environments at COLM 2026

    AIGoogle Research is presenting EnvHarness, a flexible framework that enables co-evolution between LLM agents and their training environments, at the #COLM2026 Google booth #107 today at 11:00 AM PT. The post notes that static environments limit agent growth, and EnvHarness is described as a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and adaptability.

  13. AWS Machine Learning BlogOfficialAI score27

    Share SageMaker HyperPod GPU clusters across teams with isolation and fair scheduling

    AIAWS published a reference architecture for running multiple teams on one Amazon SageMaker HyperPod EKS cluster, with each team isolated in its own Kubernetes namespace. The design combines AWS IAM Identity Center for authentication, per-team SageMaker AI domains, HyperPod Task Governance for fair resource allocation, and namespace-level cost allocation for per-team spend visibility.

  14. MarkTechPostNewsAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  15. elvisXAI score22

    Interface ring lets users control AI agents by voice from hand

    AINatura AI's Interface is a ring that lets users press and hold to speak requests to AI agents such as Claude Code, Codex, or Hermes, then release to send them. The post argues that screenless interfaces may define the next phase of agent use, since handing work to agents is currently slowed by pulling out a phone. Early-adopter pricing is $99, with shipping slated for January.

  16. Goodfire ResearchOfficialAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  17. Vercel DevelopersOfficialAI score36

    StepFun's Step 5 Preview model now available on Vercel AI Gateway

    AIVercel says StepFun's flagship Step 5 Preview, built for agentic coding, research, and finance, is now live on AI Gateway. The model offers a 1M-token context window, accepts text and image input, and uses a 600B-parameter mixture-of-experts design with 27B parameters active.

  18. Grok BotOfficialAI score38

    Grok Bot can now help run Shopify stores

    AIGrok Bot can now connect to Shopify to check orders, track inventory, and keep product listings up to date. The feature lets merchants ask it questions about their business after linking their store. Shopify's related announcement says the new connectors also allow merchants to build teams of agents to help run their business.

    Image from @bot's post
  19. Satya NadellaXAI score38

    Satya Nadella outlines Copilot as a headless "infinite SaaS factory" for agents

    AIMicrosoft CEO Satya Nadella says Copilot is being positioned as a new operating system for work, paired with a governed headless business layer that gives agents access to CRM, ERP, and other systems of record. He says Microsoft announced over 30 new Copilot skills across Dynamics 365 Sales, Service, and Customer Insights, plus Microsoft Copilot Managed Runtime for IT-governed code hosting. He describes users building custom software or Dataverse extensions through Copilot Code, though the post is an early vision with few concrete specifications.