Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri202 items
  1. Arena.aiAI score38

    Mistral Large 4 ranks in Agent Arena top 15 at -6.6% net score

    AIMistral Large 4, a preview model from Mistral AI, ranks #43 overall in Agent Arena with a -6.6% net improvement score across more than 5,000 real-world agentic sessions. That is 11 rankings above its predecessor, Mistral Medium 3.5 (-12.60%), and places it in the top 15 labs, the only European lab there. Open weights are expected at the end of October, and at its current score the model would rank #13 among open models.

    Image from @arena's post
  2. X.PINAI score49

    Apple's HomeHub ships only after LLMs finally improved Siri

    AIAccording to a source familiar with the project, Apple's homeOS was finished years ago, but the HomeHub was held back until large language models made Siri good enough to serve as its voice-driven interface. The hardware team reportedly refused to sign off on a device whose main interface was Siri, given its long-running poor performance. HomeHub is slated for an Oct 13 unveiling alongside a smart-home push with LG.

    Image from @thexpin's post

Oct 8

Oct 8Thu
  1. TechNode · AIAI score42

    XPENG names robotaxi service YOYO and opens invite-only public testing

    AIXPENG has named its robotaxi business XPENG YOYO and launched a ride-hailing mini program that lets invited members of the public test the service. The company says its first production robotaxi, based on the flagship GX model, rolled off the line in May 2026, and it has completed more than 2,000 internal test rides in Guangzhou. XPENG says YOYO uses four in-house Turing AI chips delivering 3,000 TOPS, a second-generation VLA model, and a vision-based approach without high-definition maps or LiDAR.

  2. QwenAI score22

    Free week of Qwen3.8-Max, Qwen3.8-Flash, and Wan3.0 on GMI Cloud

    AIQwen3.8-Max, Qwen3.8-Flash, and Wan3.0 are available free for a week on GMI Cloud, which is extending the offer by seven days and raising rate limits across all three models. GMI Cloud is also running a contest where three winners each receive $200 cash plus $200 in GMI credits for the most creative, most challenging, or most effort-driven projects built with Qwen or Wan.

  3. ClaudeAI score46

    Anthropic pauses Claude Startups Team and API credit offers amid demand

    AIAnthropic has paused the Claude Team plan and $1,000 API credit offers for its Claude Startups program after underestimating demand, with hundreds of thousands of applicants. Claimed offers will remain in accounts, but some approved applicants who had not yet claimed their offers will lose access as applications are re-reviewed. Startup Stack and Applied AI office hours remain available to accepted members.

  4. meng shaoAI score43

    Unsloth integrates Microsoft's mxc sandbox for Windows AI agent isolation

    AIUnsloth has integrated Microsoft's open-source mxc sandboxing system into Windows as an OS-level sandbox for isolating AI agent code execution. Its High mode provides real operating-system isolation that confines tool calls to specified directories, while its Low mode adds language-level checks that block dangerous commands and shell escapes. Both modes also strip secret environment variables and enforce resource limits such as 8GB memory and 600-second CPU time.

    Image from @shao__meng's post
  5. meng shaoAI score41

    Matt Pocock shares a framework for matching AI coding workflows to change size

    AIMatt Pocock recommends matching AI coding agent workflow weight to change size: one-shot small diffs, start medium-to-large changes with /grill-with-docs for requirements clarification, and escalate to /wayfinder for mapping and tickets only when planning becomes complex. He warns against starting with /wayfinder, since a simpler-than-expected solution can leave the generated map and tickets unnecessary.

    Image from @shao__meng's post
  6. Higgsfield AI 🧩AI score36

    Higgsfield Katana adds community presets for Claude video editing

    AIHiggsfield has released community presets for Higgsfield Katana, its AI video editing tool available inside Claude. Users can pick a preset for motion graphics, 3D animations, product launches, fashion, car, travel, or aura-farming edits, then add their own characters, products, or clothes to recreate it in Claude. More presets are coming soon.

    Video from @higgsfield's post
  7. Feng XueAI score41

    ThreatBook acquires CyberStrikeAI, a widely used AI pentesting agent

    AIThreatBook has acquired CyberStrikeAI, an open-source AI pentesting agent, after reports that attackers had used it in real intrusions. The company says it added guardrails to limit misuse but will put new capability enhancements into its commercial edition, and it argues defenders should control such tools. It also corrects an earlier threat intelligence report that linked the developer, a security engineer at Alipay, to government agencies, calling that attribution a false positive.

  8. Anil Chandra Naidu MatchaAI score29

    Open-Voyager releases open-source creative-work agent harness

    AIOpen-Voyager is announced as a free, open-source harness for creative work, described as a Codex or Claude Code equivalent. The post says it integrates with 600+ models and links to a GitHub repository for self-hosting. The background post says the original Voyager is built for video, graphics, and games, and can drive tools such as Blender, Resolve, and After Effects.

    Video from @matchaman11's post
  9. PandailyAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    AIByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  10. PandailyAI score52

    Huawei's KV cache storage faces a missing SSD endurance standard

    AIHuawei's OceanStor M900 and Nvidia's CMX move reusable inference KV cache into a shared storage tier, but no agreed SSD specification exists for it. A storage executive said endurance requirements for one design rose from 3 to 9 drive writes per day and could change again. Industry sources expect convergence to take 6 to 12 months, with another year for development and validation.

  11. PandailyAI score41

    Simplexity Robotics Trains One Robot to Tend Two CNC Lathes With 600 Trajectories

    AISimplexity Robotics says it trained a single robot to load and unload two CNC lathes on its own, using 600 real-robot trajectories and reporting a 100% success rate on the precision CNC insertion task. The work, presented at IROS 2026 on September 29, combines the SimpleWAM world action model, a DRAM memory module and DPE action scoring, with force and torque feedback for insertion recovery. The company did not say how many trials the 100% figure covers, and it does not describe the yield of the whole cell.

  12. PandailyAI score40

    Huawei Opens DevEco Studio Public Beta on HarmonyOS PCs with DevEco Code and CLI

    AIHuawei has opened its DevEco Studio for HarmonyOS PCs to public beta, alongside first public betas of the AI tool DevEco Code and the agent toolkit DevEco CLI. The beta requires HarmonyOS 7.0.0.107 or later, at least 16 GB of memory and 100 GB of storage, and runs on several MateBook models and the MatePad Edge. DevEco Code ships with Zhipu AI's GLM-5.3 and GLM-5.1 models and supports third-party model connections.

  13. PandailyAI score36

    CAIR Unveils CARES 4.0 Multimodal Clinical Agent That Suggests Rather Than Decides

    AIHong Kong's Centre for Artificial Intelligence and Robotics (CAIR), Chinese Academy of Sciences, unveiled CARES 4.0, a multimodal clinical AI agent that carries out tasks rather than only answering questions about images or video. Built on the Harness agent framework and CAIR's own CT, MRI, ultrasound, endoscopy and EEG foundation models, it has been validated at several top-tier hospitals. CAIR says the system gives suggestions with reasoning paths and sources, and that the doctor remains the final decision-maker.

  14. QbitAIAI score58

    AgentGarten lets agents evolve through code-built worlds and neural rendering

    AIMirroS released AgentGarten, which pairs executable code environments with a real-time neural renderer running above 30 fps so agents can act, observe, and learn. In a one-on-one hide-and-seek setup, the hider learned to block passages by round 4 and the seeker learned to climb ramps by round 10, guided by notes the agents wrote after each round. The authors report applying the same loop to four other tasks, including a dog-companion game, a narrow-bridge car passing task, herding, and quarry loading.

  15. Tencent HyAI score47

    Tencent Hunyuan releases ExplorationBench to measure AI scientific exploration

    AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark testing how AI systems explore through verifiable "Alien Worlds" with executable rules that conflict with familiar knowledge. Across 10 frontier systems, feedback mattered most: the best AlienCode run reached 89.0% after four rounds of probing, versus 0.5–11.0% without feedback. Answers are graded by an interpreter or proof checker rather than an LLM judge.

  16. SiliconANGLE · AIAI score38

    CoreWeave Adds RL Rollouts and Forge Platform to Target AI Inference Bottlenecks

    AICoreWeave is layering managed services over its infrastructure to address AI inference bottlenecks, including a preview capability called CoreWeave RL Rollouts that improved model reload latency by 15x versus a baseline configuration in testing. The capability is built on Nvidia's Dynamo framework, and the features are packaged into CoreWeave Forge, a platform that is free to start with paid tiers offering additional capabilities.

  17. Gizmodo · AIAI score42

    Peter Thiel Says Government AI Regulation Is the Antichrist's Work in Nashville Lectures

    AIPeter Thiel, Palantir and Founders Fund co-founder, argued in $100-per-ticket Nashville lectures that government attempts to regulate AI are evil, according to recordings obtained by Politico. He attacked former President Barack Obama and Pope Leo XIV over AI, and criticized effective altruism as a belief system that could form a one-world government to stop AI progress. The article notes Thiel's wealth is heavily invested in AI, including a $418 million AI-focused portfolio at Thiel Macro LLC.

  18. meng shaoAI score65

    Michigan's Applied Agentic Software Engineering course turns AI coding methods into five Skills

    AIThe University of Michigan's EECS 498 course Applied Agentic Software Engineering teaches a coding agent across three phases, from applying and analyzing agents to building one. Its Elephant-Goldfish Model packages a design-first workflow into five Skills, with human handoffs between each step, and the course materials are public on GitHub.

  19. IThome · AIAI score62

    Terence Tao questions OpenAI's 719 AI-generated math proofs

    AIOpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.

  20. QbitAIAI score62

    Google launches Gemini agent for office work, able to call Claude models

    AIGoogle Cloud introduced the Gemini agent, a general office agent that can search, write emails, build slides, analyze data, run code, and coordinate sub-agents. It can take on an enterprise identity with email, calendar, and account, and it selects underlying models automatically, including Anthropic's Claude. The article presents this alongside OpenAI's Dots and Meta's Muse as competing office and personal agents.

  21. LangChain BlogAI score42

    Snyk Assist: How Snyk Turned an Internal Support Agent into a Customer Feature

    AISnyk moved its internal support agent, Snyk Assist, into the core Snyk product in September 2026, giving every paying customer access. Built on LangChain and LangGraph with observability in LangSmith, the agent answers questions in plain language and can open support cases or log feature requests. It runs as a single agent behind Slack, web and API surfaces, with tools attached per user permissions.

  22. The Guardian · AIAI score42

    Anthropic bans sustained abusive or cruel behavior toward Claude

    AIAnthropic has barred users from exhibiting "sustained and needless abusive or cruel behavior" toward its models, according to a policy change first reported by The Verge. The San Francisco-based company says the ban does not apply to common user frustrations, model testing, or "dark creative themes." The change follows an August feature that lets Claude end conversations when a user is persistently harmful, which Anthropic framed as a safeguard for AI welfare.