Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. elvisXAI score34

    Elvis Saravia urges builders to focus on agent harnesses and environments

    AIElvis Saravia says AI models are already smart, but they need better harnesses and environments, with major cost implications. He recommends reading a report on how Pine Computer can help teams, and says he will test it himself and share more later. The quoted post from Stanley Wei argues that real-world AI tasks remain slow, expensive and unreliable because AI runs on computers built for humans, and announces Pine Computer.

    Image from @omarsar0's post
  2. elvisXAI score40

    Syren Video learns your style to build AI videos from prompts

    AISyren Video, a new agentic video tool, learns preferred graphics, motion, and editing rhythm from a user's library and generates new videos from a prompt. Users refine the results through chat, and the tool is free to try in a browser or through Claude MCP, per the company's announcement. The post's author says the education sector is exploring it.

  3. GuizangXAI score22

    Guizang releases a one-click Grok bot for daily AI news videos

    AIGuizang says he turned his workflow into a Grok bot that users can install with one click. The bot runs on Grok's cloud virtual machine to collect content, write code, and render a daily morning AI news video without using a local computer.

  4. CoW SwapXAI score34

    CoW Protocol launches pay-per-quote API for bots and AI agents

    AICoW Protocol has launched x402.cow.fi, a service offering pay-per-request trading quotes for bots and AI agents with no API key or sign-up required. Each quote costs $0.001, payable in USDC on Base, Ethereum, or BNB Chain, or in $COW on Base. The service is built on x402.

    Image from @CoWSwap's post
  5. QbitAINewsAI score67

    Aether AI shows CRIS-0 robot recovering from disturbances via causal reasoning

    AIAether AI, founded by UCSD assistant professor Biwei Huang, has released official demos of its CRIS-0 causal intelligence system for robots. In tests, the robot recovered from external disturbances in 9 of 10 random trials, typically within about 2 seconds, and stopped within 0.2 seconds when a human hand entered the workspace during a microwave-door task.

  6. ModelScopeOfficialAI score28

    Corvus-Gov-3B: a 3B Chinese government-domain dialogue model

    AIModelScope released Corvus-Gov-3B, a compact model tuned for Chinese policy Q&A, public-service consultation, and internal government or enterprise assistants. It was fine-tuned on one million Chinese government-domain dialogue samples and built on Llama 3.2 3B Instruct using LoRA SFT via LLaMA Factory. The model is released under Apache 2.0.

    Image from @ModelScope2022's post
  7. OpenBMBOfficialAI score28

    MiniCPM5-2B runs at 37 tok/s on iPhone Air

    AIOpenBMB reports that its MiniCPM5-2B model runs at 37 tokens per second on an iPhone Air. The post presents this as evidence that small open multimodal models can run on mobile devices without a cloud GPU, with NobodyWho noting the model is available in its Chat app.

  8. MarkTechPostNewsAI score44

    Underdog Releases Saluki 27B, a 2-Bit Qwen3.8-27B That Beats the Original at Tool Calling

    AIUnderdog has released Saluki 27B under Apache 2.0, a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB, versus 54 GB for the full BF16 model. On Underdog Bench, Saluki scores 88 against 84 for the full model, and it raises parallel tool-call accuracy to 42 from 35. It runs on stock llama.cpp, but math and reasoning drop sharply, with AIME 2025 at 79.2 versus 96.7.

  9. IThome · AINewsAI score55

    Odyssey-3 world model scores 66.1 on Physics-IQ Verified benchmark

    AIOdyssey announced the Odyssey-3 series of foundation world models, with Odyssey-3 Pro scoring 66.1 on the Physics-IQ Verified video-to-video benchmark, the highest recorded on that leaderboard. The series includes a standard version balancing physical accuracy and generation cost, and a Pro version with stronger physics prediction. The preview supports first-person and third-person navigation and lets users move the camera, take actions, or trigger events while the model predicts environmental changes in real time.

  10. PandailyNewsAI score38

    KingKong Technology Open-Sources Jumper Crab Robot Software Stack

    AIKingKong Technology has open-sourced the software stack for Jumper, a six-legged crab-style robot it designed, including its MuJoCo model, simulation scenes, reinforcement learning training and deployment tooling. Jumper has 22 degrees of freedom, measures about 400 by 400 by 200 mm, weighs about 1.8 kg and lists a maximum jump height of 400 mm or more. The mechanical CAD files, bill of materials, PCB designs and electrical schematics are not public, and RKNN inference on the real board has not yet been validated.

  11. vLLMOfficialAI score42

    vLLM Semantic Router team releases Decision 2.0 multi-question classification models

    AIThe vLLM Semantic Router team has released Decision 2.0, which answers multiple questions about one input in a single forward pass and outputs per-option probabilities. The post presents this as useful for routing and classification. A quoted post from Xunzhuo Liu says Decision 2.0 includes six open decision models ranging from 0.6B to 27B parameters, each topping same-size open models on the Jev Decision Index 0.3.

  12. QbitAINewsAI score44

    Sharpa unveils D01 humanoid robot, W02 dexterous hand, and AE01 haptic glove at IROS

    AISharpa launched D01, a fully self-developed humanoid robot with electronic skin covering the whole body and tactile coverage of the upper body, sensing forces from 0.1 to 20N at 100Hz. It also unveiled the W02 dexterous hand, which has 21 active degrees of freedom, about 30% smaller than the W01, and the AE01 exoskeleton data glove with 22 encoders for teleoperation and data collection.

Oct 8

Oct 8Thu
  1. meng shaoXAI score43

    Unsloth integrates Microsoft's mxc sandbox for Windows AI agent isolation

    AIUnsloth has integrated Microsoft's open-source mxc sandboxing system into Windows as an OS-level sandbox for isolating AI agent code execution. Its High mode provides real operating-system isolation that confines tool calls to specified directories, while its Low mode adds language-level checks that block dangerous commands and shell escapes. Both modes also strip secret environment variables and enforce resource limits such as 8GB memory and 600-second CPU time.

    Image from @shao__meng's post
  2. Higgsfield AI 🧩OfficialAI score36

    Higgsfield Katana adds community presets for Claude video editing

    AIHiggsfield has released community presets for Higgsfield Katana, its AI video editing tool available inside Claude. Users can pick a preset for motion graphics, 3D animations, product launches, fashion, car, travel, or aura-farming edits, then add their own characters, products, or clothes to recreate it in Claude. More presets are coming soon.

    Video from @higgsfield's post
  3. OpenClaw🦞OfficialAI score34

    OpenClaw shares recent feature updates and upcoming roadmap plans

    AIOpenClaw says it has added many new features and quality-of-life improvements over the past few months. A video covers new models, multiplayer, interactive dashboards, memory and skills, meetings and voice, and easier Mac setup. It also previews plans for the coming months.

  4. Simon WillisonBlogAI score22

    ttok 0.4 Adds --list-models Command for Counting Tokens with tiktoken

    AISimon Willison has released ttok 0.4, a command-line tool for counting tokens built on OpenAI's open source tiktoken library. The update adds a --list-models command, fixes a Click warning, and updates CI, and the tool can run through uvx, for example with cat file.txt | uvx ttok.

  5. meng shaoXAI score77

    Theo open-sources tsc-rs, a Rust port of the TypeScript 7 compiler

    AITheo, creator of the T3 Stack, open-sourced tsc-rs, a line-by-line Rust port of Microsoft's Go-native TypeScript 7 compiler, type checker, and language server under MIT, pinned to typescript-go commit 673a5f17. The author reports tsc-rs is about 1.61× faster than tsc 7 and about 2.95× faster than bun check on six real-app benchmarks on an Apple M4 Pro. The port passes all 181,711 ported Go tests, and CLI output matches the Go version on 120 open-source repos except for known edge cases such as monorepo rootDir and tsc -b incremental output.

    Why it matters: The post reports a benchmarked, test-verified Rust port of the TypeScript 7 compiler, with pinned upstream and stated edge cases useful for judging its compatibility.

    Image from @shao__meng's post
  6. PandailyNewsAI score57

    Shanghai AI Lab Open-Sources Intern-Decision Small Models for Structured Decisions

    AIShanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.

  7. SiliconANGLE · AINewsAI score47

    Kore.ai launches Autoloop to tune enterprise AI agents after deployment

    AIKore.ai launched Autoloop, an optimization engine that automatically adjusts AI agents built on its Kore.ai Agent Platform to meet business-set goals, including after deployment. The engine scores each proposed change against goals such as task completion, business-rule adherence, accuracy and cost. Autoloop is available now to all customers on the Artemis edition of the Kore.ai Agent Platform.

  8. TechCrunch · AINewsAI score62

    Goodfire launches inside-out monitors to catch rogue AI agents at lower cost

    AIGoodfire has launched monitors that read a model's internal signals during agent work instead of reviewing its written output. The monitors are available to Baseten customers, who can choose risks to watch and set automated responses. In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, while catching 94% of malicious hacking sessions.

  9. TechCrunch · AINewsAI score43

    Natura's $99 Interface smart ring lets users control AI agents with a finger press

    AINatura's Interface is a $99 smart ring that lets users ask AI agents to complete tasks with a finger press and also tracks heart rate, HRV, sleep, and activity. At launch it connects with Meta's Muse, Instinct, Grok Bot, Claude, ChatGPT, and more, with preorders expected next month and shipping slated for December or January. After a three- to six-month free period, Natura plans to charge a $9 monthly subscription.

  10. Jerry LiuXAI score38

    LightOn OCR-3 now on OpenDocRouter, near Gemini 3.8 Flash at lower cost

    AILightOn OCR-3 is now available on OpenDocRouter at $0.28 per 1M input tokens and $1.40 per 1M output tokens, about $3.19 per 1k pages on ParseBench. On ParseBench, the author says it sits on the Pareto frontier for open-weight OCR models, with performance similar to Gemini 3.8 Flash low at roughly 45% lower price. It is described as decent at tables, workable for charts, and quite good at grounding.

    Image from @jerryjliu0's post
  11. SantiagoXAI score34

    Atomic Agent Desktop launches with Linux support on day one

    AIAtomic Agent Desktop is a local AI-first agent available for Mac, Windows, and Linux. It offers a 4x larger context window on local models using TurboQuant, and Atomic Fusion orchestrates cloud and local models to reduce costs.

  12. 🚨 AI News | TestingCatalogXAI score62

    Atomic Agent Desktop, an open-source local AI agent app, is now available

    AIAtomic Agent Desktop is a free open-source app for macOS, Windows, and Linux that runs open models like Qwen and Gemma locally without an account. It connects to a cloud model only when selected, and its Fusion feature lets a cloud model plan a task while up to 8 local agents carry it out. The post's own text adds a setup wizard that checks RAM and suggests suitable models, and import from Claude Code, Codex, Hermes, and OpenClaw.

    Video from @testingcatalog's post
  13. Higgsfield AI 🧩OfficialAI score34

    Higgsfield launches Katana AI video editing tool inside Claude

    AIHiggsfield introduced Katana, its most powerful AI video editing tool, powered by Claude Motion and available now inside Claude via Higgsfield MCP. Users can upload a reference to create editable motion graphics, product launch videos, or aura-farming edits.

    Video from @higgsfield's post
  14. OpenRouterOfficialAI score29

    Mercury Decide now available with zero data retention on OpenRouter

    AIInception's Mercury Decide, the fastest-growing decision model on OpenRouter last week, now has a paid zero data retention (ZDR) endpoint alongside the free one. It is priced at $0.02/M input tokens, half the $0.04 list price, with output and cached input free and a 66K context window.

    Video from @OpenRouter's post
  15. Artificial AnalysisOfficialAI score38

    GPT-6 Sol (Daybreak Blue) tops Artificial Analysis Cyber Index

    AIGPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model and ranks #1 on the Index. Compared with the publicly available GPT-6 Sol, it shows its largest gains on CyberGym-E2E, the benchmark where the most safety refusals are observed.

    Image from @ArtificialAnlys's post
  16. Boris PowerXAI score46

    OpenAI's GPT-6.1-Sol leads new Arena Alignment Index for agents

    AIThe Arena Alignment Index, built from over 90K real-world agent sessions across 27 models, ranks OpenAI's GPT-6.1-Sol first with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7. GPT-6.1-Sol also posted the lowest observed rates across the index's three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. The index's authors report that newer models consistently outperform their predecessors across all four labs, suggesting broad progress in agent safety.

  17. elvisXAI score42

    Voyager: an open harness for creative AI work across video and games

    AIElvis Saravia argues that creative work needs domain-specific agent harnesses rather than coding-oriented ones, and he highlights Voyager as an open harness for video, graphics, and games. According to the quoted post, Voyager lets agents work with local files and drive apps such as Blender, DaVinci Resolve, and Unity, and it is designed to work with models like Opus, Astra, and DeepSeek.

    Video from @omarsar0's post
  18. Tessl BlogOfficialAI score38

    Mozilla.ai's cq Aims to Give Agents a Shared, Reviewable Knowledge Commons

    AIMozilla.ai's cq project proposes a shared knowledge layer where AI agents capture lessons from non-obvious fixes as structured knowledge units that other agents can later query. The default setup is local-first, using a local SQLite database so nothing leaves the machine, with an option to connect to a remote team server that adds review.

  19. LiveKitOfficialAI score22

    LiveKit Simulations lets teams test voice agents before customers do

    AILiveKit is offering free access to its Simulations product through October, letting teams check what their agent can do and find gaps before deployment. The product also lets teams test any model against their own scenarios before switching models.

    Video from @livekit's post