Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 9

TodayOct 9Fri364 items
  1. Jerry LiuAI score26

    Jerry Liu says evals now replace hand-built agent workflows

    AIJerry Liu argues that most tasks can now be solved by defining an eval and hillclimbing on it, rather than hand-coding a deterministic or agentic workflow. He says data provider companies are building evals across economic activity so frontier models can handle more work, leaving developers to define goals and success measures. He expects agent interfaces to compress most tasks into goals and eval instructions, while the most complex processes will still need explicit workflow builders.

  2. X.PINAI score60

    Suspected Guangdong attacker reportedly used Claude Code, ARTEX, GLM and DeepSeek

    AIA suspected 26-year-old in Guangdong reportedly used Claude Code, ARTEX, GLM and DeepSeek in attacks. An AI-generated résumé named South China University of Technology, but the identity is unverified and the listed phone number's owner denied involvement. The suspect reportedly sought buyers on Telegram, but no sale was reported, and ARTEX creator Autumn condemned the misuse and said he would stop releasing the tool as open source.

    Image from @thexpin's post
  3. LeiphoneAI score8

    Carbon-Silicon Dao Code Seventh Layer Sets Self-Audit Baseline and Falsification Terms

    AIThe seventh and final layer of the "Carbon-Silicon Dao Code" cross-domain migration governance framework sets a self-audit baseline, opens falsification terms, and defines the framework's applicability boundary. The article says it validates each layer's input-output consistency backward from layer seven to layer one, and it allows anyone who constructs a reproducible, traceable counterexample targeting NT1–NT4 to submit it for baseline review. It also archives the framework's documents and versions with hashes across Toutiao, Douyin, and GitHub.

  4. vLLMAI score42

    vLLM Semantic Router team releases Decision 2.0 multi-question classification models

    AIThe vLLM Semantic Router team has released Decision 2.0, which answers multiple questions about one input in a single forward pass and outputs per-option probabilities. The post presents this as useful for routing and classification. A quoted post from Xunzhuo Liu says Decision 2.0 includes six open decision models ranging from 0.6B to 27B parameters, each topping same-size open models on the Jev Decision Index 0.3.

  5. OpenAI NewsroomAI score45

    OpenAI fires three researchers over sensitive information breach, denies retaliation

    AIOpenAI says it parted ways with researchers Jasmine, Mikita, and Tomek after an internal investigation found they violated policies on handling sensitive information. The company says the decisions were not about raising safety concerns, which it says it encourages, and that it has not terminated any employee for raising concerns. OpenAI also says it is finalizing contracts with third-party safety assessors and will announce details in the coming weeks.

  6. MarkTechPostAI score44

    Google Research RRSI Guide: Mastering Self-Improving AI Agents

    AIMarkTechPost publishes a hands-on tutorial implementing RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent revise its own harness around a frozen model. The full loop drafts edits with Claude Opus on Vertex AI and scores them in Docker benchmarks, but the edit-selection rules are plain Python that the tutorial runs in a simulated environment with a calibrated noise band.

  7. indigoAI score28

    AI can build features, but defining requirements and design remains the gap

    AICurrent AI can quickly implement or replicate features, but clearly defining requirements and describing design is still missing, and the author expects this gap to persist. As requirements grow more abstract, humans may only specify goals and check results while agents handle implementation, leaving the software's logic layer as model-generated tokens.

  8. X.PINAI score46

    Manus parent Butterfly Effect raises over $500M led by Boyu Capital

    AIManus parent Butterfly Effect announced funding of over $500M led by Boyu Capital and IDG, with Tencent, HSG and ZhenFund returning, its first disclosed raise since resuming independent operations. The reported $4B post-money valuation was not confirmed. Manus also launched version 2.0 and the personal agent Cue on September 29, and is building a China-focused product team and partnerships with domestic model developers.

    Image from @thexpin's post
  9. X.PINAI score46

    Seed preprint finds DeepSeek V4 long-context retrieval varies by position

    AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

    Image from @thexpin's post
  10. LeiphoneAI score42

    Doubao Work adds Canvas feature and Doubao 2.1 Lite model

    AIDoubao Work has added a Canvas feature for complex creative tasks, placing materials, design plans and outputs on one infinite canvas where users can keep editing text, colors and layout after images are generated. The update also integrates the lightweight Doubao 2.1 Lite model, aimed at everyday Q&A, document writing, spreadsheets and PPT creation, with optimized response speed and usage consumption.

  11. GeekParkAI score62

    Krinwave raises 400 million yuan to bring brain imaging ultrasound to AI

    AIGeekPark reports that Krinwave, a Shenzhen company, has completed a new 400 million yuan round with XVC and Sequoia China as new investors. The company says its low-frequency ultrasound captures brain structures through the skull, which conventional high-frequency devices cannot, and it aims to supply hardware and its kOS software platform to brain-computer interface and NeuroAI groups.

  12. GeekParkAI score47

    Ten Days With Today AI, a Domestic Personal AI Assistant That Connects Chinese Apps

    AIToday AI, built by Teambition founder Qi Junyuan, launched its China version on September 24 and connects to Feishu, DingTalk, Tencent Docs and email. The author found it proactively sends morning and evening briefings and handles a single chat window across tasks, but struggled with misjudging task weight and gave confident yet wrong mod-installation instructions that cost an hour of testing.

  13. The Guardian · AIAI score29

    Gordon Brown urges Britain to become an innovation nation, citing £5,000 per household gain

    AIGordon Brown argues that raising UK innovation intensity to Swedish or Japanese levels could leave every household £5,000 better off and add £150bn a year to national income. He says the UK leads in universities and research papers but struggles to scale startups, with three-quarters of venture capital coming from overseas.

  14. ModelScopeAI score63

    Google releases EmbeddingGemma 2, a lightweight multimodal embedding model for on-device search

    AIGoogle released EmbeddingGemma 2, a 740M-parameter multimodal embedding model under Apache 2.0 for private, on-device search and retrieval. It maps text, code, images, video, and audio into one shared space and reports a 9.92-point gain over EmbeddingGemma 1 on MTEB Code. The post lists about 191MB active RAM for quantized text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro.

    Video from @ModelScope2022's post
  15. Joshua AchiamAI score22

    Achiam warns Ukraine war autonomy and general AI demand urgent safety focus

    AIJoshua Achiam argues that rapid battlefield evolution toward fully autonomous warfare in the Ukraine-Russia war, coinciding with the arrival of fully general AI, is among the most important subjects for safety and security advocates. He says Silicon Valley's memetic bubble is preventing serious engagement with this issue.

  16. Neroitech Inventions (NITI)AI score40

    Sui Agent Pass proposes bounded, enforceable limits on AI agent spending

    AIEvan Cheng, CEO of Mysten Labs, presented the Sui Agent Pass at Sui Basecamp in Singapore as a way to bound AI agent authority over money. Users would define allowed actions, assets, recipients, and permission duration, with the system itself enforcing those limits so that losses stop at a predefined cap even if an agent is compromised. The article argues that the real challenge is making such limits impossible to bypass when failures occur.

  17. QbitAIAI score64

    Tsinghua-linked VPP2 world action model tops RoboDojo simulation leaderboard

    AIStar Motion Era's VPP2, a world action model, ranked first on the RoboDojo simulation leaderboard with a 32.26% average success rate and 39.26 average score. The article attributes gains to staged training that separates video prediction from action learning, and reports a 58.5% zero-shot success rate on a real ALOHA dual-arm robot versus 40% for π0.5. The code is open source on GitHub.

  18. QbitAIAI score44

    Sharpa unveils D01 humanoid robot, W02 dexterous hand, and AE01 haptic glove at IROS

    AISharpa launched D01, a fully self-developed humanoid robot with electronic skin covering the whole body and tactile coverage of the upper body, sensing forces from 0.1 to 20N at 100Hz. It also unveiled the W02 dexterous hand, which has 21 active degrees of freedom, about 30% smaller than the W01, and the AE01 exoskeleton data glove with 22 encoders for teleoperation and data collection.

  19. Nace AIAI score22

    NDI 1.0 document processing model launches for coding agents at 90% lower cost

    AINACE introduces NDI 1.0, a document processing model for coding agents that it says is 90% cheaper and ranks first on the Parse Index. The company says it offers native MCP, SDK, and CLI integration for Claude Code, Codex, Hermes, OpenClaw, and PI, and supports 50 languages. NACE also states the model was trained on over 15M financial files and offers $25 in free API credits to developers.

    Video from @NaceAI's post
  20. 雷电芽衣AI score24

    ELY Office Suite, an early Rust and GPUI native office suite, is released for macOS

    AIELY Office Suite is an open-source native office suite built with Rust and GPUI, now in an early test build for macOS, with Linux and Windows planned next week. It includes Word, PPT, Excel, PDF, Markdown, and Git client components, using LibreOffice's LOK for the office kernel, MuPDF for PDF, and MarkText for Markdown rendering. The author warns the build still has many bugs and should not be used in production.

    Video from @ZacharyZhang's post