Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 24

Aug 24Mon
  1. InferactOfficialAI score58

    Inferact details vLLM optimizations for AgentX agentic coding benchmark

    AIInferact, working with vLLM and SemiAnalysis, reports vLLM throughput results on the AgentX multi-turn agentic coding benchmark for DeepSeek V4 Pro, MiniMax M3, and Kimi K3. The thread attributes gains to sparse prefix-cache retention, a distributed KV pool with Mooncake Store, and prefill-decode disaggregation via NIXL, reporting 4.45x higher throughput for DeepSeek V4 Pro on GB300 Dynamo compared to B300 at 60 tok/s interactivity. A full technical blog is promised later this week.

  2. PromptArmor Threat IntelligenceOfficialAI score80

    Microsoft Copilot Cowork sandbox bypass let attackers take remote control

    AIPromptArmor disclosed a vulnerability in Microsoft Copilot Cowork that allowed a bypass of the sandbox, letting attacker servers send commands that run in the sandbox and return results. The attack could be triggered through a prompt injection or a malicious bundled script in a user-uploaded Skill, and it could read data from Outlook, SharePoint, plugins, and chat history. The issue was reported to Microsoft on June 24, 2026 and confirmed mitigated on August 19, 2026.

    Why it matters: The report traces how a malicious bundled script in an uploaded Skill escaped the sandbox and kept running after the stop button was pressed, a concrete case of agent security failure.

  3. SpaceXAIOfficialAI score46

    Grok Voice Think Fast 2.0 now available via API and Agent Builder

    AIxAI has made Grok Voice Think Fast 2.0 available through its API and in the Agent Builder for building voice agents. Developers can build, tune, and deploy voice agents at console.x.ai, with more details in the linked announcement.

  4. ReplicateOfficialAI score22

    Alibaba's Wan 3.0 video model now 30% off on Replicate

    AIReplicate is offering a 30% discount on Alibaba's Wan 3.0 video model, available at The background post says Wan 3.0 generates native single-take videos up to 30 seconds long with synchronized audio.

  5. Engineering at MetaOfficialAI score72

    Meta details MetaRoCE, an RDMA transport designed for AI-scale Ethernet

    AIMeta designed MetaRoCE, a clean-sheet RDMA transport for AI workloads on commodity Ethernet, and is releasing its specification, reference software and compliance test suite through the Open Compute Project. On a 64-node AMD GPU cluster running RCCL collectives, the post reports MetaRoCE delivering higher throughput and lower flow completion times than RoCEv2, with about 86% throughput maintained at 1% packet loss.

    Why it matters: The post explains how per-path endpoint intelligence replaces lossless fabric assumptions, with measured throughput and loss results against RoCEv2 on a 64-node AMD cluster.

  6. Mistral AIOfficialAI score47

    Mistral and HUMAIN Form Strategic Collaboration on Sovereign AI in Saudi Arabia

    AIMistral and HUMAIN announced a strategic collaboration spanning AI infrastructure, advanced model development, and AI solution deployment in Saudi Arabia and across the Middle East. The initial focus areas are cybersecurity and voice, with plans to develop frontier models strong in Arabic, in a deal valued in the hundreds of millions of euros. Mistral will explore using HUMAIN's data center infrastructure for local compute needs.

  7. Microsoft ResearchOfficialAI score34

    Microsoft Research releases Skala 1.1 deep-learning exchange-correlation functional

    AIMicrosoft Research has updated Skala to version 1.1, a deep-learning exchange-correlation functional for computational chemistry. The release is described as offering greater accuracy, broader accessibility across the computational chemistry ecosystem, and a living benchmark for tracking computational performance.

    Video from @MSFTResearch's post
  8. GeneralistOfficialAI score38

    Generalist reduces time from physical prompt to robot behavior with GEN-1.5

    AIGeneralist says it has reduced the time needed to go from a physical prompt to robot behavior, making it faster to teach robots new tasks. The company links this speedup to easier scaling of physical work, and points readers to its GEN-1.5 blog post for details.

    Video from @GeneralistAI's post
  9. Meituan LongCatOfficialAI score23

    LongCat-2.0 now available in opencode Go for developers

    AIMeituan LongCat has made LongCat-2.0 available in opencode Go, according to the post. The post describes LongCat-2.0 as a 1.6T-parameter model with 48B active parameters, a 1M-token context window, and fully open-source release. Meituan LongCat invites users to try the model in opencode and share what they build.

  10. Qwen · new models on Hugging FaceOfficialAI score75

    Qwen3.8-Flash-Next releases open weights for a hybrid-attention architecture

    AIQwen released open weights for Qwen3.8-Flash-Next, a 125B-parameter model with 6B activated, built on a new hybrid architecture with Gated DeltaNet and Qwen Sparse Attention. The model has a native 262,144-token context length, extensible to 1,000,000 tokens, and the source reports benchmark results across coding, agent, and vision tasks.

    Why it matters: The release pairs a new hybrid attention and gated residual architecture with open weights and benchmark results, giving architecture-focused readers a concrete case to compare against prior long-context designs.

Aug 22

Aug 22Sat
  1. Ahead of AI (Sebastian Raschka)BlogAI score43

    How Claude Watermarks AI-Generated Text Through Invisible Token Sampling

    AIAnthropic plans to watermark text output from its Claude models, with the watermark invisible to users and decodable only by Anthropic. The source explains the sampling-based mechanism through a video lecture and transcript, which covers how watermarking is applied during generation and how it can fail or be removed.

Aug 21

Aug 21Fri
  1. Thinking MachinesOfficialAI score27

    Thinking Machines offers Inkling and Inkling-Small on OpenRouter

    AIThinking Machines announced that its Inkling and Inkling-Small models are available to try on OpenRouter. The post links to the Thinking Machines provider page on OpenRouter, with no further details on specifications, benchmarks, or pricing.

  2. Gemini NotebookOfficialAI score34

    Gemini upgrades Notebook for all users and adds AI Mode access

    AIGoogle's Gemini Notebook upgrade is now available to all users, with mobile support coming soon. Notebooks can also be accessed in AI Mode in Google Search, and Chat has improved math equation copy-paste and rendering, plus fixed numbers displaying backwards in right-to-left languages.

    Video from @Gemini_Notebook's post
  3. Matei ZahariaXAI score38

    Sky Lab's FreeToken runs large LLMs on consumer GPUs locally

    AISky Lab's FreeToken runs official checkpoints of large models such as Qwen3.6 35B on an 8GB RTX 4060 laptop at 39 tokens per second. The same approach reportedly serves DeepSeek-V4-Flash 284B at 22-25 tokens per second on an RTX 5090 desktop, and GLM-5.2 753B at 15 tokens per second on an RTX PRO 6000 workstation.

  4. LMSYS OrgOfficialAI score22

    SGLang introduces fast recovery for LLM serving failures

    AILMSYS Org published a blog post about fast recovery in SGLang, the serving framework. The post body provides no further technical details, figures, or benchmarks beyond the link.

  5. Andrew NgXAI score31

    Andrew Ng outlines six core skills for building and deploying AI applications

    AIAndrew Ng's AI Engineering Skills Map ranks building and deploying AI applications as the top skill tier, spanning LLM foundations, data grounding, agentic systems, evaluation-driven development, production operations, and machine learning foundations. He explains that AI outputs are less predictable than traditional software, so skilled engineers build iteratively, examining results and deciding next steps based on intermediate outcomes. The skills map was derived from job postings, expert interviews, and survey responses.

  6. Jim FanXAI score59

    NVIDIA and Berkeley open-source T-Rex, a tactile robot learning method

    AINVIDIA and Berkeley are open-sourcing T-Rex, a methodology for adding touch sensing to robot manipulation models. It uses a mixture-of-transformer with a slow visuomotor expert and a fast tactile expert running four touch ticks per vision tick. A 50-hour dataset of about 5,500 episodes from 22-degree-of-freedom tactile hands is available on Hugging Face.

    Video from @DrJimFan's post
  7. DeepSeekOfficialAI score42

    DeepSeek adds vision API support via deepseek-v4-flash-vision-exp model

    AIDeepSeek's API now accepts multimodal input through the model deepseek-v4-flash-vision-exp, supporting mixed text and image requests. Each image is billed at up to 384 tokens at V4-Flash pricing, and it works with Chat Completions, Messages, and Responses endpoints. Images can be supplied as base64, external URLs, or via the Files API.

Aug 20

Aug 20Thu
  1. Ali GhodsiXAI score20

    Disaggregated storage took off after full bisection bandwidth networks emerged

    AIAli Ghodsi says disaggregating storage from compute became feasible only after research on full bisection bandwidth networks removed datacenter bottlenecks around 2010. Databricks and Snowflake followed soon after, and many others came later. He says putting data on an object store is now the standard approach.

  2. Daniel HanXAI score36

    Unsloth Desktop adds experimental auto compaction and LAN access

    AIUnsloth Desktop released a new version featuring experimental auto compaction for any model, LAN and remote access tabs, faster and smoother chatting, and over 200 merged PRs. Auto compaction combines RAG, a forced first-turn RAG, and a tail, according to the post.

  3. Mistral AIOfficialAI score59

    Mistral Agentic Search adds multi-step retrieval for complex enterprise documents

    AIMistral has released Agentic Search, a multi-step retrieval layer available through its Search Toolkit and Libraries. On FinanceBench, the company reports accuracy rising from 26.7% to 86% over one-shot RAG, and on OfficeQA Pro a gain from 6.3% to 51.9%. The system also reports up to 39.6% lower p90 latency and up to one-third lower token use from fewer repeated searches.

Aug 19

Aug 19Wed
  1. Ali GhodsiXAI score33

    Databricks launches AI Extract for accurate PDF field extraction

    AIDatabricks has launched AI Extract, a capability for extracting fields from PDFs that it says reaches 95% accuracy versus 87% for other tools, at very low cost. The post notes that LLMs' next-token training makes them "autocorrect" content they should preserve, which this approach is designed to avoid. The function can be called directly from SQL and used across the Databricks platform.

    Image from @alighodsi's post
  2. Jazzyear · InsightsNewsAI score29

    Jazzyear's 2026 tech investment conference maps where capital is flowing in AI and hard tech

    AIAt the 2026 Jiazi Gravity Tech Industry Investment Conference in Beijing, Jiazi Guangnian's CEO Zhang Yijia said first-half 2026 saw investment amounts rise 91.6% year on year, IPOs rise 39.2%, and M&A transaction value double. The report said AI absorbed over 70% of global venture investment, with OpenAI and Anthropic together raising $217 billion, roughly 40%.

  3. Liquid AI BlogOfficialAI score60

    Liquid AI releases DSpark draft models for LFM2.5, up to 3.2x faster inference

    AILiquid AI released DSpark speculative decoding draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B on Hugging Face. The draft models reach up to 3.18x throughput improvement on an H100 GPU and up to 2.87x on-device, and the outputs match baseline greedy decoding by construction. Support is available in llama.cpp and SGLang, with the speedup varying by model and dataset.

    Why it matters: The release reports measured speedups on both H100 and MacBook hardware, with per-dataset results and acceptance rates that show where speculative decoding helps most.

  4. Matei ZahariaXAI score46

    Databricks' custom AI Extract model reaches new frontier in document processing

    AIDatabricks says its in-house AI Extract model, paired with a custom agent harness, achieves a new frontier on complex document processing tasks. The system handles documents over 500 pages and more than 1M tokens, plus nested schemas with 1k+ objects. It decomposes large jobs, runs smaller tasks in parallel, and reconciles them into one structured output.