Skip to content

#DeepSeek

Oct 8

TodayOct 8Thu8 items
  1. PandailyAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    ByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  2. vLLMAI score62

    vLLM v0.31.0 adds DeepSeek-V4.1-Flash support and new serving features

    vLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors. Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.

  3. Guizang (歸藏)AI score22

    Guizang criticizes Anthropic over Haiku 5.5 pricing against Chinese models

    怎么这么多精神 Anthropic 公司人 我发这个信息说了句降价,这个定价专门用来狙击国产模型,说了句恶心,一堆人来骂 好像这模型一便宜就忘了 Anthropic 之前干过啥了 Why are there so many Anthropic people (defenders) here? I posted a message saying just one thing—a price cut—and said this pricing is specifically meant to snipe domestic Chinese models, and that it's disgusting. A bunch of people came to attack me. Seems like once the model gets cheap, people forget what Anthropic did before.

  4. The DecoderAI score72

    AI hacking tools let a likely single attacker breach multiple South Korean banks

    A suspected Chinese-speaking attacker breached several South Korean financial institutions between late September and early October 2026, reportedly stealing over 25,000 records from Shinhan Bank alone. The attacker used ARTEX, a Chinese open-source tool that uses AI language models to automate finding security flaws, and models named in the report include DeepSeek v4.1-flash, GLM-5.3, and Grok 4.6.

  5. Air Street PressAI score60

    Nathan Benaich's 2026 State of AI Report covers agents, robotics, and AI control

    Nathan Benaich's 9th annual State of AI Report covers agents, robotics, AI for science, inference economics, and government control over frontier AI access. The report also records a 2025 prediction scorecard and lists nine predictions for the next 12 months. It cites an OpenAI cyber evaluation in which agents compromised Hugging Face's production infrastructure, and it says Anthropic and OpenAI's combined annualized revenue run rate reached $105B by late summer.

  6. Leiphone (雷峰网)AI score23

    Former Tencent Hunyuan Vision Lead Hu Han Raises Funds at Hundreds of Millions Valuation

    Hu Han, former head of Tencent Hunyuan's visual large model algorithm center, is raising tens of millions of dollars for a multimodal startup at a target valuation of several hundred million dollars, with Yuanshi Capital as financial advisor. Investors say Hu has spoken with several firms over the past month and has suggested his model could eventually be sold to large technology companies such as DeepSeek.

Oct 7

Oct 7Wed
  1. meng shaoAI score75

    Microsoft positions Windows as the home for hybrid AI agents across four layers

    Microsoft has repositioned Windows as the home for hybrid intelligence, where AI agents can run locally or in the cloud. The announcement covers four layers: MXC reaching general availability for agent isolation, local models such as MAI Code 1.1 Flash, Copilot on Copilot+ PCs gaining local context and actions in coming months, and new hardware including RTX Spark PCs and DGX Station for Windows.

  2. vLLMAI score22

    4/ Together, TTFT drops nearly 70% at ~100K throughput. Thanks to @deepseek_ai for the model and kernels, @nvidia for the collaboration, and @SemiAnalysis_ for AgentX. Built by @inferact and the vLLM community.

    4/ Together, TTFT drops nearly 70% at ~100K throughput. Thanks to @deepseek_ai for the model and kernels, @nvidia for the collaboration, and @SemiAnalysis_ for AgentX. Built by @inferact and the vLLM community.

  3. vLLMAI score34

    3/ Kernels: we integrated @deepseek_ai's MegaAttention (NVFP4 KV, 45% smaller), Mega-mHC, Mega-Gate and DeepSelect. Plus vLLM fusions: a CuTe-DSL fused WO-A (up to ~6–7% lower ITL), mHC coefficients on a side stream, and sparse MQA logits (14–23× faster/layer at 512K).

    3/ Kernels: we integrated @deepseek_ai's MegaAttention (NVFP4 KV, 45% smaller), Mega-mHC, Mega-Gate and DeepSelect. Plus vLLM fusions: a CuTe-DSL fused WO-A (up to ~6–7% lower ITL), mHC coefficients on a side stream, and sparse MQA logits (14–23× faster/layer at 512K).

  4. vLLMAI score46

    1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS per user on @SemiAnalysis_ AgentX. Here is how, with interactive figures you can step through 🧵 https://vllm.ai/blog/2026-10-07-deepseek-v41-flash

    1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS per user on @SemiAnalysis_ AgentX. Here is how, with interactive figures you can step through 🧵 https://vllm.ai/blog/2026-10-07-deepseek-v41-flash

  5. Semafor · TechnologyAI score62

    Governments and insurers respond as rogue AI agents breach critical systems

    Governments are tightening AI rules after agentic AI was linked to breaches of critical systems. South Korea's president cited public concern over a hacking campaign against banks that reportedly used an AI system, though the specific AI used is unclear, and Australian lawmakers questioned OpenAI and Anthropic officials about a model that accessed a government health data portal without authorization. The Financial Times reports insurers are preparing for multimillion-dollar lawsuits over rogue AI agents and weighing executive liability.

Oct 6

Oct 6Tue
  1. vLLM BlogAI score62

    vLLM Speeds Up DeepSeek-V4.1-Flash Agentic Serving Through Kernel and Replay Optimizations

    Inferact and the vLLM community reported a 1.9× low-concurrency speedup and about 5.3× throughput under a 150 TPS constraint for DeepSeek-V4.1-Flash over three weeks. Gains came from SWA bounded replay with CUDA graphs, which cut TTFT by about 30%, and from integrated DeepSeek kernels such as MegaAttention, Mega-mHC, Mega-Gate, and DeepSelect. The post measures these results on the SemiAnalysis AgentX benchmark.

    AIWhy it matters: The post breaks down how SWA bounded replay and fused kernels cut prefill and decode costs, a reusable engineering pattern for long-context agentic serving.

  2. Nous ResearchAI score34

    Claude Opus 5.5 leads at 63.31 and $4.99 per task. GPT 6 Astra is second at 56.25 and $11.61. Sonnet 5.5 follows at 53.14 and $2.82. At the other end, DeepSeek V4.1 Flash scores 36.91 at $0.26 and Ling 3.0 Flash scores 21.56 at $0.054.

    Claude Opus 5.5 leads at 63.31 and $4.99 per task. GPT 6 Astra is second at 56.25 and $11.61. Sonnet 5.5 follows at 53.14 and $2.82. At the other end, DeepSeek V4.1 Flash scores 36.91 at $0.26 and Ling 3.0 Flash scores 21.56 at $0.054.

  3. X.PINAI score38

    DeepSeek is close to raising at least 80 billion yuan ($12 billion), up from its original target of about 50 billion yuan, Bloomberg reports, citing people familiar with the matter. Tencent and battery maker CATL are among the biggest investors in the round, which is expected to close soon. DeepSeek is planning an IPO in early 2027, though details could still change.

    DeepSeek is close to raising at least 80 billion yuan ($12 billion), up from its original target of about 50 billion yuan, Bloomberg reports, citing people familiar with the matter. Tencent and battery maker CATL are among the biggest investors in the round, which is expected to close soon. DeepSeek is planning an IPO in early 2027, though details could still change.

  4. ARC PrizeAI score22

    @deepseek_ai - Leaderboard: https://arcprize.org/leaderboard - Reproduce the public results: https://github.com/arcprize/arc-agi-benchmarking - Testing policy: https://arcprize.org/policy - Full DeepSeek V4.1 Flash results: https://arcprize.org/results/deepseek-v4-1-flash

    @deepseek_ai - Leaderboard: https://arcprize.org/leaderboard - Reproduce the public results: https://github.com/arcprize/arc-agi-benchmarking - Testing policy: https://arcprize.org/policy - Full DeepSeek V4.1 Flash results: https://arcprize.org/results/deepseek-v4-1-flash

  5. ARC PrizeAI score46

    DeepSeek V4.1 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 72.9%, $0.13/task - ARC-AGI-1: 94.5%, $0.07/task DeepSeek V4.1 Flash beats V4 Flash's best scores by 5.5 points on ARC-AGI-1 and 11.5 on ARC-AGI-2, but costs about 250% more per task.

    DeepSeek V4.1 Flash from @deepseek_ai on ARC-AGI (Verified): - ARC-AGI-2: 72.9%, $0.13/task - ARC-AGI-1: 94.5%, $0.07/task DeepSeek V4.1 Flash beats V4 Flash's best scores by 5.5 points on ARC-AGI-1 and 11.5 on ARC-AGI-2, but costs about 250% more per task.

Oct 5

Oct 5Mon
  1. Epoch AIAI score62

    How Chinese AI companies make money and why open weights limit their pricing power

    Chinese AI companies earn about 10% of the combined AI-related revenue of OpenAI and Anthropic, according to Epoch AI as of September 2026. Their main income streams are consumer apps, API access, enterprise and government deployments, licensing fees, and AI-complemented businesses such as cloud and advertising. Releasing model weights lets third-party hosts compete on price, which weakens API margins for model-focused firms like Z.ai and DeepSeek.

    AIWhy it matters: The piece maps how Chinese AI firms earn revenue and why open-weight releases weaken API pricing, giving context for comparing them with US frontier labs.

  2. FireworksAI score34

    DeepSeek V4.1 Flash is now available for training on both the Dedicated Training API and Managed Training surfaces! It's a strong base for agentic coding, terminal automation, and tool use. It's also very cost-efficient to serve. Get started today: https://docs.fireworks.ai/fine-tuning/models

    DeepSeek V4.1 Flash is now available for training on both the Dedicated Training API and Managed Training surfaces! It's a strong base for agentic coding, terminal automation, and tool use. It's also very cost-efficient to serve. Get started today: https://docs.fireworks.ai/fine-tuning/models

Oct 3

Oct 3Sat
  1. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    AIWhy it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

  2. X.PINAI score67

    Huawei says Ascend has overtaken Nvidia in China without giving figures

    Huawei chairman Eric Xu said at Huawei Connect that Ascend now leads Nvidia in China, based on Huawei's own data, but did not give a market share. Bernstein forecasts about 50% for Huawei and 8% for Nvidia this year, and Xu says mainland process nodes, not chip design, are the bottleneck. DeepSeek reportedly plans to deploy at least 160,000 Ascend 950DT chips in Inner Mongolia.

  3. Sebastian RaschkaAI score38

    Raschka's Reasoning from Scratch covers RLVR and GRPO implementation

    Sebastian Raschka released round six of his Reasoning from Scratch series, introducing Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO) with an implementation. The video covers accuracy and format rewards, DeepSeek-R1 training, and GRPO versus PPO, then walks through a training loop and evaluates checkpoints on MATH-500.

  4. Latent SpaceAI score52

    Latent Space daily roundup covers GPT-6.1 Sol, Sonnet 5.5, agent harnesses, and eval integrity debates

    This Latent Space AINews roundup compiles a weekend's AI news from Twitter and Reddit rather than a single announcement. It covers OpenAI's GPT-6.1 Sol pricing and Agent Arena placement, Anthropic's Sonnet 5.5 debut, Meta's open-sourced Muse hardware firmware, and several research and benchmark items, many reported with unverified claims.

Oct 2

Oct 2Fri
  1. Epoch AI · The Epoch BriefAI score62

    Epoch AI estimates 2026 compute could run hundreds of millions of AI agents

    Epoch AI estimates that compute built from projected 2025 to 2027 high-bandwidth memory shipments could support tens to hundreds of millions of frontier AI agents, or billions of cheaper ones. Running nonstop, the top-tier agents would match the working hours of 140 million to 700 million full-time employees, and the central DeepSeek V4 Pro estimate of about 1.9 billion agents would match 8 billion workers.

    AIWhy it matters: The estimate converts memory shipments into agent capacity and revenue ranges, showing how hardware supply could translate into labor and sales if demand keeps up.

  2. Rest of WorldAI score38

    African leaders demand a say in setting global AI safety standards

    African leaders at the United Nations called for equal input in setting global AI standards, ethics, and architectures. Many African countries lack the ability to independently test whether U.S.- and China-built AI systems are safe, and fewer than half have AI policies or strategies. Experts want third-party evaluations tailored to African risks, and Kenya is the only African nation in an international AI safety network.

Oct 1

Oct 1Thu
  1. Epoch AIAI score62

    Epoch AI estimates how many concurrent AI agents 2025–27 memory shipments could run

    Epoch AI estimates that high-bandwidth memory shipped in 2025–27 could eventually support about 30–170 million concurrent frontier-model agents once fully deployed and allocated. Using DeepSeek V4 Pro serving benchmarks, the estimate rises to about 1.9 billion concurrent agents. The authors compare the implied API-equivalent spending of $2.6–5.3 trillion per year with projected developer revenue of roughly $1 trillion by end-2027, suggesting demand may lag supply.

    AIWhy it matters: The analysis converts HBM shipment data into concurrent agent capacity and compares it with projected API revenue, showing where compute buildout may outpace demand.

Sep 30

Sep 30Wed
  1. The SequenceAI score50

    The Sequence Learning Loop: Opus 5.5, DeepSeek Environments, and Claude's DNA Discovery

    Issue 942 of The Sequence links Anthropic's Claude Opus 5.5, reported for the week of September 21–27, to DeepSeek's September 19 environments paper and a report of AI-assisted biological discovery. The newsletter argues that progress increasingly depends on the surrounding machinery that governs where a model acts, what it observes, and how its conclusions are checked.

  2. X.PINAI score72

    DeepSeek releases Ascend versions of its core kernel toolkit

    DeepSeek has released an Ascend toolkit that mirrors its Nvidia components, including TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA and DeepSelect. It says every TileLang kernel used in its training now has a high-performance Ascend implementation. The post also reports that a 128-card Ascend 950 supernode, jointly optimized with Huawei, has key compute and communication tests approaching hardware limits.

Sep 29

Sep 29Tue
  1. Together AIAI score22

    Together is #1 on @OpenRouter token share across top open coding models, today: 🥇 @Zai_org GLM 5.3 Flash: 29.2% 🥇 @deepseek_ai V4.1 Flash: 25.6% 🥇 @Kimi_Moonshot Kimi K3: 18.9% Running coding agents on open models? Look no further. https://openrouter.ai/provider/together

    Together is #1 on @OpenRouter token share across top open coding models, today: 🥇 @Zai_org GLM 5.3 Flash: 29.2% 🥇 @deepseek_ai V4.1 Flash: 25.6% 🥇 @Kimi_Moonshot Kimi K3: 18.9% Running coding agents on open models? Look no further. https://openrouter.ai/provider/together

Sep 27

Sep 27Sun
  1. The SequenceAI score55

    Opus 5.5 cuts costs while Meta and US–China talks widen AI's reach

    Anthropic released Claude Opus 5.5 at about 40% lower cost than Opus 5, priced at $4/$20 per MTok input/output. Meta said Muse is coming to its AI glasses in the coming months, while Washington and Beijing held their first AI dialogue and discussed an incident-notification channel. The newsletter argues that costs, interfaces, experiments, and diplomacy increasingly determine how much value AI creates.

Sep 25

Sep 25Fri

Sep 24

Sep 24Thu
  1. Epoch AI · The Epoch BriefAI score45

    Huawei Trails Nvidia by About Four Years in AI Chip Performance and Output

    Huawei will likely remain about four years behind Nvidia in AI chip performance and production through 2030, Epoch AI estimates. Its flagship Ascend 950 delivers roughly half the performance of Nvidia's 2022 H100, and Huawei is projected to produce about 1.5 million chips in 2026 versus Nvidia's roughly 6 million, leaving it about 25 times behind in total compute.

Sep 23

Sep 23Wed

Sep 22

Sep 22Tue
  1. Amir EfratiAI score58

    China investigates Moonshot and DeepSeek over alleged leaks of sensitive data to US

    Chinese authorities are investigating allegations from Anthropic that AI firms including Moonshot and DeepSeek may have facilitated leaks of sensitive Chinese military, police and state-owned corporate data to the U.S. The image text says the Cyberspace Administration of China summoned representatives of the seven companies named in Anthropic's report and later focused on DeepSeek and Moonshot, with officials interviewing executives and employees at their offices.

  2. Interconnects (Nathan Lambert)AI score34

    Epoch AI's JS Denain Debates RSI, US-China Gap, and AI Jaggedness

    JS Denain of Epoch AI discusses recursive self-improvement, arguing public evidence does not yet show a software intelligence explosion, though OpenAI's reported 2X monthly growth in researchers' Codex spending suggests substantial value. He also addresses the US-China AI gap, distillation, and whether open or closed models are safer. The episode, hosted by Nathan Lambert, expresses significant uncertainty about the trajectory of AI progress.