DeepSeek-V4.1-Flash is free on WorkBuddy until October 9th
AIWorkBuddy is offering free access to DeepSeek-V4.1-Flash until October 9th. The post invites users to apply the model to their tasks while the offer lasts.
Updated
Updated
AIWorkBuddy is offering free access to DeepSeek-V4.1-Flash until October 9th. The post invites users to apply the model to their tasks while the offer lasts.
AIChinese authorities are investigating allegations from Anthropic that AI firms including Moonshot and DeepSeek may have facilitated leaks of sensitive Chinese military, police and state-owned corporate data to the U.S. The image text says the Cyberspace Administration of China summoned representatives of the seven companies named in Anthropic's report and later focused on DeepSeek and Moonshot, with officials interviewing executives and employees at their offices.
AIJS Denain of Epoch AI discusses recursive self-improvement, arguing public evidence does not yet show a software intelligence explosion, though OpenAI's reported 2X monthly growth in researchers' Codex spending suggests substantial value. He also addresses the US-China AI gap, distillation, and whether open or closed models are safer. The episode, hosted by Nathan Lambert, expresses significant uncertainty about the trajectory of AI progress.
AIOllama says some requests to the deepseek-v4.1-flash model were charged at an incorrect rate over the last few hours, causing higher-than-expected usage for certain users. Affected accounts have had their monthly or weekly and session usage reset, and any extra usage consumed due to the error has been restored.
AIFireworks AI reports that an oracle router choosing among 18 models per DeepSWE task reaches 97.6% at $1.88 per task, versus 74.1% at $6.52 for GPT-6 Astra alone. The oracle is hindsight-based, so the authors say a production router must predict the best model before the task starts, which is the harder problem.
AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.
AISemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.
AISenseTime chief scientist Lin Dahua argues that native multimodal AI, which processes language, vision and other information in one shared model, is essential for AI to move beyond coding into industries and the physical world. SenseTime released the open-source SenseNova U1 in April and U1.5 Lite nearly four months later, and reported first-half 2026 revenue of RMB 2.91 billion, up 23.4% year-on-year. Lin's claim that a breakthrough could come within two years is the source's prediction, not a confirmed result.
AIFormer DeepSeek kernel engineer Shengyu Liu argues that AI industrializing software production could reduce programming to a recreational craft and erode students' engineering skills. His central concern is less whether AI can outthink humans than whether access to it stays widespread or gets concentrated in a few corporations. The post, cited by X.PIN, contrasts this with Western warnings about AI escaping human control.
AIv0 now supports multiple AI models through Vercel's AI Gateway, spanning frontier, cheap, open, and fast options. Users can pick from models including Claude, GPT, Kimi, GLM, Grok, and DeepSeek to build their apps.
AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.
Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.
AIFireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.
AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.
Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.
AILM Studio now offers DeepSeek V4.1 Flash, hosted in the United States with Zero Data Retention enabled by default. The model is accessible through a dedicated LM Studio model page.
AILM Studio has launched its Bionic version, now live for users. The announcement is tied to DeepSeek-V4.1-Flash, the smallest model in DeepSeek's new architecture family, which adds native visual understanding and targets faster inference and higher throughput.
AIOllama says DeepSeek-V4.1-Flash is now fully rolled out on its cloud, hosted in the US and Europe. Prompts and responses are not logged or trained on, and per-token pricing matches the DeepSeek API, including off-peak pricing. The post repeats DeepSeek's claim that the model is more capable, faster, and more cost effective than prior DeepSeek models, including DeepSeek-V4-Pro.
AIDeepSeek-V4.1-Flash is now rolling out to Pro plan subscribers, according to Ollama. The background post says it is first available on Ollama's cloud for Max and Team accounts, with capacity being added to reach all subscribers.
AIModel page:
AIOllama is rolling out DeepSeek-V4.1-Flash on its cloud, starting with Max and Team accounts. The company says it is quickly adding capacity to extend access to all subscribers.
AISebastian Raschka says DeepSeek V4.1 contains a major architecture overhaul using an encoder-decoder setup, and he argues it could have been named V5. The attached diagrams compare DeepSeek V4-Flash (284B) with DeepSeek V4.1-Flash (552B), which has 1M supported context and a 10-layer encoder. The attached charts report a global KV cache per token of 890 bytes for V4.1-Flash, versus 3,514 for V4-Flash and 48,068 for DeepSeek-V3.2.
AILewis Tunstall joked that DeepSeek released a model just after Hugging Face's TRL library removed support for seq2seq models. The post is a lighthearted remark with no further technical details or figures.
AISGLang and Miles ship day-0 inference and RL support for DeepSeek V4.1 Flash, with weights now available. The model is natively multimodal with 552B backbone parameters, 16B active during decode and 8B during prefill, and supports up to 1M context. V4.1 adds shared compressed KV across layers, a two-stage sparse indexer, and a 196B Engram lookup memory.
AIDeepSeek says it will work with the open-source community on inference support for DeepSeek-V4.1-Flash and explore more deployment options. The company is inviting organizations planning large-scale deployments with 2,000 GPUs and a storage cluster to get in touch. The model and a technical report are published on Hugging Face.
AIDeepSeek says its more efficient V4.1-Flash architecture lets it serve more users at lower cost and pass the savings on through lower API prices. Peak/off-peak pricing continues, with off-peak rates at 50% of peak rates. The new pricing takes effect at 04:00 UTC on September 10, 2026.
AIDeepSeek says V4.1-Flash is now live on its API with native multimodal support, accessed through the model name deepseek-flash. The older V4-Flash and V4-Flash-Vision-Exp are retired, while deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. Requests to deepseek-v4-pro will route to V4.1-Flash at V4.1-Flash rates starting 04:00 UTC on Sept 14, 2026, until V4.1-Pro launches.
AIDeepSeek says its V4.1-Flash model needs only 1/4 the HBM and 1/8 the SSD storage for its KV cache compared with the previous generation. Because cache-hit charges often make up a large share of agent costs, the company says the compressed cache significantly reduces those costs.
AIDeepSeek has introduced a 552B-parameter MoE model built on a new Causal Encoder–Decoder architecture, activating 8B parameters for input and 16B for output. The company says new pre-training methods and larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
AIDeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native multimodal visual understanding. The API now serves it under the model name deepseek-flash, while V4 Flash and V4 Flash Vision Exp were retired and routed to V4.1 Flash. API prices were reduced with the release, and V4 Pro remains available after September 14, 2026.
Why it matters: The release lists benchmark results alongside API model-name changes and retirements, so developers can check both capability claims and migration steps.
AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.
Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.
AIMagic says its new pretraining recipe matches DeepSeek V4 Pro's pretraining while using 50x less compute, roughly half the FLOPs used for GPT-3, or about $0.5M on GB200. The post, which congratulates the team, suggests that during recursive self-improvement, automated AI researchers may be less bottlenecked by compute than expected.
AIDeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.
Why it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.
AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).
AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.
AIDeepSeek's API now accepts multimodal input through the model deepseek-v4-flash-vision-exp, supporting mixed text and image requests. Each image is billed at up to 384 tokens at V4-Flash pricing, and it works with Chat Completions, Messages, and Responses endpoints. Images can be supplied as base64, external URLs, or via the Files API.
AIDeepSeek's Files API is now live and free to use. Users upload an image once and reference it by file_id, saving request bandwidth and avoiding re-uploads across requests.
AIDeepSeek says its V4-Flash-Vision-Exp works smoothly across agent frameworks, combining visual understanding with a range of tools. The post presents this multimodal capability as a way to unlock more practical agent workflows.
AIDeepSeek has made its experimental multimodal model DeepSeek-V4-Flash-Vision-Exp available on the DeepSeek API Platform. The company says it matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge. On multimodal agent benchmarks it improves substantially over V4-Flash and approaches Opus-4.8, and DeepSeek Harness 0.1.1 was released the same day with support for the new model.
AIDeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.
Why it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.
AIThe author argues that agents editing their own harnesses and training loops are only bounded self-improvement so far. Recursion would require the system to also raise and keep an honest evaluation bar, which current evidence does not show.
AIDeepSeek has released DeepSeek Harness v0.1 in Developer Preview, opening the codebase under the MIT license for developers building agent harnesses. The harness is built on the Cordis meta-framework and treats models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI as plugins that can be mixed, matched, replaced, and extended.
Why it matters: The source specifies the MIT license and a plugin-based architecture covering models, tools, and sessions, which helps developers assess extensibility before adopting it.