Skip to contentSkip to stories
Updated

#China

Oct 9

  1. ModelScopeAI score60

    Qwen-Image-2.1-Turbo cuts image generation and editing to 8 denoising steps

    AIModelScope announces Qwen-Image-2.1-Turbo, an accelerated checkpoint that keeps the 7B visual architecture and runs image generation and editing in 8 denoising steps. The source says it uses CFG=1 and prefix KV caching to reuse text and reference-image context across steps, supports 2048 resolution with square, portrait, landscape, and widescreen presets, and loads through QwenImage21Pipeline in Diffusers. It is released under the Qwen Research License Agreement.

    Why it matters: The source names a concrete speedup path, 8 sampling steps and CFG=1 with prefix KV caching, which matters to anyone weighing image generation latency.

    Image from @ModelScope2022's post

Oct 8

  1. SemiAnalysisAI score72

    SemiAnalysis argues China's AI safety regime is speed-first, not frontier-focused

    AISemiAnalysis argues China's real AI safety approach prioritizes rapid development, regulating AI applications and outputs rather than frontier models. Its dataset of 857 releases from nine Chinese developers found only 31 (3.6%) with any published safety result, and only 9 available at launch. The author also reports that technical experts favor binding frontier rules, but none of their demands has been adopted in binding Chinese instruments.

    Why it matters: The piece tests China's stated AI safety position against its releases, statements, and rules, offering a checkable case for how US pacing debates should read Beijing.

Oct 5

  1. Epoch AIAI score62

    How Chinese AI companies make money and why open weights limit their pricing power

    AIChinese AI companies earn about 10% of the combined AI-related revenue of OpenAI and Anthropic, according to Epoch AI as of September 2026. Their main income streams are consumer apps, API access, enterprise and government deployments, licensing fees, and AI-complemented businesses such as cloud and advertising. Releasing model weights lets third-party hosts compete on price, which weakens API margins for model-focused firms like Z.ai and DeepSeek.

    Why it matters: The piece maps how Chinese AI firms earn revenue and why open-weight releases weaken API pricing, giving context for comparing them with US frontier labs.

Sep 21

  1. Tencent HyAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

Aug 26

  1. Unsloth AIAI score78

    Unsloth explains how to run Qwen3.8-Flash-Next locally on 75GB RAM

    AIUnsloth announces that Qwen3.8-Flash-Next can be run locally through its GGUF quantizations. The source says the 1-bit version needs 75GB of RAM or unified memory, and that the 125B MoE model is reported to outperform Claude-Opus-4.6 (Max).

    Why it matters: The source gives concrete local hardware requirements, quantization sizes, and a guide, showing how a 125B MoE model can run on a 75GB RAM setup.

    Image from @UnslothAI's post

Aug 25

  1. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.3-Flash, a natively multimodal model with 320B parameters

    AIZ.ai released GLM-5.3-Flash on Hugging Face, the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters. The source says it outperforms GLM-5.2 across benchmarks at one-tenth the price and approaches Claude Opus 4.8 on coding and agentic benchmarks. It adopts a hybrid sparse and linear attention architecture to reduce long-context serving costs.

    Why it matters: The release shows a hybrid sparse and linear attention design aimed at cutting long-context serving costs, which is useful for comparing efficiency trade-offs.

  2. Z.ai (GLM) · new models on Hugging FaceAI score72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    AIZ.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Jul 27

  1. Kimi.aiAI score86

    Moonshot AI releases Kimi K3 weights and technical report

    AIMoonshot AI is releasing the model weights and technical report for Kimi K3, a 2.8T-parameter MoE model with native visual understanding and a 1M-token context window. The post says the new architecture delivers 2.5x the intelligence per unit of compute, and the company is also opening high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale.

    Why it matters: The source names the model size, context window, and released weights, which helps readers compare its scale and openness with other frontier releases.

    Image from @Kimi_Moonshot's post

Jul 23

  1. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-Base, a 122B deep research agent model

    AIBAAI has released AREX-Base, a 122B-total, 10B-activated Mixture-of-Experts deep research agent built on Qwen3.5-122B-A10B with a 262,144-token context. The model uses an inner research loop and an outer self-improvement loop, and the source reports it scoring 82.5 on BrowseComp and 85.4 on GAIA, under Apache 2.0.

    Why it matters: The release pairs a 122B-parameter deep research agent with benchmark tables against frontier and open models, letting readers compare its search-agent results directly.

Jun 15

  1. Z.ai Release NotesAI score62

    Z.ai Release Notes: GLM-5.2 Adds 1M Lossless Context for Long Tasks

    AIZ.ai's release notes list GLM-5.2 as supporting 1M lossless context, with improved long-horizon task performance and reduced context drift and goal forgetting. The company says GLM-5.2 achieves open-source SOTA performance on coding and long-horizon task benchmarks. The page also includes the newer GLM-5.3 and GLM-5.3-Flash entries, which are listed above GLM-5.2.

    Why it matters: The page lists a dated series of Z.ai model releases, showing how the coding and long-horizon agent line has evolved from GLM-4.5 through GLM-5.2.

Apr 24

  1. Ahmad Al-DahleAI score82

    Ahmad Al-Dahle says DeepSeek-V4's efficient 1M context is its key bet

    AIAhmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.

    Why it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.

Jan 1

  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score75

    Moonshot AI releases open-source multimodal agent model Kimi K2.5

    AIMoonshot AI released Kimi K2.5, an open-source native multimodal agentic model built by continual pretraining on about 15 trillion mixed visual and text tokens. The model card reports a 1T-parameter Mixture-of-Experts architecture with 32B activated parameters and a 256K context length, and it lists benchmark results against GPT-5.2, Claude 4.5 Opus, Gemini 3 Pro, DeepSeek V3.2, and Qwen3-VL-235B-A22B-Thinking. Weights and code are released under a Modified MIT License, with API access on the Moonshot platform.

    Why it matters: The model card gives a full benchmark table against GPT-5.2, Claude 4.5 Opus, and Gemini 3 Pro, useful for comparing open multimodal agent models.

That’s everything