Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face
Original titletencent/Simple-Attention-Sparsification
AISummary
Tencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss.
The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build.
The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.
Source: Tencent · new models on Hugging Face · huggingface.coPublished · added here