Skip to content
View original post on X: SGLang· 58/100AI score58/100

SGLang v0.5.21 adds native decisions API and new model support

AISummary

SGLang has released v0.5.21 with a native Decisions API that turns an LLM or VLM into a low-latency classifier and scorer.

The release also lets /v1/score rerank search or RAG results in one call, lets PD instances switch between prefill and decode without restarting, and adds support for models including DeepSeek-V4.1 Flash, Kimi K3, and GLM-5.3-Flash on AMD MI355X. The announcement reports a 22% faster first token on long prompts for DeepSeek-V4.1 Flash and 20.6% higher prefill throughput for Kimi K3 in PD serving.

Post on XView on X
@sgl_project

SGLang v0.5.21 landed! Native decisions API is here 🎉

Some of our favorite updates:

  • Decisions API turns an LLM/VLM into a low-latency classifier and scorer
  • /v1/score can now rerank search or RAG results in one go
  • PD instances can switch between prefill and decode with no restart needed
  • DeepSeek-V4.1 Flash gets 22% faster first token on long prompts
  • Kimi K3 gets 20.6% higher prefill throughput in PD serving
  • GLM-5.3-Flash now runs on AMD MI355X with FP8 / MXFP4 MoE and MTP
  • You can now run MiniMax H3 inside @ComfyUI with SGLang-Diffusion backend

New models include DeepSeek-V4.1 Flash, GigaChat 3.5, MiMo-V2.6, Ling-3.0-flash-VL, IQuest-Q1, Qwen-Image 2.1, FLUX 3 Action, and more.

Full release notes👇

Source: SGLang · x.comPublished · added here