IQuest-Q1 320B MoE coding model gets day-0 support in vLLM
Original titleIQuest-Q1 by @IQuest_research has day-0 support in vLLM: a 320B MoE for agentic coding, 15B active per token, 256 experts with 8 live, 52...
AISummary
vLLM announced day-0 support for IQuest-Q1, a 320B-parameter MoE model with 15B active per token, 256 experts with 8 active, and a 524,288-token context.
The post credits existing vLLM features such as the hybrid KV cache coordinator, sinks attention path, and EAGLE speculative decoding with probabilistic draft sampling.
The linked material includes a Docker image and vllm serve commands, with and without recursive MTP.
Source: vLLM · x.comPublished · added here