LFM2.5-VL-DSpark speeds up vision-language model decoding on GPUs and edge devices
Original titleLFM2.5-VL-DSpark: Accelerating vision-language models on edge and beyond
AISummary
Liquid AI released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, delivering decoding throughput gains of up to 2.66× on GPUs and 3.13× on edge devices. The drafter adds about 280M parameters, an 8.9% increase in the deployed model's parameter count, and is available on Hugging Face with support in llama.cpp, SGLang, and MLX-VLM.
Source: Liquid AI Blog · liquid.aiPublished · added here