LiquidAI LFM2.5-2.6B-DSpark Speeds Up LFM2.5 Decoding With Speculative Drafting
Original titleLiquidAI/LFM2.5-2.6B-DSpark
AISummary
Liquid AI released LFM2.5-2.6B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-2.6B target, on Hugging Face.
In SGLang on a single H100 with batch size 1, mean decoding throughput rises from 323 to 864 tokens per second, about 2.67x, and on an Apple M4 Max via Metal it rises from 61 to 139 tokens per second, about 2.27x.
Because the target verifies every proposed token, the output matches what LFM2.5-2.6B would generate alone.
Source: Liquid AI · new models on Hugging Face · huggingface.coPublished · added here