Skip to content
Read the original: Liquid AI Blog· Pick60/100AI score60/100

Liquid AI releases DSpark draft models for LFM2.5, up to 3.2x faster inference

Original titleLFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook

AISummary

Liquid AI released DSpark speculative decoding draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B on Hugging Face. The draft models reach up to 3.18x throughput improvement on an H100 GPU and up to 2.87x on-device, and the outputs match baseline greedy decoding by construction. Support is available in llama.cpp and SGLang, with the speedup varying by model and dataset.

AIWhy it matters

The release reports measured speedups on both H100 and MacBook hardware, with per-dataset results and acceptance rates that show where speculative decoding helps most.

Read the original liquid.ai

Source: Liquid AI Blog · liquid.aiPublished · added here