Liquid AI released an experimental speculative decoding draft model called LFM2.5-VL-DSpark for its LFM2.5-VL-3B vision-language system, Hugging Face reported on September 24, 2026. The new drafter introduces a speculative decoding path that accelerates token generation without altering the target model's output quality.
The drafter adds approximately 280 million parameters, expanding the 3-billion-parameter target model's footprint by 8.9 percent. Built with a four-layer, attention-only decoder stack, it taps hidden states from fixed layers of the target network to draft candidate token blocks. Because image patches and text tokens share a projected representation before those layers, the drafter processes vectors of identical dimensionality regardless of input modality. Liquid AI trained the network over 10 epochs on vision-language supervised fine-tuning data, recommending runtime block sizes of eight or nine tokens.
Benchmark tests across six tasks from the MMSpec benchmark demonstrated decode speedups reaching 3.13x on edge devices and 2.66x on Nvidia H100 GPUs. On an Apple M5 Max running MLX, decode speeds increased between 2.30x and 3.13x, yielding end-to-end latency gains from 1.56x to 2.62x. On an Apple M3 Ultra using llama.cpp, decoding accelerated by 1.57x to 2.14x, while end-to-end speed rose by 1.30x to 1.77x. Datacenter runs on the H100 delivered end-to-end latency improvements between 1.64x and 2.27x.
Vision encoding and prompt prefill remain unaccelerated by the system, capping total latency gains under Amdahl's law. Because image processing and initial token ingestion consume large amounts of wall time on edge hardware, speculative decoding only reduces time spent during autoregressive token generation.
Liquid AI published the open-weight checkpoints on Hugging Face in Safetensors and GGUF formats. Day-one runtime support arrived across three software frameworks: SGLang via pull request 40651, llama.cpp via pull request 29339, and MLX-VLM via pull request 2280.
