Home › AI › Liquid AI Launches LFM2.5-VL-DSpark to
AI

Liquid AI Launches LFM2.5-VL-DSpark to Speed Up Vision Inference

Liquid AI published an open-weight speculative decoding drafter that boosts vision-language decode speeds by up to 3.13x.

WHAT YOU NEED TO KNOW
  • LFM2.5-VL-DSpark adds 280 million parameters, an 8.9 percent increase over the 3-billion-parameter base model.
  • Decoding speeds improved by up to 3.13x on an Apple M5 Max and up to 2.66x on an Nvidia H100.
  • Integrations launched on day one for llama.cpp, MLX-VLM, and SGLang in Safetensors and GGUF formats.

Liquid AI released an experimental speculative decoding draft model called LFM2.5-VL-DSpark for its LFM2.5-VL-3B vision-language system, Hugging Face reported on September 24, 2026. The new drafter introduces a speculative decoding path that accelerates token generation without altering the target model's output quality.

The drafter adds approximately 280 million parameters, expanding the 3-billion-parameter target model's footprint by 8.9 percent. Built with a four-layer, attention-only decoder stack, it taps hidden states from fixed layers of the target network to draft candidate token blocks. Because image patches and text tokens share a projected representation before those layers, the drafter processes vectors of identical dimensionality regardless of input modality. Liquid AI trained the network over 10 epochs on vision-language supervised fine-tuning data, recommending runtime block sizes of eight or nine tokens.

Benchmark tests across six tasks from the MMSpec benchmark demonstrated decode speedups reaching 3.13x on edge devices and 2.66x on Nvidia H100 GPUs. On an Apple M5 Max running MLX, decode speeds increased between 2.30x and 3.13x, yielding end-to-end latency gains from 1.56x to 2.62x. On an Apple M3 Ultra using llama.cpp, decoding accelerated by 1.57x to 2.14x, while end-to-end speed rose by 1.30x to 1.77x. Datacenter runs on the H100 delivered end-to-end latency improvements between 1.64x and 2.27x.

Vision encoding and prompt prefill remain unaccelerated by the system, capping total latency gains under Amdahl's law. Because image processing and initial token ingestion consume large amounts of wall time on edge hardware, speculative decoding only reduces time spent during autoregressive token generation.

Liquid AI published the open-weight checkpoints on Hugging Face in Safetensors and GGUF formats. Day-one runtime support arrived across three software frameworks: SGLang via pull request 40651, llama.cpp via pull request 29339, and MLX-VLM via pull request 2280.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →