HomeAILiquid AI Releases Distilled 4-Bit LFM
AI

Liquid AI Releases Distilled 4-Bit LFM2.5 Model Checkpoints

Liquid AI has published four 4-bit LFM2.5 checkpoints that use quantization-aware distillation to recover accuracy lost during model compression.

WHAT YOU NEED TO KNOW
  • Liquid AI released QAD Q4_0 GGUF checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B.
  • The distilled checkpoints retain between 96.5% and 97.4% of their respective BF16 baseline performance.
  • Decode throughput was measured across four hardware targets: MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, and Raspberry Pi 5.

Liquid AI released updated 4-bit checkpoints for four models in its LFM2.5 family on August 19, 2026, according to an announcement published on Hugging Face. The release covers the LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B models formatted as GGUF files.

The checkpoints rely on quantization-aware distillation, a technique where a high-precision teacher model distills into a quantized student model. Liquid AI stated that the checkpoints retain the low memory footprint and throughput of native Q4_0 GGUFs while recovering 97 percent of the average accuracy lost relative to BF16 baselines.

Benchmark evaluations

Liquid AI tested the distilled checkpoints against post-training quantization files across a benchmark suite assessing reasoning, instruction-following, tool use, and agent capabilities. The evaluations included GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, and BFCLv4, calculated as the mean across five repeats. Math testing used GSM8K for the 230M and 350M models, and AIME25 for the 1.2B-Instruct and 2.6B models.

The distilled checkpoints retained 97.1 percent of BF16 baseline performance for the 230M model, 96.5 percent for the 350M model, 97.4 percent for the 1.2B-Instruct model, and 96.6 percent for the 2.6B model.

Hardware and runtime support

Liquid AI measured decode throughput across four hardware targets: an Apple MacBook Pro, a NucBox EVO-X2, a Samsung Galaxy S26 Ultra, and a Raspberry Pi 5. The MacBook Pro and NucBox executed GPU inference, while the Galaxy S26 Ultra and Raspberry Pi 5 used Arm CPU inference.

The 230M and 350M checkpoints matched Q5_K_M quality within evaluation variance at 4 to 33 percent higher decode throughput. The 1.2B and 2.6B checkpoints matched Q4_K_M quality at 3 to 14 percent higher throughput. For the 230M and 1.2B models, the checkpoints also matched Unsloth's UD-Q4_K_XL post-training quantization checkpoints.

The files run in llama.cpp and other runtimes supporting GGUF Q4_0 artifacts, and are hosted directly on Hugging Face.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →