HomeAIHugging Face Details Local Agent Deplo
AI

Hugging Face Details Local Agent Deployment With LFM2.5-2.6B

Liquid AI and Hugging Face detailed LFM2.5-2.6B, a 2.6-billion-parameter model designed to run multi-step AI agents entirely on local device hardware.

WHAT YOU NEED TO KNOW
  • LFM2.5-2.6B operates in under 2.5 GB of memory with a 128K context window.
  • Decode speeds reach 220 tok/s on Apple M5 Max and 113 tok/s on AMD Ryzen CPUs.
  • Post-training uses a four-stage pipeline including direct agentic reinforcement learning.

Hugging Face detailed the deployment of Liquid AI's LFM2.5-2.6B model on August 4, 2026, outlining an architecture designed to execute AI agents entirely on edge devices like laptops and smartphones. The model requires less than 2.5 gigabytes of memory while supporting native tool calling and multi-step workflows.

Pre-training for LFM2.5-2.6B covered approximately 34 trillion tokens, followed by a mid-training phase that expanded its context window to 128,000 tokens. Post-training converted the base model into an agent across four stages: two rounds of supervised fine-tuning, domain-specific teacher specialization, multi-domain on-policy distillation, and agentic reinforcement learning.

The reinforcement learning infrastructure separates model optimization, inference, and environment execution across distinct services. A sandbox environment hosts black-box agent harnesses like OpenClaw and Hermes Agent, using a proxy component to capture token-level trajectories for training sample validation without altering harness code.

Inference speeds and hardware

Benchmark data published by Hugging Face shows the 2.6-billion-parameter model reaching decode speeds of 220 tokens per second on an Apple M5 Max processor and 113 tokens per second on an AMD Ryzen AI Max+ 395 CPU. On mobile phone hardware, the model operates at 30 tokens per second, while high-concurrency deployment on a single Nvidia H100 GPU generates roughly 15,000 output tokens per second.

Software integration includes day-one support for inference frameworks such as llama.cpp, MLX, vLLM, SGLang, and ONNX, along with compatibility with Hugging Face transformers versions 5.0.0 and above. Both LFM2.5-2.6B and its base variant LFM2.5-2.6B-Base are available on Hugging Face alongside a WebGPU browser research agent demo.

Benchmark performance

Evaluations compared LFM2.5-2.6B against larger models including Gemma 4 E2B, Gemma 4 E4B, Qwen3.5-4B, and Qwen3.5-9B across STEM, instruction following, and agentic tasks. The model led the evaluation group on instruction-following benchmarks and topped tool-use tests against all competitors except the 9.7-billion-parameter Qwen3.5-9B, though larger models maintained a lead in coding tasks.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →
IN THE AI INDEX

Models named in this story, with their current rank on the index:

Qwen3.5-9B · #122 overallSee the full AI Model Rankings →