HomeAIGRPO Tuning Lifts LFM2.5-350M Structur
AI

GRPO Tuning Lifts LFM2.5-350M Structured Output to 29.7%

A 100-step training recipe using TRL improved LFM2.5-350M schema compliance on the IFStruct benchmark from 22.6% to 29.7%.

WHAT YOU NEED TO KNOW
  • Fine-tuning LFM2.5-350M with GRPO for 100 steps increased overall IFStruct accuracy from 22.6% to 29.7%.
  • The recipe trained approximately 6 million parameters using about 500 samples on a 16-gigabyte GPU setup.
  • JSON pass rates rose from 18.0% to 31.9%, while average inference latency shifted from 1,453 to 1,518 milliseconds.

A 100-step reinforcement learning run raised the structured-output accuracy of Liquid AI's 350-million-parameter model from 22.6% to 29.7%, Hugging Face reported on Sept. 3, 2026.

Researchers fine-tuned the LFM2.5-350M model using Group Relative Policy Optimization through the TRL library. The workflow trained roughly 6 million parameters—about 1.66% of the model—by applying LoRA adapters across the hybrid attention and convolution architecture.

Training relied on about 500 samples from the nvidia/Nemotron-RL-instruction_following-structured_outputs dataset. Hugging Face appended fenced code block instructions to 40% of the prompts and converted a disjoint 20% into top-level array tasks. Three weighted rewards scored completions on format compliance, top-level field counts, and JSON Schema validation.

Evaluation took place on the 2,000-sample IFStruct benchmark using llama.cpp on an Apple M5 Max MacBook Pro. In the baseline run, the base model passed 452 tasks, scoring 18.0% on JSON and 27.2% on YAML. After 100 steps of GRPO, the fine-tuned checkpoint passed 594 tasks, lifting the JSON pass rate to 31.9% while YAML accuracy reached 27.5%.

The fine-tuned model recorded an average latency of 1,518 milliseconds, compared to 1,453 milliseconds for the base checkpoint. The final 29.7% score remains below Qwen3.5-2B's 33.15% result on IFStruct, though the training process ran on a single 16-gigabyte GPU.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →