Skip to content
Read the original: Hugging Face· huggingface·Published PickAI score67/100

Hugging Face guide shows how to train agent models across multiple harnesses with RL

Original titleThe same model, with the same weights, scores 62% in one agent harness and 33% in another.

AISummary

Hugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness.

The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses.

Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

AIWhy it matters

The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

Read the original x.com

Source: Hugging Face · x.com