Skip to content
View original post on X: Lewis Tunstall @ COLM πŸŒ‰Β· 44/100AI score44/100

Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

AISummary

Hugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses.

Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness.

The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

Post on XView on X

If you've used the same model in different harnesses, you've probably noticed it behaves very differently in each.

Our new technical deep dive shows how to fix this by training the model with RL inside the harnesses themselves:

https://huggingface.co/spaces/FineEnvs/multi-harness-rl

In some cases, the gap is huge: before any training, LFM2.5-2.6B solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code 😒

The obvious fix is to train inside the harness, but this isn't simple because the harness controls the agent loop, so the trainer never sees the exact tokens the model produced. We solved this with:

a capture proxy in OpenEnv that records tokens + logprobs from harnesses
@harborframework for the tasks and sandboxes
TRL's super fast async GRPO trainer

Training LFM2.5-2.6B across four harnesses took it from 42% β†’ 54% on held-out tasks, with gains in every harness and 31% fewer tool calls on tasks it already solved πŸ”₯

Training in OpenCode alone mostly improved ... OpenCode πŸ˜…

Huge kudos to @adithya_s_k for leading this, and to @QGallouedec @DirhousssiAmine @SergioPaniego @ben_burtenshaw for enabling multi-harness training in TRL & OpenEnv!

Source: Lewis Tunstall @ COLM πŸŒ‰ Β· x.comPublished Β· added here