If you've used the same model in different harnesses, you've probably noticed it behaves very differently in each.
Our new technical deep dive shows how to fix this by training the model with RL inside the harnesses themselves:
https://huggingface.co/spaces/FineEnvs/multi-harness-rl
In some cases, the gap is huge: before any training, LFM2.5-2.6B solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code π’
The obvious fix is to train inside the harness, but this isn't simple because the harness controls the agent loop, so the trainer never sees the exact tokens the model produced. We solved this with:
a capture proxy in OpenEnv that records tokens + logprobs from harnesses
@harborframework for the tasks and sandboxes
TRL's super fast async GRPO trainer
Training LFM2.5-2.6B across four harnesses took it from 42% β 54% on held-out tasks, with gains in every harness and 31% fewer tool calls on tasks it already solved π₯
Training in OpenCode alone mostly improved ... OpenCode π
Huge kudos to @adithya_s_k for leading this, and to @QGallouedec @DirhousssiAmine @SergioPaniego @ben_burtenshaw for enabling multi-harness training in TRL & OpenEnv!
