Jasper's guide shows how reward tweaks shape search agent behavior
Original titleTraining a search agent is great for RL because every lever impact model behavior in legible ways. Jasper's guide shows how small updates...
AISummary
Jasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency.
The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.
Source: Tinker · x.comPublished · added here