Sundial put the effort in getting RLVR right: training on verified fixes instead of artificial errors, smart rewards that disincentivize hacking, and deterministic scoring.
The result — a trained Inkling-Small that fixes 83.7% of LaTeX errors in under 1 second and $.0013.
