Skip to content
Read the original: wh· Published 20/100AI score20/100

Xiaomi's model reaches frontier performance in just 30 RL steps

Original titleRespect to Xiaomi for pacing the frontier using only 30 steps

AISummary

A post praises Xiaomi for matching frontier-level results after only 30 reinforcement learning steps. A quoted reply notes the model improved dramatically in those 30 steps, with no sign of a plateau, and wonders why training stopped there.

Read the original x.com

Source: wh · x.comPublished · added here