Miles v0.1.2 is out.
This release makes Miles steadier and easier to operate in production.
Score centering keeps async RL stable and GPUs busy.
Miles can now run natively on Kubernetes via an experimental backend. RL jobs become ordinary cluster workloads, and orchestration can restart while training continues.
Training gets more flexible, too. Models that learn from each other, like a model and its grader, can train and improve together in one run.
torchtitan from @PyTorch joins Megatron and FSDP as a third training backend, giving teams another PyTorch-native option for training at scale.
DeepSeek-V4.1-Flash and MiMo-V2.6-Flash-RL land in main.
Full release notes 👇
