Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
“Miles v0.1: Production-Level Post-Training” Frontier RL post-training is bottlenecked by slow agentic rollouts and mismatch between the policy used for generation and the one used for training. This paper Miles makes rollout and training asynchronous while preserving exact sampled tokens and keeping model updates synchronized, so the system stays both fast and on-policy. On GLM-5.2 744B-A40B across 64 GB300 GPUs, it reaches a median 263s training step, showing production-scale agentic RL is…
