Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
“MiMo-V2.6 Scaling Reinforcement Learning Towards Self-Improvement” Xiaomi MiMo just completed their largest RL scaling run so far, openly. And it came along with a technical report, outlining how they've scaled RL across compute, environments, and grading. tl;dr The run costs $2.6M for Pro and $0.9M for Flash, with ~44% spent on rollouts, 41-44% on training, and 14% on grading. Each RL step uses 1,568 prompts x 16 rollouts, producing ~25K trajectories and 2.7-3.7B tokens, with contexts…
