Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
This should be talked more about. GLM finds that PPO > GRPO for agents that compacts and therefore breaks the turn history.
“We therefore move from group-wise optimization to a critic-based PPO formulation that learns from individual rollouts, relying on a critic to estimate token-level advantages rather than group-relative comparisons. This single-rollout formulation fits compaction naturally, as it places no constraint on how many traces a prompt produces or on their relative lengths” It is so true that PPO fits better in multi-trace credit assignment for agenti RL. Still surprised how @Zai_org speeds on carrying this through at scale! 🚀 🐐