Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Bullish on this trend of making post-training more accessible. A new post-training era is upon us. If you work on agentic RL, long-context tasks (a big focus today) are expensive, inefficient, and don't scale well. I've been diving into RL envs and evals for long-context tasks, and I can see this being useful. In agent RL, rollouts use most of the tokens. Every turn re-reads the whole growing context, including tool outputs, files, and earlier turns. Tinker just cut the price of those…
Tinkerers have been busy scaling up long-context RL! We’ve made significant improvements to Tinker’s efficiency to support those, and are passing these on with price cuts up to 70%. GLM-5.3-Flash and DeepSeek-v4.1-Flash are also live for cost-efficient long-context work. https://t.co/in9eAAFO6O