Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Recommended read. And I agree that the RSI is also a systems engineering problem. Self-improving agents need better research environments. RSIGym gives a research agent training, inference, evals, and sandboxes as services it can call. The agent spends its budget on experiments instead of rebuilding infra. With Opus 5 as the researcher, the improved system went from 17.67% to 50.33% on SWE-bench Verified. Also cool to see a way to measure the quality of co-evolution between harnesses and…
💡Our view: RSI is a systems engineering problem, not just a model problem. Progress depends on the environment a research agent works in: what resources it can call, what it can change, and how it runs experiments. That environment should reflect real production workflows and be easy for agents to use. At @Evolvent_AI, we built RSIGym around Everything as a Service. Agents can call established research services and iterate on data, training settings, and harness code. We also introduce RSI-Index to measure how well frontier agents jointly improve a target model’s weights and harness. Across 6 research agents and 5 benchmarks, Opus 5 leads at 0.4809. 🧵 We’re releasing the code, experiment configurations, research trajectories, and evaluation logs. Try your own improvement methods with RSIGym. 🌐 Website: https://t.co/vXiqt3YUN5 💻 GitHub: https://t.co/chjL41y9OH 📄 arXiv: https://t.co/UpDi0y6fIV