Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
When OpenAI built SWE-bench Verified in 2024, human annotators screened 1,699 real GitHub issues. They flagged 38.3% as underspecified. The set that became the standard yardstick for coding agents kept only issues a person could work from. Real bugs skip that filter. They show up as "checkout is broken" in Slack, and somebody has to make the bug happen again before anyone can fix it. Reproduction has always been slow work. As coding agents get faster at writing the fix, it becomes a bigger…
Everyone got a coding agent. Nobody got a QA engineer. Until today. Meet Ship, your Autonomous Quality Engineer. It tests deployed PRs, reproduces bugs from Slack and Linear, and hands Claude or Codex the context to fix them. Try it free: https://t.co/iKhsWrVxu9 https://t.co/lgTklYYuwT