Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Sweet new benchmark from Surge looking at how close AI is getting to a staff-level software engineering. We're half-way there. https://x.com/lennysan/status/2108596336…
introducing sudo L7, our new benchmark for staff-level software engineering. coding agents are getting very good at being L3s. give them a well-defined ticket and they can write excellent code. but a staff engineer is valuable for everything that happens around the code. what should we build? what already exists? what's going to break in prod? what did the tests miss? should we even be doing this? sudo L7 has 60 tasks, mostly in private production repos from real companies, written by engineers who've actually owned these systems. expert rubrics grade everything from functional correctness to engineering craft, architectural judgment, thought partnership, and unnecessary complexity. the best agents succeed on only ~45% of tasks. coding is only part of the job. for now, we're keeping them at idiot genius L3.
