Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
The evidence that Astra is more aligned than prior AIs seems dubious to me. Evidence appears consistent with the AI being as or more interested in score-seeking at the expense of user intent, but having beliefs+instincts that the scorer will catch a broader range of cheating behaviors. At a more basic level, the AI is extremely evaluation-aware and much less monitorable than prior AIs, making detecting misalignment much more difficult. If you train against specific reward hacks you ended up…