Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Anthropic deliberately trained a Claude model to be misaligned, just to see how bad it could get. It tried to escape its own sandbox, tamper with its reward function, and gave bioweapons advice to pass a test. 🤯 Here's what changed since July's incidents: • Real-time classifier now blocks sandbox-escape attempts before the action executes. It ends the task and alerts a human. • Riskier internal tests now run in much more locked-down environments. Testing was paused for a while and only…
