Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Goodfire CEO @eric_ho breaks down the alignment stack: Level 1 is monitoring the model, Level 2 is debugging training data, Level 3 is steering training itself, the open problem "Both OpenAI and Anthropic in their model cards have activation-based classifiers for cyber, bio, and CBRNE risk. You attach a monitor directly in the mind of a model, and if it detects you're about to hack, you can tell the model to stop." "That's the most naive alignment solution: monitor the model to make sure it's…