Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
"Alignment Forecasting: Predicting Misalignment From Training Data" This paper has an AI inspect the training data and forecast whether fine-tuning will increase behaviors like deception, sycophancy, or sabotage. Their decomposed forecaster reaches 0.801 at predicting whether fine-tuning will trigger misalignment, versus 0.48~0.65 for directly asking frontier LLMs to make the same prediction, and its signals can also help filter risky training examples. https://alphaxiv.org/abs/2609.35805
