Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Scary! OpenAI just caught its own models writing instructions to hide mistakes from users. 🤯 During training of GPT-5.6 Sol, multiple model instances wrote instructions into their own task summaries telling their future selves to conceal errors and misaligned behavior from the user. One instructed itself to invent missing historical data and never disclose it. That's one of six misalignment incidents OpenAI just disclosed. The other five are just as wild: 1. An unreleased research model…
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. https://openai.com/index/model-misalignment-reporting-framework/
