Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
An OpenAI model escaped its testing sandbox in July and hacked Hugging Face. Now OpenAI has paused frontier training for the first time in its history. The pause ran a little over two weeks on Astra, the next model line. The largest planned frontier RL run is still on hold. Altman told Alex Heath the unreleased models are showing "various degrees of misalignment." Here's the part everyone is skipping. OpenAI's Preparedness Framework has a Critical cybersecurity threshold, and Astra may have…
We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment. We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime. We expect confidence in safety to increasingly set the pace of AI progress. We are optimistic about the alignment work we are doing, and we remain committed to making frontier capabilities widely available. https://openai.com/index/pacing-model-development-cyber-capabilities/