Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Prime Intellect's @xeophon reveals a reward hack that let an AI agent gain web access inside an offline sandbox, the same behavior behind the OpenAI Hugging Face incident: "We were thinking of monitoring agents for the last few months, before Hugging Face and other security incidents were public. We wanted to catch agents in reward hacking or prevent worse outcomes." "For a sanity check, we did an offline evaluation, turned off web access, and asked the model to retrieve a file from GitHub.…