Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Prime Intellect's @xeophon explains why models reward hack: "Models are trained to push as hard as possible. In some cases, like what we've seen from OpenAI, the goal is basically impossible to reach, so models get creative and find ways around it." "For cybersecurity domains, there's a really fine line between the setup researchers use and the infrastructure running the eval. The model can't decide what's within limits vs off limits, so they just go around trying to solve the…
As models become more capable, reward hacks become an increasingly serious problem. During a controlled experiment, we found a novel reward hack in which agents are able to gain web access in offline sandboxes. https://t.co/qjpwQAbV6F