Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Google's AI calls being alone with a Hindu "completely safe" and responds to being alone with a Christian by suggesting you call 911. The gap is a fossil record of the model's safety training. Safety tuning works by drilling a model on thousands of flagged prompts until it produces careful, rehearsed answers. Red teams have spent years stress-testing prompts about Jewish and Muslim people, because those groups are the most frequent targets of hate speech online. So for those phrasings, the…