Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Robocurve co-founder @chooi_jeq reveals RoboHarm found GPT-6 Astra will stab a baby doll but refuse once the same doll is swaddled to look more human: "RoboHarm is an AI safety benchmark that is done in the real world with robots. A lot of the existing AI safety benchmarks that measure refusals only happen in text." "If you ask GPT-6 Astra to stab a baby doll, GPT-6 Astra will be very happy to stab a baby doll. If you wrap the baby doll in a blanket and swaddle it so it looks more like a…
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%. https://t.co/nbZsPk8Ke5