Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
This is a new open-source benchmark that tests whether AI agents can turn scientific diagrams into editable PowerPoint slides. The goal is for an agent to preserve the diagram's meaning, layout, and details. This might be subtle, but it's critical: If you have an agent that changes the direction of a single arrow, the final diagram would be useless. That's precisely what this benchmark tests.
The next test for AI agents is not just whether they can write code. It is whether they can look, understand, code, inspect—and fix. Today we’re releasing PPTBench, a benchmark for Visual Coding through scientific diagram slides reconstruction. 500 tasks. 36 configurations. 18,000 reconstructions. 64.03% failed the semantic check. Only 2.57% passed all gates with no recorded defect. Visual coding is still wide open. 🤓 Read agent sessions on AgentGit: https://t.co/XQppVwg6vV Full breakdown 👇 Project: https://t.co/8nkiEZKFq9 Paper: https://t.co/PC757AHyMG Github: https://t.co/jnsSLsYfdl