Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Experimenting with automated evals, and whether more compute = better checks (LLM judges). Yes, kind of!
Smarter models wrote better evals, and higher effort led to better evals for small models like Luna.