Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
trajectory labeling is really several questions, not one pass/fail - nice framing here that's how Jev as a judge works in LangSmith evals: difficulty and correctness each get their own typed answer, in one pass, on every trace https://langchain.com/blog/jev-is-now-av…
This is a very clever usecase for Jev-like decision models. Trajectory labelling at scale is key for RSI. Why is that? Every trajectory needs to be assessed across several dimensions. 1. How difficult was the task? 2. Was the agent’s solution correct? 3. Did the agent consider a diverse set of design choices before deciding to pick a specific route? 4. Is the agent’s thinking directionally correct? Imagine answering these questions for millions of agents each of which spawns tens or hundreds of subagents! Token cost for trajectory labelling needs to be driven down a lot to do this at scale. This work is unique and solves a rather under-looked aspect of RL research.