Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
90% vs. 92.5% accuracy. About 150x lower cost. That’s what @vesko_st on our team found comparing TypeSafe’s Jev with Opus on a 40-question long-context reading test. We're offering Jev in @lightfld to do AI automations for cheaper than a deterministic SFDC workflow automation.
Everyone has been impressed by TypeSafe AI’s Jev — I wanted to see for myself how good it is — and how useful it could be for a product like @lightfld. I started with text understanding, because it is the foundation of everything a decision model does: it has to understand the input, and it has to understand and follow the instructions you give it. The short takeaway: TypeSafe's tiny "Jev" classifier is surprisingly good — roughly on the level of Claude Sonnet on classification tasks, while being a couple of orders of magnitude cheaper to run. I looked at two kinds of work. First, general text understanding: three widely used reading-comprehension and commonsense benchmarks, where Jev performed on par with Opus. Because those may well sit in Jev's training data, I also hand-authored a fresh benchmark over a single random Wikipedia article — every wrong option drawn from the text, so nothing could be memorized. There, Jev landed alongside Sonnet and Opus (it tops CommonsenseQA and MMLU-CF, ties Opus on RACE-H, and is ~150× cheaper and ~10× faster on long-context questions). Then I tried tasks closer to sales and customer relationships. There aren't many good public benchmarks, so I dug up older customer-service and negotiation datasets. On those, Jev is a bit better than Claude Haiku and slightly below Sonnet 5. Big caveat: none of these come from real B2B sales, so the findings are directional rather than a description of what we actually do at @lightfld . Overall, the results look good enough that I think small, inexpensive classification models like Jev will have real uses, including in our own product. Link to the full write-up below.