Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
This is the best blog post on LLM inference I've seen this year. | Tech Twitter
Press Space to continue
This is the best blog post on LLM inference I've seen this year.
They achieved 10x latency and >1400 tokens/sec by moving speculative decode onto two 2GB SRAM/chip Corsairs, a small cost on top of a standard GPU setup on gpt-oss-120b.
This performance at this price is insane.