Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
"The Surprising Effectiveness of Shared Memory in Looped Transformers" Each loop for looped transformers still maintains its own KV cache, which leads to a growing memory. This paper introduces Looped Prediction Transformers, where the first loop builds a shared KV cache, while later loops reuse it and keep only a small local window. Surprisingly, sharing memory also improves model quality. At five loops, the hybrid model achieves lower perplexity and higher downstream accuracy than a…
