Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
“Why Does Post-Training Quantization Work?” This paper shows that pretrained LLMs are inherently robust to quantization because layers actively counteract accumulated quantization errors, while high-dimensional LM-head geometry preferentially preserves top-token predictions. This lets 4-bit models stay close to full precision despite substantial hidden-state perturbations. https://alphaxiv.org/abs/2609.11716
