Press Space to continue
Finding signal on Twitter is more difficult than it used to be. We curate the best tweets on topics like AI, startups, and product development every weekday so you can focus on what matters.
Press Space to continue
Press Space to continue
Fascinating. Current LLM design is like "hardware Stockholm syndrome". Models were designed to fit CUDA warps and HBM limits as a compromise. Custom chiplets (and interconnects)!
"The current dominance of GPUs in deep learning is largely accidental". To what extent can we accept that the genuinely optimal architecture for models fits cleanly onto a power-of-2 SM based system. Working on GPUs we assume the optimal matrix size or d_model must be a power-of-2. I reckon these are unlikely the *actual* information theoretic optimal for LLMs. There is an extremely large amount of implicit assumptions baked into model architectures based on the hardware they are being trained/served on. Training/Inference silicon disaggregation is just the start. I think we will see lots of advancement with massive model disaggregation using chiplets making packaging and interconnects more important than ever! https://t.co/km5EtX09fv https://t.co/DpJrH1idRw