Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
(news.ycombinator.com)
1.
2.
My first impressions on ROCm and Strix Halo
(news.ycombinator.com)
3.
Nvidia says it can shrink LLM memory 20x without changing model weights
(venturebeat.com)