85.3 GFlops: Optimizing FP32 Matrix Multiplication on a Single AMD Zen 3 Core
(news.ycombinator.com)
1.
2.
3.
AI Compute Extensions (ACE) Specification
(news.ycombinator.com)
4.
[x86] AI Compute Extensions (ACE) Specification
(news.ycombinator.com)
5.
Inference cost at scale with napkin math
(news.ycombinator.com)
6.
They’re made out of weights
(news.ycombinator.com)
7.
"They're made out of weights"
(news.ycombinator.com)
8.
Training an LLM in Swift, Part 1: Taking matrix mult from Gflop/s to Tflop/s
(news.ycombinator.com)
9.
Anatomy of High-Performance Matrix Multiplication (2008) [pdf]
(news.ycombinator.com)
10.
Is Matrix Multiplication Ugly?
(news.ycombinator.com)