85.3 GFlops: Optimizing FP32 Matrix Multiplication on a Single AMD Zen 3 Core
(news.ycombinator.com)
1.
2.
3.
AI Compute Extensions (ACE) Specification
(news.ycombinator.com)
4.
[x86] AI Compute Extensions (ACE) Specification
(news.ycombinator.com)
5.
Inference cost at scale with napkin math
(news.ycombinator.com)
6.
They’re made out of weights
(news.ycombinator.com)
7.
"They're made out of weights"
(news.ycombinator.com)
8.
Math Jokes in Alice in Wonderland
(news.ycombinator.com)
9.
Training an LLM in Swift, Part 1: Taking matrix mult from Gflop/s to Tflop/s
(news.ycombinator.com)
10.
Analog computing from waste heat
(technologyreview.com)
11.
Anatomy of High-Performance Matrix Multiplication (2008) [pdf]
(news.ycombinator.com)
12.
Galactic Algorithm
(news.ycombinator.com)
13.
Mark's Magic Multiply
(news.ycombinator.com)
14.
80386 Multiplication and Division
(news.ycombinator.com)
15.
Spaced repetition for efficient learning (2019)
(news.ycombinator.com)
16.
Spaced Repetition for Efficient Learning
(news.ycombinator.com)
17.
Is Matrix Multiplication Ugly?
(news.ycombinator.com)
18.
Product of Additive Inverses
(news.ycombinator.com)
19.
Cracovians: The Twisted Twins of Matrices
(news.ycombinator.com)