Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
(news.ycombinator.com)
1.
2.
Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data (2024)
(news.ycombinator.com)
3.
Matrix Multiplications on GPUs Run Faster When Given "Predictable" Data
(news.ycombinator.com)
4.
CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
(news.ycombinator.com)
5.
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication Through RL
(news.ycombinator.com)
Today's top topics:
acer
google
openai
android authority
ifa 2026
apple
samsung
anthropic
dell
machine learning