Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
(news.ycombinator.com)
1.
2.
Fixed three bugs that made Qwen3.5-122B a daily driver on Mac Studio
(news.ycombinator.com)
3.
Performance per dollar is getting faster and cheaper
(news.ycombinator.com)
4.
GLM5.2 on AMD MI355X at 2626 tok/s/node at over 2x lower cost than Blackwell
(news.ycombinator.com)
5.
I made a kernel 2.2x faster. It made my training loop 3x slower
(news.ycombinator.com)
6.
Mechanism of age-related accumulation of mtDNA mutations in human blood
(feeds.nature.com)
7.
The DNA virome varies with human genes and environments
(feeds.nature.com)
8.
Freemediaheckyeah
(news.ycombinator.com)
10.
Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0
(news.ycombinator.com)
11.
Deploying DeepSeek on 96 H100 GPUs
(news.ycombinator.com)
12.
Glyn: Type-safe PubSub and Registry for Gleam actors with distributed clustering
(news.ycombinator.com)
13.
PJ5 TTL CPU
(news.ycombinator.com)