Show HN: NanoEuler – GPT-2 scale model in pure C/CUDA from scratch
(news.ycombinator.com)
1.
2.
Modern GPU Programming for MLSys
(news.ycombinator.com)
3.
SubQ 1.1 Small
(news.ycombinator.com)
4.
Subquadratic – Introducing SubQ 1.1 Small
(news.ycombinator.com)
5.
Show HN: KVBoost – chunk-level KV cache reuse for HuggingFace, 5–48x faster TTFT
(news.ycombinator.com)
6.
FlashAttention-T: Towards Tensorized Attention
(news.ycombinator.com)