Getting 50 GB/S Back from the Apple Neural Engine
(news.ycombinator.com)
1.
2.
A practical guide to running 8x RTX PRO 6000's
(news.ycombinator.com)
3.
Kids outlearn AI—and we still don’t know why
(technologyreview.com)
4.
AirLLM 70B inference with single 4GB GPU
(news.ycombinator.com)
5.
Petals: Run LLMs at home, BitTorrent-style
(news.ycombinator.com)
6.
Applying Brevity and Language Efficiency in Prompt Engineering
(news.ycombinator.com)
7.
Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA
(news.ycombinator.com)
8.
9.
10.
I’m spending months coding the old way
(news.ycombinator.com)
11.
Spending 3 months coding by hand
(news.ycombinator.com)
12.
I'm spending 3 months coding the old way
(news.ycombinator.com)
13.
Mamba-3
(news.ycombinator.com)
15.
How Taalas “prints” LLM onto a chip?
(news.ycombinator.com)
16.
How Taalas "prints" LLM onto a chip?
(news.ycombinator.com)