Benchmarking Pocket-Scale Inference
(news.ycombinator.com)
1.
2.
Why your local LLM feels dumber than it is
(news.ycombinator.com)
3.
Show HN: Shoehorn – Quantize any model down to run on your machine
(news.ycombinator.com)
4.
Qwen 3.8 27B
(news.ycombinator.com)
5.
Qwen 3.8 27B is out: open weights, best local dense model yet
(news.ycombinator.com)
6.
Compression is prediction
(news.ycombinator.com)
7.
Compression Is Prediction
(news.ycombinator.com)
8.
Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong
(news.ycombinator.com)
9.
The 4-Bitter Lesson: Balancing Stability and Performance in NVFP4 RL
(news.ycombinator.com)
10.
Asymmetric Quantization: Near-Lossless Retrieval with 97% Storage Reduction
(news.ycombinator.com)
11.
Integer Quantization: Deep Dive
(news.ycombinator.com)
12.
KVarN: Native vLLM KV-cache quantization back end by Huawei
(news.ycombinator.com)
13.
Show HN: I successfully failed at one-shot-ing a video codec like h.264
(news.ycombinator.com)
14.
15.
16.
17.
Quantization from the Ground Up
(news.ycombinator.com)
18.
TurboQuant: Redefining AI efficiency with extreme compression
(news.ycombinator.com)