91.
92.
93.
Universal Claude.md – cut Claude output tokens
(news.ycombinator.com)
94.
Universal Claude.md – cut Claude output tokens by 63%
(news.ycombinator.com)
95.
97.
Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon
(news.ycombinator.com)
98.
Run a 1T parameter model on a 32gb Mac by streaming tensors from NVMe
(news.ycombinator.com)
99.
One All-in-One AI Platform, Endless Business Possibilities for Just $85
(feeds.feedburner.com)
100.
Show HN: A deterministic middleware to compress LLM prompts by 50-80%
(news.ycombinator.com)
101.
Mamba-3
(news.ycombinator.com)
102.
Unsloth Studio
(news.ycombinator.com)
103.
My Journey to a reliable and enjoyable locally hosted voice assistant (2025)
(news.ycombinator.com)
104.
My Journey to a reliable and enjoyable locally hosted voice assistant
(news.ycombinator.com)
106.
How to run Qwen 3.5 locally
(news.ycombinator.com)
107.
Files are the interface humans and agents interact with
(news.ycombinator.com)
108.
Filesystems Are Having a Moment
(news.ycombinator.com)
109.
Uploading Pirated Books via BitTorrent Qualifies as Fair Use, Meta Argues
(news.ycombinator.com)
110.
Unsloth Dynamic 2.0 GGUFs
(news.ycombinator.com)
111.
Show HN: Llama 3.1 70B on a single RTX 3090 via NVMe-to-GPU bypassing the CPU
(news.ycombinator.com)
112.
How Taalas “prints” LLM onto a chip?
(news.ycombinator.com)
113.
How Taalas "prints" LLM onto a chip?
(news.ycombinator.com)
114.
The path to ubiquitous AI (17k tokens/sec)
(news.ycombinator.com)
115.
116.
117.
Sampling at negative temperature
(news.ycombinator.com)
118.
Yann LeCun: Meta ‘fudged a little bit’ when benchmark-testing Llama 4 model
(feeds.feedburner.com)
119.
120.
Show HN: GPULlama3.java Llama Compilied to PTX/OpenCL Now Integrated in Quarkus
(news.ycombinator.com)
Today's top topics:
googlebook
apple
gemini
google
mac mini
m6 chip
chromeos
mac studio
android
m5 ultra