Homebench – Benchmark local LLMs for speed, memory, and quality
(news.ycombinator.com)
1.
2.
AirLLM 70B inference with single 4GB GPU
(news.ycombinator.com)
3.
Petals: Run LLMs at home, BitTorrent-style
(news.ycombinator.com)
4.
5.
6.
Ollama: All Aboard Open Models
(news.ycombinator.com)
7.
Same model, same Q4_K_M label: 5.02, 5.07 and 5.27 bits per weight
(news.ycombinator.com)
8.
Benchmarking 15 "E-Waste" GPUs with Modern Workloads
(news.ycombinator.com)
9.
10.
Mark Zuckerberg Just Got Rather Badly Humiliated
(futurism.com)
11.
12.
Show HN: Smart model routing directly in Claude, Codex and Cursor
(news.ycombinator.com)
13.
What I'm Finding About LLM Code Style and Token Costs
(news.ycombinator.com)
14.
RubyLLM: A Ruby framework for all major AI providers
(news.ycombinator.com)
15.
RubyLLM: A single, beautiful Ruby framework for all major AI providers
(news.ycombinator.com)
16.
In the Weights is your new AI-centric vanity search
(techcrunch.com)
17.
LLMs Are Complicated Now
(news.ycombinator.com)
18.
Two Qwen3 models on one DGX Spark: the residency math
(news.ycombinator.com)
19.
20.
21.
Running local models is good now
(news.ycombinator.com)
22.
Applying Brevity and Language Efficiency in Prompt Engineering
(news.ycombinator.com)
23.
Will Meta's $14 Billion Bet on AI Ever Pay Off?
(slashdot.org)
24.
How to setup a local coding agent on macOS
(news.ycombinator.com)
25.
How to Setup a Local Coding Agent on macOS
(news.ycombinator.com)
26.
27.
A 10 year old Xeon is all you need
(news.ycombinator.com)
28.
A 10 year old Xeon is all you need (for 26B-A4B MTP Drafters without GPU)
(news.ycombinator.com)
29.
Odysseus – self-hosted AI workspace
(news.ycombinator.com)
30.
Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA
(news.ycombinator.com)