Run Qwen3.8 27B locally: real numbers from my Mac Studio
(news.ycombinator.com)
1.
2.
DFlash 2: Keep Drafting Parallel
(news.ycombinator.com)
3.
Unsloth Dynamic 3.0 GGUFs
(news.ycombinator.com)
4.
Show HN: Shoehorn – Quantize any model down to run on your machine
(news.ycombinator.com)
5.
Llama.cpp v0.1.0
(news.ycombinator.com)
6.
Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP
(news.ycombinator.com)
7.
llama.cpp
(news.ycombinator.com)
8.
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
(news.ycombinator.com)
9.
Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp
(news.ycombinator.com)
10.
11.
12.
Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
(news.ycombinator.com)
13.
Meta Muse Glimmer – Open weights 30B local coding model
(news.ycombinator.com)
14.
Meta Muse Glimmer – open weights 30B local coding model
(news.ycombinator.com)
15.
Homebench – Benchmark local LLMs for speed, memory, and quality
(news.ycombinator.com)
16.
17.
18.
Running local models is good now
(news.ycombinator.com)
19.
How to setup a local coding agent on macOS
(news.ycombinator.com)
20.
Odysseus – self-hosted AI workspace
(news.ycombinator.com)
21.
Liquid AI reveals 8B-A1B MoE trained on 38T
(news.ycombinator.com)
22.
Social Animus
(news.ycombinator.com)
23.
A Comma and a Question Mark, Redux: Quick Terminal Helpers Using Pi
(news.ycombinator.com)
24.
A Comma and a Question Mark
(news.ycombinator.com)
25.
26.
DeepSeek-V4-Flash means LLM steering is interesting again
(news.ycombinator.com)
27.
What's in a GGUF, besides the weights – and what's still missing?
(news.ycombinator.com)
28.
Running local models on an M4 with 24GB memory
(news.ycombinator.com)
29.
DeepSeek 4 Flash local inference engine for Metal
(news.ycombinator.com)
30.
Show HN: Adam – An embeddable cross-platform AI agent library
(news.ycombinator.com)
Today's top topics:
openai
google
apple
samsung
generative ai
android authority
chatgpt
hugging face
meta
nvidia