Skip to content
Tech News
clear
Topics: Today This Week This Month This Year
1.
Shapelearn Qwen 3.8 27B (13.1 GB VRAM) (news.ycombinator.com)
2.
Getting 50 GB/S Back from the Apple Neural Engine (news.ycombinator.com)
3.
Georgi Gerganov on llama.cpp/ggml future after Nvidia acquisition of HuggingFace (news.ycombinator.com)
4.
A practical guide to running 8x RTX PRO 6000's (news.ycombinator.com)
5.
Run Qwen3.8 27B locally: real numbers from my Mac Studio (news.ycombinator.com)
6.
Kids outlearn AI—and we still don’t know why (technologyreview.com)
7.
Vomit: Clean up Claude 5's token output with a separate LLM (news.ycombinator.com)
8.
Clean up Claude 5's token vomit with a separate LLM (news.ycombinator.com)
9.
DFlash 2: Keep Drafting Parallel (news.ycombinator.com)
10.
Unsloth Dynamic 3.0 GGUFs (news.ycombinator.com)
11.
Show HN: Shoehorn – Quantize any model down to run on your machine (news.ycombinator.com)
12.
Llama.cpp v0.1.0 (news.ycombinator.com)
13.
Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP (news.ycombinator.com)
14.
llama.cpp (news.ycombinator.com)
15.
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp (news.ycombinator.com)
16.
Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp (news.ycombinator.com)
17.
Meta's 'Open' Muse Glimmer Model Can Run On a Single Computer (slashdot.org)
18.
Meta returns to open source with Muse Glimmer, an Apache 2.0 licensed 30B parameter AI model optimized for agents — available now (venturebeat.com)
19.
Meta's 'open source' Muse Glimmer model can run on a single computer (engadget.com)
20.
Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows (news.ycombinator.com)
21.
Meta Muse Glimmer – Open weights 30B local coding model (news.ycombinator.com)
22.
Meta Muse Glimmer – open weights 30B local coding model (news.ycombinator.com)
23.
Building a Rust Inference Engine That Matches Llama.cpp (news.ycombinator.com)
24.
Homebench – Benchmark local LLMs for speed, memory, and quality (news.ycombinator.com)
25.
AirLLM 70B inference with single 4GB GPU (news.ycombinator.com)
26.
Petals: Run LLMs at home, BitTorrent-style (news.ycombinator.com)
27.
Unlimited AI tokens aren't unlimited after all as US Army burns through supply (arstechnica.com)
28.
Meta testing StoryKit, an AI iPhone app that creates ‘personalized children’s stories’ (9to5mac.com)
29.
Same model, same Q4_K_M label: 5.02, 5.07 and 5.27 bits per weight (news.ycombinator.com)
30.
Benchmarking 15 "E-Waste" GPUs with Modern Workloads (news.ycombinator.com)
Today's top topics: apple googlebook google siri ai mac studio mac mini gemini ios 27 m6 chip openai
View all today's topics →