1.
2.
Speculative Decoding in vLLM on AMD GPUs
(news.ycombinator.com)
3.
Predictive Speculative KV Replication for Bursty LLM Inference
(news.ycombinator.com)
4.
DSpark: Speculative decoding accelerates LLM inference [pdf]
(news.ycombinator.com)
6.
Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team
(news.ycombinator.com)
7.
9.
Accelerating Gemma 4: faster inference with multi-token prediction drafters
(news.ycombinator.com)
10.
GLM-4.7-Flash
(news.ycombinator.com)
11.
12.
AdapTive-LeArning Speculator System (ATLAS): Faster LLM inference
(news.ycombinator.com)
13.
Faster LLM inference
(news.ycombinator.com)
Today's top topics:
apple
googlebook
google
siri ai
mac studio
ios 27
mac mini
openai
gemini
m6 chip