AirLLM 70B inference with single 4GB GPU
(news.ycombinator.com)
1.
2.
The Economics of Speculative Decoding
(news.ycombinator.com)
3.
Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution
(news.ycombinator.com)
4.
Attention Residuals
(news.ycombinator.com)
Today's top topics:
apple
openai
google
microsoft
samsung
android authority
iphone
android
app store
universal clipboard