AirLLM 70B inference with single 4GB GPU
(news.ycombinator.com)
1.
2.
The Economics of Speculative Decoding
(news.ycombinator.com)
3.
Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution
(news.ycombinator.com)
Today's top topics:
openai
anthropic
apple
ai safety
google
iphone 18 pro
ios 27
dario amodei
artificial intelligence
nvidia