Speculative Decoding in vLLM on AMD GPUs
(news.ycombinator.com)
1.
2.
“Next-token predictor” is the wrong mental model for LLMs
(news.ycombinator.com)
3.
"Next-token predictor" is the wrong mental model for LLMs
(news.ycombinator.com)
4.
Stop Thinking of LLMs as Next-Token Predictors
(news.ycombinator.com)
5.
How to build a diffusion language model
(news.ycombinator.com)
6.
Continuous Diffusion Language Models (CDLM's)
(news.ycombinator.com)
7.
Autoregressive Language Model on the 6502 Processor
(news.ycombinator.com)
8.
9.
Autoregressive next token prediction and KV Cache in transformers
(news.ycombinator.com)
10.
Towards end-to-end automation of AI research
(feeds.nature.com)
11.
Normalizing Flows Are Capable Generative Models
(news.ycombinator.com)