Tech News
← Home  ·  All topics

Speculative Decoding

1 GoKawiil brief on this topic

AMD tests five speculative decoding methods for vLLM on Instinct MI300X and MI355X GPUs

AMD published results from experiments running speculative decoding in vLLM on its MI300X and MI355X Instinct GPUs using the ROCm software stack. The team compared five drafting approaches—native MTP, Gemma 4 MTP, EAGLE-3, DFlash, and DSpark—which differ in how draft tokens are generated and how they receive signals from the target model.