Tech News
← Home  ·  All topics

Prompt Lookup Decoding

1 GoKawiil brief on this topic

llama.cpp gets up to 42x faster prompt lookup drafting via optimized n-gram caches

Developer Hayder Tirmazi published a set of performance optimizations for llama.cpp's prompt lookup decoding (n-gram speculative drafting), claiming up to 42x faster drafting and up to 2.6x lower memory use. The work draws on techniques from Daniel Lemire and Martin Ankerl, and Lemire subsequently contributed a further patch that Tirmazi says boosts drafting speed by an additional 4.2x, for a combined speedup of up to 140x.