llama.cpp prompt-lookup drafting made up to 42× faster
Hayder Tirmazi’s Sept 26 write-up speeds llama.cpp’s n-gram prompt-lookup drafting up to 42× and cuts peak memory ~2.6× via reference reads, dense outer maps, sorted-vector followers, and a Lemire constmap for the static cache—acceptance rate unchanged. Gains concentrate on repetitive drafting (code edits, structured output), not general chat.