llama.cpp's 42x drafting result is not a 42x GPU upgrade
Do not buy or postpone a GPU based on the new “42x faster” llama.cpp headline. Hayder Tirmazi’s September 26 write-up reports faster prompt lookup drafting, not answers arriving 42 times faster. It is useful work. It is not a benchmark comparing the machines on your shopping list.
That distinction matters for the upgrade questions showing up this week. One r/LocalLLM buyer is choosing between expanding a GPU PC and moving to a unified-memory box for coding. Another buyer adding a second GPU is also weighing a case, motherboard, and power-supply upgrade. In the r/LocalLLaMA discussion of this optimization, a reader asks the right question: what is the net effect on inference speed?
My advice: separate “my model does not fit” from “my model fits, but I spend too long waiting.” Do not treat this drafting result as an answer to either question without testing your workload.
What got faster
Prompt lookup decoding proposes likely next tokens using text that has appeared before, rather than running a second neural network to propose them. The original method looks for matching token sequences in the prompt. Tirmazi’s llama.cpp implementation overview also describes caches built from the current context, previous runs, and a separate text collection.
The new work makes those cache operations cheaper. The important limit is
in the author’s Experimental Setup section: llama-lookup-stats replays
a text file as simulated model output. It measures drafting time, matches,
and cache loading. It is not timing a coding assistant finishing a task.
The benchmark repository uses WikiText-103 text and runs the lookup test on the CPU. The published results identify an Apple M4 Pro as the test machine. These are the author’s measurements, not tests we reproduced. The author’s update reports an additional optimization taking the drafting speedup up to 140x. That larger number still describes drafting, not end-to-end generation.
The repository pins its variants to commits in the author’s llama.cpp fork. Do not assume your installed app includes them because it uses llama.cpp. Ask which build was tested before trying to reproduce the result.
What this means for your next purchase
If your work involves editing supplied code or answering questions from documents, prompt lookup is worth testing. The original method targets tasks whose output reuses text from the input, not just tasks with long prompts. Its experiments also found different gains for different chat turns. That is a reason to test your own prompts, not to borrow its speed numbers for your machine.
If your main problem is fitting the model and context, keep that as a separate buying decision. The new benchmark’s memory accounting covers the lookup test process and its caches. It does not establish a smaller memory requirement for your chosen model. I would not choose a lower-memory configuration on that evidence.
Start with our model-fit tool and home inference guide to narrow the options. The fit tool is an estimate, not a measured speed result. If an RTX 5090 and a Mac Studio M4 Max are on your list, use their comparison page for the tracked hardware details. This lookup experiment does not pick a winner between them. If you are adding another card instead, read the multi-GPU guide before ordering the rest of the build.
The test I would ask for
Before spending, ask a seller or someone with the proposed setup to run a small set of your real tasks. Use the same model file, quantization, context length, and output limit. Record the runtime version and whether drafting is enabled. Keep a baseline without the optimization.
Measure these separately:
- The wait before output starts. Include a long input like the code or documents you actually use. Record whether the model was already loaded and whether the prompt was cached.
- The speed while output arrives. Do not substitute a drafting rate for the model’s generation rate.
- Time to a usable result. For coding, check that the edit works. A fast response that needs another round is not the outcome you are buying.
- Memory during the task. Test your intended context and number of simultaneous requests, not just whether the model opens.
llama-bench’s documentation
separates prompt processing (pp), text generation (tg), and combined
tests (pg). Use those labels when comparing results. Then time the real
task through your app as well. Our
runtime setup guide is a starting
point if you do not yet have a baseline.
Buy more hardware when a measured limit gets in the way of work you care about. Test software improvements first when the machine already fits the job. This result gives you something specific to investigate, not a reason to multiply every local AI benchmark by 42.