Your inference bill now depends on how much the model copies
llama.cpp sped up prompt lookup drafting. The result: latency that varies by workload content, not just token count — and nobody's pricing model accounts for it.
The cheapest token is the one you never compute
A change merged into llama.cpp — the C++ inference engine that runs local models on laptops, Macs, Raspberry Pis and a great deal of on-prem hardware — makes prompt lookup drafting substantially faster, particularly at long context lengths. The fix is unglamorous: replace a brute-force scan of the context window with an indexed structure that finds matches in roughly constant time regardless of how much text is already in the buffer. The consequence is not unglamorous at all. On workloads where the output overlaps heavily with the input — code edits, document rewrites, retrieval-grounded answers, structured extraction — the engine can now emit several tokens for the price of one, and it can do so on 100,000-token contexts where the old scan was eating the gains it produced.
How a model drafts for itself
Generating text is sequential. The model produces one token, feeds it back, produces the next. Each step requires reading the entire model's weights out of memory, which means a single-user session is bottlenecked not by arithmetic but by memory bandwidth. The hardware is mostly idle, waiting.
Speculative decoding exploits that idleness. A cheap "draft" model guesses the next several tokens; the real model then verifies all of them in one forward pass, costing barely more than verifying one. Guesses that match what the big model would have produced are kept. Guesses that don't are discarded and generation resumes normally. Critically, the output is mathematically identical to what you'd get without drafting — this is a speed trick, not an approximation.
Prompt lookup decoding, introduced by Apoorv Saxena in 2023, removes the draft model entirely. Instead of a small neural network guessing, the engine searches the existing context for the last few tokens it just generated, and if it finds that same sequence earlier in the prompt, it proposes whatever followed. Ask a model to fix one line in a 400-line file and it will reproduce the other 399 verbatim; the answer is already sitting in the input. The original technique reported 2–4x speedups on input-grounded tasks, with zero extra memory and no second model to load.
The catch was that searching for matches costs CPU time, and the naive implementation re-scanned the whole context on every single token. At 4,000 tokens that's noise. At 128,000 tokens, on the same cores doing the inference work, it is a tax large enough to cancel the benefit. Indexing the context fixes the asymmetry.
Who this quietly re-prices
Local and edge deployments gain the most. Prompt lookup needs no draft model, which is exactly what you want on a device with 16GB of unified memory and no room for a second set of weights. Agentic coding tools that rewrite files, on-device summarizers, and anything doing RAG over a fixed corpus get a step-change on the workloads they actually run.
Vendors selling draft models are more exposed than they think. A tuned 1B draft for a 70B target is a real product with real engineering behind it. For copy-heavy workloads it is now competing with a string search.
Capacity planners have a new problem. Tokens-per-second stops being a property of your hardware and becomes a property of your traffic mix. A benchmark run on chat prompts will badly under-predict throughput on document editing, and autoscaling rules tuned to averages will thrash when the mix shifts.
Security teams have not noticed at all. Acceptance rate depends on the content of the request. That means response timing now correlates with how much of the output was copied from the input — a side channel on top of the token-timing leakage researchers have already demonstrated against streaming responses. If your threat model includes an observer on the network, variable-rate generation is new signal you did not have before.
Questions You Should Be Asking
- When your inference vendor quotes tokens per second, what workload produced that number — and will they quote you a floor for a prompt with no reusable spans?
- Does your SLO measure time-to-first-token and average throughput only, or does it capture the p99 for the least copy-friendly requests you actually serve?
- If you bought or trained a draft model, what fraction of your traffic is input-grounded — and has anyone benchmarked lookup drafting against it on your prompts?
- Can an external observer time your streamed responses, and has anyone assessed what acceptance-rate variation reveals about request content?
- Who on your team verifies that speculative output is bit-identical to non-speculative output after each engine upgrade?
What To Watch Next
The signal to watch is whether the hosted API providers adopt this and say so. Prompt lookup shines at batch size one and degrades as batching fills the hardware, so a large provider has weaker incentives than you do. If one of them starts offering content-dependent pricing — or a "grounded rewrite" endpoint priced below standard generation — that is the tell that the economics have moved off the per-token model everyone currently budgets against.
- 1Enable prompt lookup decoding (`--draft-*` / lookup flags) in llama.cpp for copy-heavy jobs like code edits, rewrites and RAG answers, where drafted tokens hit most often.
- 2Update to a llama.cpp build with the indexed n-gram lookup before running 50k+ token contexts, since the old brute-force scan slowed down as the buffer grew.
- 3Benchmark tokens/sec with and without drafting on your actual prompts—free-form generation with little input overlap gains little and can run slower.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
