A long prompt can slow tokens that a model was already generating because prompt processing and token generation share a serving worker. The prompt is computed in a prefill phase; each active response then needs repeated decode steps. A vLLM report records decode throughput falling sharply during long prefills, even while its KV cache had room left. That points to a latency problem, but it does not yet establish the cause.

Prefill and decode ask different things of the GPU

When a request arrives, prefill processes the prompt and builds the key–value (KV) cache used by later attention steps. The model can process many prompt tokens together, making prefill compute-intensive. Decode uses that cache to produce the response one token at a time. Each next token depends on the preceding one, and generation often becomes sensitive to moving model weights and cache data through memory. Prompt length alone does not tell you how much prefill work remains: a prefix-cache hit can reuse KV state for matching earlier tokens.

A serving system batches work from different requests to keep the accelerator busy. If one request has a large uncached prompt while others are decoding, the runtime must decide how much prompt work to admit alongside the next-token steps. vLLM’s current documentation describes chunked prefill: it divides a large prompt into pieces, prioritizes pending decode requests, then schedules prefill work within the remaining token budget. A chunk is smaller than the full prompt, but its token count is not a promise about how long that GPU step will take.

That distinction matters. A token budget limits scheduled work by count; it does not directly cap execution time. The cost of a chunk can vary with model architecture, attention path, batch shape, parallelism, and the hardware running it. If a step takes longer than expected, a decode request can miss its next-token target even though it remains active and the KV cache is not full. vLLM’s scheduler documentation explains the intended chunking policy; the incident below shows why operators still need latency measurements from their own model and hardware combination.

The reported slowdown is large, but its cause is open

In issue #54919, a user described a two-node DGX Spark deployment running a quantized Qwen3.8-Flash-Next model. With several long-context agent requests in flight, aggregate decode throughput reportedly dropped from about 45–107 tokens per second to 0.5–5 for periods of three to seven minutes. The report says the waiting queue was empty, KV-cache use remained below capacity, and there was no out-of-memory error, preemption, or restart. Throughput recovered when the long prefill finished.

The report is not a controlled benchmark. The issue’s environment-capture field is still a placeholder, the latest measurements used a locally modified scheduler, and the author said a public synthetic reproducer was not yet available. A later commenter added a separate data point across several vLLM versions, reporting a similar drop during uncached prefill and normal decode when the same prompts were full prefix-cache hits. That strengthens the case for investigating prefill interference, but comments in an issue are still reports to reproduce—not a maintainer-confirmed diagnosis.

The key observation is narrower than “vLLM has a scheduler bug”: long-prefill work coincided with very slow decode in these setups, then decode recovered when prefill ended. The issue does not isolate whether the time went to model kernels, attention metadata, memory traffic, communication between devices, scheduling gaps, or an interaction among them. The report itself lists several of these as questions for further instrumentation. Read the issue and its follow-up measurements for the exact configurations and limitations.

Token throughput can hide a bad experience

Aggregate tokens per second is useful for capacity planning, but it can conceal pauses in individual streams. A service may produce many prompt tokens while a user sees their answer stop mid-sentence. For interactive workloads, measure time to first token (TTFT) and inter-token latency (ITL)—the time between successive output tokens—per request, and report tail values such as p95 and p99. vLLM exposes per-request fields for queue time, TTFT, generation time, mean ITL, and overall output throughput; these fields describe different parts of the request and should not be treated as interchangeable.

For a diagnosis, align those request timings with the engine’s per-iteration timings and the actual scheduled prefill-token count. Record active and waiting requests, cached versus newly processed prompt tokens, KV-cache utilization, preemptions, and any retries. vLLM documents a per-request mean ITL; calculate p95 and p99 from streamed token timestamps or suitable histograms rather than assuming that mean captures pauses. On the hardware side, capture GPU utilization, memory bandwidth, clocks and power, plus interconnect activity for multi-GPU runs. High utilization alone is not proof of useful progress: it says the device is busy, not which operation is consuming the time.

A small test matrix can separate the likely causes

Start with three runs on the same model, server version, hardware, and output length: decode-only traffic; the same decode traffic plus an uncached long prompt; and the same prompt served as a confirmed prefix-cache hit. Repeat each run, then vary the prefill chunk or token budget while keeping the request mix fixed. If only uncached prefill triggers the pause, that narrows the search to the work performed while building the prompt state. It still does not identify which kernel or system resource is responsible.

Next, change one factor at a time: model version, attention backend, speculative decoding, tensor or expert parallelism, and chunk size. Keep exact software commits and launch settings with the traces. A useful report includes the baseline and stressed p50/p95/p99 ITL, TTFT, per-step duration, cached-token counts, and resource counters on one timeline. Without that alignment, a throughput screenshot cannot distinguish an oversized prefill step from memory migration, communication overhead, or a client retry loop.

Latency isolation is the real scheduling target

Chunked prefill is a way to share the GPU more carefully, not a hard real-time guarantee. Recent research explores choosing each prefill chunk against the next-token deadlines of active requests, aiming to use available slack without letting a prompt monopolize a step. In a September preprint, SLOWeave reports 39% higher goodput than its strongest fixed-chunk baseline on mixed traffic, and 38% on long-context traffic, under a 25 ms time-per-output-token objective. Those are results from the paper’s evaluated workloads, not proof that the method fixes issue #54919 or is ready for every serving stack.

The practical standard is straightforward: adding a long prompt should not cause unbounded gaps in responses that are already streaming. Operators need to set a latency objective, measure whether each decode step meets it, and test the workload mix under the exact runtime and hardware they deploy. If the pause disappears with a cache hit, a smaller chunk, or a different parallelism setting, that is a clue for the next experiment—not yet the explanation.

At larger scale, a team can place prefill and decode on separate workers, shifting some contention into KV-cache transfer and request routing. EyesTech’s earlier account of GLM’s prefill/decode design describes that architectural option. For a single shared worker, however, the open question remains whether the runtime’s scheduling decisions bound the time active users wait between tokens—not just how many tokens fit in the next batch.

Categorized in:

Frontier AI & Models,

Last Update: October 1, 2026