Skip to content

DSV4: prefill priority starves in-flight decodes for the whole chunked-prefill burst #483

Description

@huhoo

What happened

On a DSV4 deployment serving long-context coding traffic, Scheduler._schedule_next_batch gives prefill unconditional priority:

batch = (
    self.prefill_manager.schedule_next_batch(self.prefill_budget)
    or self.decode_manager.schedule_next_batch()
)

A long prompt is prefilled in chunks, so it occupies that many consecutive scheduler slots. While those run, requests that are already decoding are never scheduled — the or short-circuits before the decode manager is ever consulted.

Correction (see the comment below). The first version of this issue quoted 384-token chunks, 0.388 s per chunk and a 267 s stall. Those were measured under an older --swa-full-tokens-ratio 0.02 on this box; production runs 0.12, which gives a 3456-token chunk budget. The corrected, controlled measurement is below. The finding holds — it is stronger, and it now comes from a controlled experiment rather than from reading bursts out of live traffic.

Controlled measurement, current configuration

Production config unchanged (--swa-full-tokens-ratio 0.12, --max-running-requests 50, --num-tokens 327680, --cuda-graph-max-bs 64). Two requests, temperature: 0, no other traffic:

  • victim — a short prompt asking for ~700 tokens of output, i.e. a request that decodes for a while.
  • attacker — a 200,019-token prompt, fired 6 s after the victim, asking for 3 tokens.
run victim wall victim output victim decode rate
victim alone 11.8 s 699 tokens 59.4 tok/s
victim + one 200K prefill 90.4 s 699 tokens 7.7 tok/s

A 7.7x collapse in the decoding request's throughput. The attacker's own numbers: 200,019 tokens in 58 chunks over 77 s, i.e. 1.35 s per 3456-token chunk (2598 tok/s prefill).

The engine log shows the mechanism directly:

14:03:48  DECODE  x19     6 s    <- victim decoding normally
14:03:56  prefill x58    77 s    <- 58 CONSECUTIVE prefill steps, no decode between
14:05:13  DECODE  x16     5 s    <- victim resumes

The victim's decode stops for the entire 77 s the other request is prefilling. 90.4 - 11.8 = 78.6 s of added latency against a 77 s burst — the stall is essentially total, not partial.

This is not a capacity problem: during the burst the engine reported rr=1 at token usage 0.46, i.e. the KV pool was less than half full while requests queued.

What I expected

An in-flight decode should make progress during a long prefill burst. The scheduler already carries the marker for this gap:

# TODO: support other policies: e.g. DECODE first

Proposed fix

--decode-interleave-every N: prefill keeps priority, but after N consecutive prefill steps one decode step is taken if a decode is actually runnable. It never idles a slot to keep the promise, and unset reproduces the historical order bit-for-bit.

This has now been deployed on this box and measured. With N=8, same experiment, streaming the victim's output and taking the gap between consecutive chunks:

attacker victim wall max inter-token gap gaps > 5 s gaps > 20 s
victim alone — 11.7 s 0.03 s 0 0
flag unset 79.5 s 90.4 s 77 s 1 1
flag = 8 79.7 s 90.6 s 13.03 s 7 0

Three independent checks agree this is the mechanism, not noise:

  1. 58 chunks / 8 predicts 7 interleaved decode steps; measured gaps > 5 s = 7.
  2. The bound predicted from the measured per-chunk cost (8 x 1.35 s = 10.8 s) against the observed 13.03 s.
  3. Total wall time is unchanged (90.4 s -> 90.6 s) by design: this moves decode from "all after the prefill" to "spread through it". It is a responsiveness fix, not a throughput fix.

Two caveats I want to state plainly rather than bury:

  1. This is a latency/responsiveness fix. It does not make a long job finish sooner — that needs a faster prefill (a larger chunk, cheaper per-token prefill), which is a separate axis. Interleaving and a larger chunk are complementary: a larger chunk shortens the burst but still leaves decodes unscheduled for its whole duration.
  2. Measuring this is easy to get wrong. The engine's Decode batch log line is throttled to every decode_log_interval (20 by default) decode steps, while Prefill batch logs every step — so ~7 interleaved decode steps produce no log line at all and interleaving looks broken when it is working. The numbers above come from streaming the response, not from the log.

PR: #484.

Environment

  • FreeToken: 0.1.2, source build. Base is 4800af0 (internal fork close to main); the branch is rebased onto main at 68a81ff.
  • Checkpoint: DeepSeek-V4-Flash, local snapshot (ModelScope, 2026-08-26). From its config.json: model_type=deepseek_v4, architectures=[DeepseekV4ForCausalLM], 43 hidden layers, 4096 hidden, 256 routed experts, max_position_embeddings=1048576, bfloat16.
  • GPUs: 8x NVIDIA GeForce RTX 4090, 24564 MiB each, driver 595.58.03
  • CPU: 2x Intel Xeon Platinum 8468V, 48 cores/socket (192 logical)
  • System RAM: 1007 GiB
  • OS: Ubuntu 22.04.5 LTS, kernel 5.15.0-174-generic, Python 3.10

Command

ft serve \
  --model /root/models/deepseek-v4-flash \
  --tensor-parallel-size 8 \
  --disable-pynccl \
  --moe-backend offload \
  --moe-cache-auto \
  --kv-reserve-tokens 327680 \
  --num-tokens 327680 \
  --swa-full-tokens-ratio 0.12 \
  --memory-ratio 0.90 \
  --max-running-requests 50 \
  --cuda-graph-max-bs 64 \
  --max-prefill-length 4096 \
  --expert-load parallel \
  --expert-prefetch 2 \
  --enable-cache-report \
  --decode-log-interval 20 \
  --decode-interleave-every 8 \
  --served-model-name deepseek-v4-flash \
  --host 0.0.0.0 \
  --port 1919 \
  --moe-prefill-hit-d2d

The chunk budget above is what this geometry resolves to: --swa-full-tokens-ratio 0.12 -> 308 window pages, 253 reserved at --max-running-requests 50, 55 free -> 3456 tokens. --max-prefill-length 4096 does not change it: on the DSV4 path the pool's chunk budget governs, which is the subject of #115.

Anything else

The same box also showed a per-stream decode collapse (rr=1 17.6 ms/step vs rr=2 213.7 ms/step, 12x). Part of that was a separate defect — --cuda-graph-max-bs 1 forced rr>=2 out of the CUDA graph. But the prefill bursts above are what kept rr pinned near 1 and left the box looking idle while requests queued.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions