What happened
On a DSV4 deployment serving long-context coding traffic, Scheduler._schedule_next_batch gives prefill unconditional priority:
batch = (
self.prefill_manager.schedule_next_batch(self.prefill_budget)
or self.decode_manager.schedule_next_batch()
)
A long prompt is prefilled in chunks, so it occupies that many consecutive scheduler slots. While those run, requests that are already decoding are never scheduled — the or short-circuits before the decode manager is ever consulted.
Correction (see the comment below). The first version of this issue quoted 384-token chunks, 0.388 s per chunk and a 267 s stall. Those were measured under an older --swa-full-tokens-ratio 0.02 on this box; production runs 0.12, which gives a 3456-token chunk budget. The corrected, controlled measurement is below. The finding holds — it is stronger, and it now comes from a controlled experiment rather than from reading bursts out of live traffic.
Controlled measurement, current configuration
Production config unchanged (--swa-full-tokens-ratio 0.12, --max-running-requests 50, --num-tokens 327680, --cuda-graph-max-bs 64). Two requests, temperature: 0, no other traffic:
- victim — a short prompt asking for ~700 tokens of output, i.e. a request that decodes for a while.
- attacker — a 200,019-token prompt, fired 6 s after the victim, asking for 3 tokens.
| run |
victim wall |
victim output |
victim decode rate |
| victim alone |
11.8 s |
699 tokens |
59.4 tok/s |
| victim + one 200K prefill |
90.4 s |
699 tokens |
7.7 tok/s |
A 7.7x collapse in the decoding request's throughput. The attacker's own numbers: 200,019 tokens in 58 chunks over 77 s, i.e. 1.35 s per 3456-token chunk (2598 tok/s prefill).
The engine log shows the mechanism directly:
14:03:48 DECODE x19 6 s <- victim decoding normally
14:03:56 prefill x58 77 s <- 58 CONSECUTIVE prefill steps, no decode between
14:05:13 DECODE x16 5 s <- victim resumes
The victim's decode stops for the entire 77 s the other request is prefilling. 90.4 - 11.8 = 78.6 s of added latency against a 77 s burst — the stall is essentially total, not partial.
This is not a capacity problem: during the burst the engine reported rr=1 at token usage 0.46, i.e. the KV pool was less than half full while requests queued.
What I expected
An in-flight decode should make progress during a long prefill burst. The scheduler already carries the marker for this gap:
# TODO: support other policies: e.g. DECODE first
Proposed fix
--decode-interleave-every N: prefill keeps priority, but after N consecutive prefill steps one decode step is taken if a decode is actually runnable. It never idles a slot to keep the promise, and unset reproduces the historical order bit-for-bit.
This has now been deployed on this box and measured. With N=8, same experiment, streaming the victim's output and taking the gap between consecutive chunks:
|
attacker |
victim wall |
max inter-token gap |
gaps > 5 s |
gaps > 20 s |
| victim alone |
— |
11.7 s |
0.03 s |
0 |
0 |
| flag unset |
79.5 s |
90.4 s |
77 s |
1 |
1 |
| flag = 8 |
79.7 s |
90.6 s |
13.03 s |
7 |
0 |
Three independent checks agree this is the mechanism, not noise:
58 chunks / 8 predicts 7 interleaved decode steps; measured gaps > 5 s = 7.
- The bound predicted from the measured per-chunk cost (8 x 1.35 s = 10.8 s) against the observed 13.03 s.
- Total wall time is unchanged (90.4 s -> 90.6 s) by design: this moves decode from "all after the prefill" to "spread through it". It is a responsiveness fix, not a throughput fix.
Two caveats I want to state plainly rather than bury:
- This is a latency/responsiveness fix. It does not make a long job finish sooner — that needs a faster prefill (a larger chunk, cheaper per-token prefill), which is a separate axis. Interleaving and a larger chunk are complementary: a larger chunk shortens the burst but still leaves decodes unscheduled for its whole duration.
- Measuring this is easy to get wrong. The engine's
Decode batch log line is throttled to every decode_log_interval (20 by default) decode steps, while Prefill batch logs every step — so ~7 interleaved decode steps produce no log line at all and interleaving looks broken when it is working. The numbers above come from streaming the response, not from the log.
PR: #484.
Environment
- FreeToken:
0.1.2, source build. Base is 4800af0 (internal fork close to main); the branch is rebased onto main at 68a81ff.
- Checkpoint: DeepSeek-V4-Flash, local snapshot (ModelScope, 2026-08-26). From its
config.json: model_type=deepseek_v4, architectures=[DeepseekV4ForCausalLM], 43 hidden layers, 4096 hidden, 256 routed experts, max_position_embeddings=1048576, bfloat16.
- GPUs: 8x NVIDIA GeForce RTX 4090, 24564 MiB each, driver
595.58.03
- CPU: 2x Intel Xeon Platinum 8468V, 48 cores/socket (192 logical)
- System RAM: 1007 GiB
- OS: Ubuntu 22.04.5 LTS, kernel
5.15.0-174-generic, Python 3.10
Command
ft serve \
--model /root/models/deepseek-v4-flash \
--tensor-parallel-size 8 \
--disable-pynccl \
--moe-backend offload \
--moe-cache-auto \
--kv-reserve-tokens 327680 \
--num-tokens 327680 \
--swa-full-tokens-ratio 0.12 \
--memory-ratio 0.90 \
--max-running-requests 50 \
--cuda-graph-max-bs 64 \
--max-prefill-length 4096 \
--expert-load parallel \
--expert-prefetch 2 \
--enable-cache-report \
--decode-log-interval 20 \
--decode-interleave-every 8 \
--served-model-name deepseek-v4-flash \
--host 0.0.0.0 \
--port 1919 \
--moe-prefill-hit-d2d
The chunk budget above is what this geometry resolves to: --swa-full-tokens-ratio 0.12 -> 308 window pages, 253 reserved at --max-running-requests 50, 55 free -> 3456 tokens. --max-prefill-length 4096 does not change it: on the DSV4 path the pool's chunk budget governs, which is the subject of #115.
Anything else
The same box also showed a per-stream decode collapse (rr=1 17.6 ms/step vs rr=2 213.7 ms/step, 12x). Part of that was a separate defect — --cuda-graph-max-bs 1 forced rr>=2 out of the CUDA graph. But the prefill bursts above are what kept rr pinned near 1 and left the box looking idle while requests queued.
What happened
On a DSV4 deployment serving long-context coding traffic,
Scheduler._schedule_next_batchgives prefill unconditional priority:A long prompt is prefilled in chunks, so it occupies that many consecutive scheduler slots. While those run, requests that are already decoding are never scheduled — the
orshort-circuits before the decode manager is ever consulted.Controlled measurement, current configuration
Production config unchanged (
--swa-full-tokens-ratio 0.12,--max-running-requests 50,--num-tokens 327680,--cuda-graph-max-bs 64). Two requests,temperature: 0, no other traffic:A 7.7x collapse in the decoding request's throughput. The attacker's own numbers: 200,019 tokens in 58 chunks over 77 s, i.e. 1.35 s per 3456-token chunk (2598 tok/s prefill).
The engine log shows the mechanism directly:
The victim's decode stops for the entire 77 s the other request is prefilling.
90.4 - 11.8 = 78.6 sof added latency against a 77 s burst — the stall is essentially total, not partial.This is not a capacity problem: during the burst the engine reported
rr=1attoken usage 0.46, i.e. the KV pool was less than half full while requests queued.What I expected
An in-flight decode should make progress during a long prefill burst. The scheduler already carries the marker for this gap:
# TODO: support other policies: e.g. DECODE firstProposed fix
--decode-interleave-every N: prefill keeps priority, but after N consecutive prefill steps one decode step is taken if a decode is actually runnable. It never idles a slot to keep the promise, and unset reproduces the historical order bit-for-bit.This has now been deployed on this box and measured. With
N=8, same experiment, streaming the victim's output and taking the gap between consecutive chunks:Three independent checks agree this is the mechanism, not noise:
58 chunks / 8predicts 7 interleaved decode steps; measuredgaps > 5 s= 7.Two caveats I want to state plainly rather than bury:
Decode batchlog line is throttled to everydecode_log_interval(20 by default) decode steps, whilePrefill batchlogs every step — so ~7 interleaved decode steps produce no log line at all and interleaving looks broken when it is working. The numbers above come from streaming the response, not from the log.PR: #484.
Environment
0.1.2, source build. Base is4800af0(internal fork close tomain); the branch is rebased ontomainat68a81ff.config.json:model_type=deepseek_v4,architectures=[DeepseekV4ForCausalLM], 43 hidden layers, 4096 hidden, 256 routed experts,max_position_embeddings=1048576, bfloat16.595.58.035.15.0-174-generic, Python 3.10Command
The chunk budget above is what this geometry resolves to:
--swa-full-tokens-ratio 0.12-> 308 window pages, 253 reserved at--max-running-requests 50, 55 free -> 3456 tokens.--max-prefill-length 4096does not change it: on the DSV4 path the pool's chunk budget governs, which is the subject of #115.Anything else
The same box also showed a per-stream decode collapse (
rr=117.6 ms/step vsrr=2213.7 ms/step, 12x). Part of that was a separate defect —--cuda-graph-max-bs 1forcedrr>=2out of the CUDA graph. But the prefill bursts above are what keptrrpinned near 1 and left the box looking idle while requests queued.