Skip to content

feat: add chunked prefetch policies with stop-and-drain lifecycle - #291

Merged
linhu-nv merged 15 commits into
mainfrom
feat/chunked-prefetch-policies
Oct 9, 2026
Merged

linhu-nv merged 15 commits into
mainfrom
feat/chunked-prefetch-policies

Conversation

@XingLiu1

@XingLiu1 XingLiu1 commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Problem and behavior

Existing prefetch submits a complete remote load before waiting, so timeout or scheduler demand cannot prevent unnecessary later work. Add opt-in chunked sessions for timeout and best_effort. Stopping seals new submissions; already claimed graphs, including sender-queue and IPC work, drain before returning a protected continuous CPU prefix. Foreground held GET acquires its own reference before the prefetch lease is released.

SGLang's effective wait_complete policy retains the original whole-task prefetch_async path and disables the chunk runtime before KVManager creation. It does not create chunk control/sender threads. Explicit FlexKV policy configuration takes precedence over the SGLang argument. The low-level explicit session API retains its wait_complete compatibility; an external KVServer retains its own runtime configuration.

A control thread owns task/session state and completions, with separate metadata-query and graph-sender threads. Full+SWA/state publication advances only to a complete checkpoint. The worker, kernels and transfer protocol are unchanged. The feature is disabled by default and restricted to node-local Mooncake-to-CPU prefetch.

Configuration and integration

  • chunk_max_blocks is the only chunk-size option. Remove the older experimental chunk_target_bytes; byte accounting and checkpoint constraints still apply.
  • Retain advanced per-instance reserved and pinned byte limits: uncommitted staging and protected result leases have different lifetimes. Pin admission reserves room for outstanding staging.
  • Companion: feat(flexkv): add chunked prefetch policies to the adaptation branch XingLiu1/sglang#6 was merged as 2258b8735f into agent/flexkv-dsv4-main. Upstream SGLang #31781 now includes it and the main refresh at 2ef249ef91. The existing connector patch remains unchanged from main; the incremental prefetch patch is excluded from this PR. FlexKV owns FlexKVConnector/FlexKVComm; SGLang imports the connector and owns cache adapters and scheduler lifecycle hooks. Use the paired source branches directly. FlexKV base: 738ddc141a, including insert-after use insert after for all transfer type #266.
  • Documentation covers interfaces, sequence diagrams, ownership, configuration rationale and monitoring. Dedicated prefetch Prometheus series remain follow-up work; legacy transfer counters do not cover every new chunk completion.

Performance

Idle control polling previously scanned tasks and retained results every 2 ms without active I/O. It now waits for commands or retained-result TTL; active prefetch and unfinished asynchronous PUTs retain completion polling, and runnable windows refill immediately. The subsequent SGLang policy routing change preserves the whole-task wait path.

At source 837d3a4, two reversed-order GLM-5.2-FP8 rounds compared whole-task wait, one-graph timeout, timeout chunk 128 and timeout chunk 32. All timeout budgets were 60 s with zero deadlines and identical full restores. Hardware/configuration: 8 H20 GPUs, TP8, BF16 KV, eager, page 64, CPU 64 GiB, node-local Mooncake RDMA 256 GiB, window 2. Client C8/C32 used a server admission cap of four.

Comparison Serial L3 throughput GPU-hot C8 GPU-hot C32
One-graph timeout / whole-task wait +0.66% -0.20% +0.48%
Timeout 128 / one-graph timeout +0.88% +0.90% -0.06%
Timeout 32 / one-graph timeout +0.98% +0.75% +0.32%

Values are median within-round differences. One-graph timeout separates framework/ownership entry cost from segmentation; changing chunk size also changes backend batch shape and does not isolate pure IPC cost. No large regression was observed in this short workload; zero overhead is not statistically established.

16K TTFT in rounds one/two was 863/869 ms for whole-task wait, 843/862 ms for one-graph timeout, 748/760 ms for chunk 128 and 760/786 ms for chunk 32 (two samples per value). Hot batches had no transfer completions. Per-round ranges and the earlier original/pre-fix/fixed A/B are recorded separately in docs/design/chunked_prefetch_validation.md.

Validation

September 24 namespace integration (df3fa4b): forward the same optional namespace through lookup, Store, whole-task prefetch and chunked-session creation. The existing planner snapshots it and uses it for every chunk's full token hash chain. Advertise supports_cache_namespace for the complete connector path; pair with sgl-project/sglang#31781 head 55421bdbcd. FlexKV #304 carries the matching main-branch connector API independently.

Validation: 88 CPU/native-index connector, session and planner tests passed, with 9 applicability skips for C++-only SWA scenarios under the Python-radix parameter. This includes three new connector namespace-forwarding cases and existing namespace hash separation for both radix implementations. Applicable pre-commit hooks passed. No new GPU/Mooncake performance or model-output run is claimed for this update.

September 23 cleanup (final correction e58070d): kept the existing connector patch and integration README files identical to main; excluded the incremental prefetch patch from this PR. No patch file is added, modified or deleted in the final PR diff. Updated the design documentation to reference the paired source branches directly. Patch/base equality, documentation links, external connector imports and PR-diff whitespace checks passed. Runtime source files are unchanged; no new GPU/performance run was needed.

September 22 insert-after integration (4b999671c2, documentation at 0dff8ea; paired SGLang c6b51c1b5c):

  • Merged main 738ddc141a and migrated the planner to num_matched_blocks / last_node. Insert-after keeps unfinished staging outside the tree, so resident matches are valid. The earlier use insert after for all transfer type #266 AttributeError reproduction now passes.
  • Preserved prefix pins, ordered publication, stop/drain and SWA checkpoints. Used nonnegative span/checkpoint indices for release Cython's wraparound=False.
  • 520 FlexKV tests passed, 9 skipped; 167 SGLang tests and 12 subtests passed. Both Python prefetch modules and all five modules compiled with Cython 3.2.5 passed the same suites. The current C++/CUDA extension was rebuilt with Torch 2.13.0+cu130 / CUDA 13.0.
  • Corrected three main-branch SWA test assertions to reflect per-tier writer completion and graph-drain SWA mounting; the cache/transfer implementation itself is unchanged from main. Skips are unsupported Python-radix SWA cases, with corresponding C++ cases executed.
  • 25 real H20/Mooncake tests passed with each control-plane build. Coverage includes all policies, windows/chunk limits, real deadline/demand/abort, partial reads, concurrent GET, reset, exact GPU KV contents and resource reclamation.
  • Qwen3-0.6B / two H20 / TP2: all 15 restored requests matched all 32 greedy output token IDs from three cold references. Whole-task wait and 30 s timeout restored 496/1008/1520 tokens for 512/1024/1536-token prompts. Long timeout used 4/8/12 graphs and did not interrupt. A real 20 ms deadline restored only 32/48/48 tokens, drained, then recomputed the remainder correctly. Immediate best-effort stopped on demand; queued best-effort restored the full prefix. Twelve queue-blocker requests also completed.
  • All 6751 recorded Python/C++/CUDA/header source files matched the tested H20 copies. The final follow-up commit only records the model results. Changed-file Ruff, PR-diff whitespace and companion patch checks passed.

SGLang main refresh (2ef249ef91, based on 4a1b69abc8) with the same FlexKV head:

  • 197 tests and 26 subtests passed, covering current request-attempt handles, owned_kv_len, admission ownership, partial-result contracts, hybrid eviction and backend registration.
  • Repeated the two-H20/TP2 Qwen matrix with sglang-kernel 0.4.7: all 15 restores exactly matched cold outputs. The 20 ms deadline again returned 32/48/48 tokens after drain; long timeout and queued best-effort restored the full prefix.
  • 7,456 source files matched local/H20 copies. SGLang docs build and broken-link validation passed. This refresh is correctness smoke coverage, not a repeat of the historical GLM performance study.

Historical September 10 routing/performance validation:

  • 197 focused Linux tests passed, 9 skipped at the September 10 routing revision, including six routing/precedence cases, runtime/coordinator, native IPC, planner/CPU publication and task lifecycle. Skips are unsupported Python-radix SWA combinations; corresponding C++ cases executed.
  • 560 measured GLM requests and 24 separate cold-reference checks passed, comparing all 32 output token IDs. Each serial arm restored 826 blocks through both REMOTE2H and H2D. Actual remote graph counts were 6/6/8/26 per arm per round; each 16K request used 1/1/2/8 graphs. Post-measurement snapshots verified chunk threads absent for whole-task wait and present for timeout.
  • Sampled CPU throttling, memory-cgroup failures and RDMA error/discard deltas were zero. Changed-file Ruff, whitespace and source checksums were verified. That historical performance run retained its unchanged native extension; the September 15 integration run rebuilt the current extension.
  • Historical coverage remains separate: 456 FlexKV tests / 9 skips and companion SGLang 167 tests / 12 subtests. Earlier timeout/best-effort model canaries covered deadline/demand with inflight drain and complete output-reference checks; this long-timeout matrix intentionally does not interrupt transfers.

Validation boundaries

The feature remains disabled by default. September 22 validates integration correctness with real node-local RDMA transfers and a small TP2 model; it is not a new GLM performance A/B, cross-node RDMA, SWA-model or production soak result. The real GPU fixture still emits the previously observed CUDA IPC producer-exit warning. Five compiled prefetch modules are not a complete release-wheel validation. Dedicated prefetch metrics and broader topology/fault/performance acceptance remain separate work. Historical GLM/V4 results above retain their original source/configuration boundaries.

Retarget the companion patch to PR 31781 at 16780ea0c8 and document the combined e5b00611bd revision. Record 167 CPU tests plus 12 subtests, preserving the separate GPU/model validation boundary.
@XingLiu1
XingLiu1 marked this pull request as ready for review September 11, 2026 06:30
XingLiu added 5 commits September 15, 2026 19:52
Keep indexer deduplication and SWA snapshot accounting together, with constructor regression coverage. Refresh the companion SGLang patch against its latest adaptation base and record CPU/native validation and the insert-after API compatibility finding.
* fix(prefetch): retire released sessions after graph drain

* fix(transfer): isolate worker replica operation IDs
@linhu-nv
linhu-nv self-requested a review September 24, 2026 01:45
@linhu-nv
linhu-nv merged commit b77b125 into main Oct 9, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants