Repository navigation
feat: add chunked prefetch policies with stop-and-drain lifecycle - #291
Merged
Merged
Conversation
3 of 5 tasks
Retarget the companion patch to PR 31781 at 16780ea0c8 and document the combined e5b00611bd revision. Record 167 CPU tests plus 12 subtests, preserving the separate GPU/model validation boundary.
This was referenced Sep 8, 2026
added 6 commits
September 8, 2026 16:21
Point SGLang integration at XingLiu1/sglang#6, targeting agent/flexkv-dsv4-main for review. The adaptation branch was restored; the review head has the same tested source tree.
XingLiu1
marked this pull request as ready for review
September 11, 2026 06:30
added 5 commits
September 15, 2026 19:52
Keep indexer deduplication and SWA snapshot accounting together, with constructor regression coverage. Refresh the companion SGLang patch against its latest adaptation base and record CPU/native validation and the insert-after API compatibility finding.
* fix(prefetch): retire released sessions after graph drain * fix(transfer): isolate worker replica operation IDs
linhu-nv
self-requested a review
September 24, 2026 01:45
linhu-nv
approved these changes
Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem and behavior
Existing prefetch submits a complete remote load before waiting, so timeout or scheduler demand cannot prevent unnecessary later work. Add opt-in chunked sessions for
timeoutandbest_effort. Stopping seals new submissions; already claimed graphs, including sender-queue and IPC work, drain before returning a protected continuous CPU prefix. Foreground held GET acquires its own reference before the prefetch lease is released.SGLang's effective
wait_completepolicy retains the original whole-taskprefetch_asyncpath and disables the chunk runtime before KVManager creation. It does not create chunk control/sender threads. Explicit FlexKV policy configuration takes precedence over the SGLang argument. The low-level explicit session API retains itswait_completecompatibility; an external KVServer retains its own runtime configuration.A control thread owns task/session state and completions, with separate metadata-query and graph-sender threads. Full+SWA/state publication advances only to a complete checkpoint. The worker, kernels and transfer protocol are unchanged. The feature is disabled by default and restricted to node-local Mooncake-to-CPU prefetch.
Configuration and integration
chunk_max_blocksis the only chunk-size option. Remove the older experimentalchunk_target_bytes; byte accounting and checkpoint constraints still apply.2258b8735fintoagent/flexkv-dsv4-main. Upstream SGLang #31781 now includes it and the main refresh at2ef249ef91. The existing connector patch remains unchanged from main; the incremental prefetch patch is excluded from this PR. FlexKV ownsFlexKVConnector/FlexKVComm; SGLang imports the connector and owns cache adapters and scheduler lifecycle hooks. Use the paired source branches directly. FlexKV base:738ddc141a, including insert-after use insert after for all transfer type #266.Performance
Idle control polling previously scanned tasks and retained results every 2 ms without active I/O. It now waits for commands or retained-result TTL; active prefetch and unfinished asynchronous PUTs retain completion polling, and runnable windows refill immediately. The subsequent SGLang policy routing change preserves the whole-task wait path.
At source
837d3a4, two reversed-order GLM-5.2-FP8 rounds compared whole-task wait, one-graph timeout, timeout chunk 128 and timeout chunk 32. All timeout budgets were 60 s with zero deadlines and identical full restores. Hardware/configuration: 8 H20 GPUs, TP8, BF16 KV, eager, page 64, CPU 64 GiB, node-local Mooncake RDMA 256 GiB, window 2. Client C8/C32 used a server admission cap of four.Values are median within-round differences. One-graph timeout separates framework/ownership entry cost from segmentation; changing chunk size also changes backend batch shape and does not isolate pure IPC cost. No large regression was observed in this short workload; zero overhead is not statistically established.
16K TTFT in rounds one/two was 863/869 ms for whole-task wait, 843/862 ms for one-graph timeout, 748/760 ms for chunk 128 and 760/786 ms for chunk 32 (two samples per value). Hot batches had no transfer completions. Per-round ranges and the earlier original/pre-fix/fixed A/B are recorded separately in
docs/design/chunked_prefetch_validation.md.Validation
September 24 namespace integration (
df3fa4b): forward the same optional namespace through lookup, Store, whole-task prefetch and chunked-session creation. The existing planner snapshots it and uses it for every chunk's full token hash chain. Advertisesupports_cache_namespacefor the complete connector path; pair with sgl-project/sglang#31781 head55421bdbcd. FlexKV #304 carries the matching main-branch connector API independently.Validation: 88 CPU/native-index connector, session and planner tests passed, with 9 applicability skips for C++-only SWA scenarios under the Python-radix parameter. This includes three new connector namespace-forwarding cases and existing namespace hash separation for both radix implementations. Applicable pre-commit hooks passed. No new GPU/Mooncake performance or model-output run is claimed for this update.
September 23 cleanup (final correction
e58070d): kept the existing connector patch and integration README files identical to main; excluded the incremental prefetch patch from this PR. No patch file is added, modified or deleted in the final PR diff. Updated the design documentation to reference the paired source branches directly. Patch/base equality, documentation links, external connector imports and PR-diff whitespace checks passed. Runtime source files are unchanged; no new GPU/performance run was needed.September 22 insert-after integration (
4b999671c2, documentation at0dff8ea; paired SGLangc6b51c1b5c):738ddc141aand migrated the planner tonum_matched_blocks/last_node. Insert-after keeps unfinished staging outside the tree, so resident matches are valid. The earlier use insert after for all transfer type #266AttributeErrorreproduction now passes.wraparound=False.SGLang main refresh (
2ef249ef91, based on4a1b69abc8) with the same FlexKV head:owned_kv_len, admission ownership, partial-result contracts, hybrid eviction and backend registration.sglang-kernel 0.4.7: all 15 restores exactly matched cold outputs. The 20 ms deadline again returned 32/48/48 tokens after drain; long timeout and queued best-effort restored the full prefix.Historical September 10 routing/performance validation:
Validation boundaries
The feature remains disabled by default. September 22 validates integration correctness with real node-local RDMA transfers and a small TP2 model; it is not a new GLM performance A/B, cross-node RDMA, SWA-model or production soak result. The real GPU fixture still emits the previously observed CUDA IPC producer-exit warning. Five compiled prefetch modules are not a complete release-wheel validation. Dedicated prefetch metrics and broader topology/fault/performance acceptance remain separate work. Historical GLM/V4 results above retain their original source/configuration boundaries.