Skip to content

[Bug]: qvac-fabric MoE expert cache asserts on Vulkan for caches > ~2.5 GB (bank tensor exceeds 1 GiB suballocation block) β€” --fit auto-sizes it to 10% of experts, so defaults crash on large MoE modelsΒ #4435

Description

@yuranich

πŸ› Bug Description

Component: qvac-fabric-llm.cpp (v10549.0.0, 57ef02c), Vulkan backend. Filed here because issues are disabled on that repo.

The MoE expert cache (--moe-cache-mib, QVAC-23773) aborts at context creation on Vulkan as soon as the cache is larger than ~2.5 GB. Because --fit sizes the cache automatically to 10 % of the expert weights (common_fit_auto_moe_cache, moe_cache_auto = true in common_params), the default llama-cli / llama-completion / llama-server invocation crashes on any large MoE model on Vulkan; users have to know to pass --moe-cache-mib 0.

common_params_fit_impl: automatic MoE cache = 7345.00 MiB (10% of 73450.00 MiB expert weights)
llama_moe_cache:    Vulkan0 MoE cache size =  7344.73 MiB, slots = 2506
ggml/src/ggml-backend.cpp:1568: GGML_ASSERT(!cache_bound || cached_tensor->buffer != NULL) failed
#6  ggml_backend_sched_split_graph
#7  llama_context::graph_reserve
#8  llama_context::resolve_fused_ops
#9  llama_context::sched_reserve
#10 llama_context::llama_context

Root cause (as far as I could trace it)

llama_moe_cache (src/llama-moe-cache.cpp) builds one "bank" tensor per expert projection (ggml_dup_tensor_layout with ne[2] = n_slots + 1) and hands the whole context to ggml_backend_alloc_ctx_tensors_from_buft; the per-layer cached tensors are ggml_view_3d views into those banks. The ggml-vulkan buffer type reports max_size = suballocation_block_size = 1 GiB by default (ggml-vulkan.cpp, overridable with GGML_VK_SUBALLOCATION_BLOCK_SIZE). Once a single bank exceeds that, the allocation is split and the views end up with buffer == NULL, which the scheduler asserts on.

For this model (ffn_down_exps Q5_1 = 600 MiB per layer / 512 experts β‰ˆ 1.17 MiB per slot) the down bank crosses 1 GiB at ~880 slots:

--cpu-moe --moe-cache-mib slots down bank result
2048 698 0.80 GiB runs (5.8 tok/s)
2500 ~853 0.98 GiB runs
2700 ~921 1.05 GiB assert
3072 … 8192, and --fit auto (7345 MiB, 2506 slots) assert
--fit auto with GGML_VK_SUBALLOCATION_BLOCK_SIZE=4294967296 2506 2.9 GiB runs (10.2 tok/s, hit rate 22–86 %)

CUDA (same build, build-cuda) runs every size, so this is specific to the Vulkan buffer-size limit.

πŸ”„ Steps to Reproduce

# unsloth/Qwen3.8-Flash-Next-GGUF UD-Q4_K_XL (103.7 GiB, 512 experts), any Vulkan GPU with < 100 GB VRAM
llama-completion -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 16384 -t 8 --fit-target 512 -n 32 -no-cnv -p "hi"
# or explicitly:
llama-completion -m ... --cpu-moe --moe-cache-mib 3072 -n 32 -no-cnv -p "hi"

βœ… Expected

Either allocate the banks in chunks no larger than ggml_backend_buft_get_max_size() (or size n_slots per bank so each stays under it), or have common_fit_auto_moe_cache / llama_moe_cache clamp the budget to what the backend can allocate and log a warning, instead of aborting on defaults.

πŸ’» Environment

Intel Core Ultra 7 265KF, 128 GB DDR5-4800, NVIDIA RTX 3090 24 GB, driver 595.91, Vulkan 1.4.329, Ubuntu 26.04. qvac-fabric-llm.cpp v10549.0.0 built with -DGGML_VULKAN=ON (glslc 2026.1); same failure from the @qvac/llm-llamacpp 0.52.0 prebuild (cache β‰₯ 3 GB). The v10297 build shipped by @qvac/sdk 0.19.1 has no MoE cache and is unaffected.

πŸ“Ž Additional context

  • Even when it runs, the cache is slower than the static --fit expert split on this box (Vulkan 10 vs 22 tok/s, CUDA 18–22 vs 24.9 tok/s), so --moe-cache-mib 0 is currently the right default here; that is a separate tuning question.
  • llama-server with any cache size (including 2048 MiB, with or without --no-warmup) additionally aborts on the first request with ggml-backend.cpp:259 GGML_ASSERT(... "tensor access out of bounds") in ggml_backend_sched_compute_splits β†’ ggml_backend_tensor_set_async; llama-completion and the addon do not. Logged separately if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    NLPllm and embedbugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions