π Bug Description
Component: qvac-fabric-llm.cpp (v10549.0.0, 57ef02c), Vulkan backend. Filed here because issues are disabled on that repo.
The MoE expert cache (--moe-cache-mib, QVAC-23773) aborts at context creation on Vulkan as soon as the cache is larger than ~2.5 GB. Because --fit sizes the cache automatically to 10 % of the expert weights (common_fit_auto_moe_cache, moe_cache_auto = true in common_params), the default llama-cli / llama-completion / llama-server invocation crashes on any large MoE model on Vulkan; users have to know to pass --moe-cache-mib 0.
common_params_fit_impl: automatic MoE cache = 7345.00 MiB (10% of 73450.00 MiB expert weights)
llama_moe_cache: Vulkan0 MoE cache size = 7344.73 MiB, slots = 2506
ggml/src/ggml-backend.cpp:1568: GGML_ASSERT(!cache_bound || cached_tensor->buffer != NULL) failed
#6 ggml_backend_sched_split_graph
#7 llama_context::graph_reserve
#8 llama_context::resolve_fused_ops
#9 llama_context::sched_reserve
#10 llama_context::llama_context
Root cause (as far as I could trace it)
llama_moe_cache (src/llama-moe-cache.cpp) builds one "bank" tensor per expert projection (ggml_dup_tensor_layout with ne[2] = n_slots + 1) and hands the whole context to ggml_backend_alloc_ctx_tensors_from_buft; the per-layer cached tensors are ggml_view_3d views into those banks. The ggml-vulkan buffer type reports max_size = suballocation_block_size = 1 GiB by default (ggml-vulkan.cpp, overridable with GGML_VK_SUBALLOCATION_BLOCK_SIZE). Once a single bank exceeds that, the allocation is split and the views end up with buffer == NULL, which the scheduler asserts on.
For this model (ffn_down_exps Q5_1 = 600 MiB per layer / 512 experts β 1.17 MiB per slot) the down bank crosses 1 GiB at ~880 slots:
--cpu-moe --moe-cache-mib |
slots |
down bank |
result |
| 2048 |
698 |
0.80 GiB |
runs (5.8 tok/s) |
| 2500 |
~853 |
0.98 GiB |
runs |
| 2700 |
~921 |
1.05 GiB |
assert |
3072 β¦ 8192, and --fit auto (7345 MiB, 2506 slots) |
|
|
assert |
--fit auto with GGML_VK_SUBALLOCATION_BLOCK_SIZE=4294967296 |
2506 |
2.9 GiB |
runs (10.2 tok/s, hit rate 22β86 %) |
CUDA (same build, build-cuda) runs every size, so this is specific to the Vulkan buffer-size limit.
π Steps to Reproduce
# unsloth/Qwen3.8-Flash-Next-GGUF UD-Q4_K_XL (103.7 GiB, 512 experts), any Vulkan GPU with < 100 GB VRAM
llama-completion -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 16384 -t 8 --fit-target 512 -n 32 -no-cnv -p "hi"
# or explicitly:
llama-completion -m ... --cpu-moe --moe-cache-mib 3072 -n 32 -no-cnv -p "hi"
β
Expected
Either allocate the banks in chunks no larger than ggml_backend_buft_get_max_size() (or size n_slots per bank so each stays under it), or have common_fit_auto_moe_cache / llama_moe_cache clamp the budget to what the backend can allocate and log a warning, instead of aborting on defaults.
π» Environment
Intel Core Ultra 7 265KF, 128 GB DDR5-4800, NVIDIA RTX 3090 24 GB, driver 595.91, Vulkan 1.4.329, Ubuntu 26.04. qvac-fabric-llm.cpp v10549.0.0 built with -DGGML_VULKAN=ON (glslc 2026.1); same failure from the @qvac/llm-llamacpp 0.52.0 prebuild (cache β₯ 3 GB). The v10297 build shipped by @qvac/sdk 0.19.1 has no MoE cache and is unaffected.
π Additional context
- Even when it runs, the cache is slower than the static
--fit expert split on this box (Vulkan 10 vs 22 tok/s, CUDA 18β22 vs 24.9 tok/s), so --moe-cache-mib 0 is currently the right default here; that is a separate tuning question.
llama-server with any cache size (including 2048 MiB, with or without --no-warmup) additionally aborts on the first request with ggml-backend.cpp:259 GGML_ASSERT(... "tensor access out of bounds") in ggml_backend_sched_compute_splits β ggml_backend_tensor_set_async; llama-completion and the addon do not. Logged separately if useful.
π Bug Description
Component: qvac-fabric-llm.cpp (v10549.0.0,
57ef02c), Vulkan backend. Filed here because issues are disabled on that repo.The MoE expert cache (
--moe-cache-mib, QVAC-23773) aborts at context creation on Vulkan as soon as the cache is larger than ~2.5 GB. Because--fitsizes the cache automatically to 10 % of the expert weights (common_fit_auto_moe_cache,moe_cache_auto = trueincommon_params), the defaultllama-cli/llama-completion/llama-serverinvocation crashes on any large MoE model on Vulkan; users have to know to pass--moe-cache-mib 0.Root cause (as far as I could trace it)
llama_moe_cache(src/llama-moe-cache.cpp) builds one "bank" tensor per expert projection (ggml_dup_tensor_layoutwithne[2] = n_slots + 1) and hands the whole context toggml_backend_alloc_ctx_tensors_from_buft; the per-layer cached tensors areggml_view_3dviews into those banks. The ggml-vulkan buffer type reportsmax_size = suballocation_block_size= 1 GiB by default (ggml-vulkan.cpp, overridable withGGML_VK_SUBALLOCATION_BLOCK_SIZE). Once a single bank exceeds that, the allocation is split and the views end up withbuffer == NULL, which the scheduler asserts on.For this model (
ffn_down_expsQ5_1 = 600 MiB per layer / 512 experts β 1.17 MiB per slot) the down bank crosses 1 GiB at ~880 slots:--cpu-moe --moe-cache-mib--fitauto (7345 MiB, 2506 slots)--fitauto withGGML_VK_SUBALLOCATION_BLOCK_SIZE=4294967296CUDA (same build,
build-cuda) runs every size, so this is specific to the Vulkan buffer-size limit.π Steps to Reproduce
β Expected
Either allocate the banks in chunks no larger than
ggml_backend_buft_get_max_size()(or sizen_slotsper bank so each stays under it), or havecommon_fit_auto_moe_cache/llama_moe_cacheclamp the budget to what the backend can allocate and log a warning, instead of aborting on defaults.π» Environment
Intel Core Ultra 7 265KF, 128 GB DDR5-4800, NVIDIA RTX 3090 24 GB, driver 595.91, Vulkan 1.4.329, Ubuntu 26.04. qvac-fabric-llm.cpp v10549.0.0 built with
-DGGML_VULKAN=ON(glslc 2026.1); same failure from the@qvac/llm-llamacpp0.52.0 prebuild (cache β₯ 3 GB). The v10297 build shipped by@qvac/sdk0.19.1 has no MoE cache and is unaffected.π Additional context
--fitexpert split on this box (Vulkan 10 vs 22 tok/s, CUDA 18β22 vs 24.9 tok/s), so--moe-cache-mib 0is currently the right default here; that is a separate tuning question.llama-serverwith any cache size (including 2048 MiB, with or without--no-warmup) additionally aborts on the first request withggml-backend.cpp:259 GGML_ASSERT(... "tensor access out of bounds")inggml_backend_sched_compute_splits β ggml_backend_tensor_set_async;llama-completionand the addon do not. Logged separately if useful.