The reference system is a 128 GB Strix Halo with Radeon 8060S (gfx1151),
such as the Framework Desktop. The ROCm build uses the standard binary names
and selects the ROCm backend by default.
For a container setup, see the maintained ROCm toolbox. It can also be managed with AI Toolbox Cockpit.
For a native Ubuntu build you need HIP, hipBLAS, hipBLASLt, rocBLAS, rocWMMA, and hipCUB development files. The Ubuntu 26.04 setup used these packages:
sudo apt-get update
sudo apt-get install -y hipcc rocminfo rocm-smi \
libamdhip64-dev libhipblas-dev libhipblaslt-dev librocblas-dev \
librocwmma-dev libhipcub-dev
sudo usermod -aG render,video "$USER"Log out and back in after changing groups. rocminfo must report gfx1151
and be able to open /dev/kfd before DwarfStar can run.
Some packaged rocWMMA headers omit rocwmma/internal/. If compilation fails
there, install the complete headers matching your ROCm installation, or use
the container. Do not mix header versions as a general workaround.
Check the memory pool reported by rocminfo. Some 128 GB configurations expose
only about 62 GB to the GPU, which is insufficient for resident Flash Q2 plus
runtime buffers. Firmware and kernel GTT/TTM settings control this limit.
The native reference setup used these memory parameters:
amdgpu.gttsize=126976 ttm.pages_limit=32505856 ttm.page_pool_size=32505856
They are a system-specific starting point, not an allocation budget for DwarfStar. Preserve existing boot options and consult your kernel's settings before changing them. Keep RAM available for the OS even when the GPU can address most of it. Do not disable the IOMMU merely to copy another host's configuration; doing so changes device isolation.
make strix-halo
./download_model.sh ds4f-q2
./ds4 --rocmmake rocm is an alias. Use the current 0731 Q2 download for a first run;
larger mixed and Q4 models have substantially higher memory requirements.
Flash's ROCm resident and pipeline paths should not be confused with the GLM
SSD-streaming path.
The reference Q2 setup uses SSD streaming to leave room for its graph and KV state. Begin with automatic cache sizing and a small context:
./download_model.sh glm53-q2
./ds4 --rocm -m gguf/GLM-5.3-Flash-Q2.gguf \
--ssd-streaming --ctx 4096GLM 5.2 also supports ROCm streaming. Full-model GLM 5.2 inference requires it; distributed layer slices can be resident. See SSD streaming before adjusting the cache budget.
Both GLM 5.3 Flash and DeepSeek Flash Vision Experimental support images on
ROCm. Add the matching encoder with --vision FILE, as described in
models and vision.
For a model-free routed-kernel check, use make test-mxfp4-rocm.
Full-model validation is described in testing.