Skip to content

GLM-5.3-Flash FP8 Optimized Inference Support #1015

Description

@zzyu17

As stated in docs/MODELS.md (line 46):

| `glm53-fp8` | 305 GiB | Packaged native weights only; inference not implemented

inference for GLM-5.3-Flash FP8 model is not yet supported in the current ds4 engine.

I tried simply patching ds4.c by routing GLM Q8_0 experts through the existing generic Metal ID kernels and adding Q8_0 coverage to the F32 reference path (see https://github.com/zzyu17/ds4/commit/61b0d30c30d05cd118335f60e044a8f772225c01 for details), which gives the below results for a GLM-5.3-Flash Q8_0 GGUF on M3 Ultra 512G.

Image

The performance is quite good, with ~390 token/s prefill and ~21 token/s generation. However, an optimized (Metal) inference kernel for GLM-5.3-Flash FP8 model should perform even better, compared to the generic Metal kernel patch.

So, looking for the optimized (Metal) inference support for GLM-5.3-Flash FP8 model, eagerly!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions