-
Notifications
You must be signed in to change notification settings - Fork 2.1k
Pull requests: antirez/ds4
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Metal: optimize Qwen3.8 kernels, MTP state and SSD MoE scheduling (Up tp 40%)
#1056
opened Sep 15, 2026 by
GiorgioOppo
Loading…
server: advertise the loaded GLM model id in /v1/models
#1055
opened Sep 15, 2026 by
spencergilbert
Loading…
Fix GLM-5.3 Metal decode masking for padded selections
#1051
opened Sep 14, 2026 by
IngeniousIdiocy
Loading…
Qwen tool-call error: say where the tool name belongs
#1050
opened Sep 14, 2026 by
aovestdipaperino
Loading…
Remove unsupported Hugging Face download resume promise
#1048
opened Sep 14, 2026 by
pminervini
Loading…
Metal: add SSD expert streaming for Qwen3.8 Flash Next
#1047
opened Sep 14, 2026 by
GiorgioOppo
Loading…
Improve Metal router softplus accuracy for small logits
#1044
opened Sep 14, 2026 by
emilianbold
Loading…
metal: Q4_K group-6 expert table by default for DeepSeek V4.1 Flash (+2.3% decode on M3 Ultra, greedy-identical)
#1043
opened Sep 14, 2026 by
adriangalilea
Loading…
v41: fuse the single-box decode glue (HC, MoE, attention): +20% on M3 Ultra on top of #1041, bit-exact
#1042
opened Sep 14, 2026 by
adriangalilea
Loading…
v41: queue single-box decode layers and commit each without waiting (+37% decode on M3 Ultra, bit-exact)
#1041
opened Sep 14, 2026 by
adriangalilea
Loading…
ROCm: add DeepSeek V4.1 Flash support for Strix Halo (gfx1151)
#1036
opened Sep 13, 2026 by
kyuz0
Contributor
Loading…
engram: parallelize DeepSeek V4.1 Flash decode reads on macOS
#1035
opened Sep 13, 2026 by
Dango233
Loading…
metal: reduce decode synchronization for DeepSeek V4.1 Flash SSD streaming
#1034
opened Sep 13, 2026 by
Dango233
Loading…
metal: add opt-in slab residency for DeepSeek V4.1 Flash (M2 192GB 0.25tk/s -> 13tk/s)
#1033
opened Sep 13, 2026 by
Dango233
Loading…
Fix CUDA long-context smoke test linking
#1030
opened Sep 12, 2026 by
Matthley
Loading…
3 tasks done
rocm: model weights in host RAM on integrated APUs (Strix Halo)
#1028
opened Sep 12, 2026 by
Piega
Loading…
Enable GLM 5.3 tensor parallelism on the ROCm backend
#1024
opened Sep 10, 2026 by
davidcanar
Loading…
server: opt-in auto-reduction of over-budget image histories (DS4_VISION_KEEP_IMAGES)
#1022
opened Sep 10, 2026 by
nazerim
Loading…
server: auto-discard reload-loop checkpoints (frontier-contradicted)
#1021
opened Sep 10, 2026 by
nazerim
Loading…
kvstore: stamp checkpoints with a behavioral tokenizer fingerprint
#1020
opened Sep 10, 2026 by
nazerim
Loading…
Add OdinLink (odl_tb5) TP transport: --transport odl
#1018
opened Sep 10, 2026 by
davidcanar
Loading…
Previous Next
ProTip!
Filter pull requests by the default branch with base:main.