b10164
b10164
View on GitHubView PackagePublished: Jul 28, 2026

Release Notes

ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration (#22675)

  • ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration

  • cuda: added SSD CICD fixes for CUDA / HIP / MUSA / MSVC.

  • ggml-cuda: review comments fixed.

  • ggml-cuda: Fuse M matrix materialization into pre_matmul kernel and enabled test.

  • ggml-cuda: test updates and fixes

  • ggml-cuda: test updates to remove hardcoding of tensor initialise data limits.

  • ggml-cuda: ssd minor review comment fixed.

  • ggml-cuda: ssd minor CICD fixed.

  • CUDA SSD: Fixes correctness by promoting s0_stride_seq to int64_t, improves memory coalescing in ssm_ssd_prepare_dt_kernel, and boosts efficiency by merging B_weighted and C_scaled; also addresses prior review comments.

  • cuda: fix sdata read-write race in prepare_dt fallback scan loop

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI: