b10142
b10142
View on GitHubView PackagePublished: Jul 27, 2026

Release Notes

mtmd: Add Vision Support for Minimax-M3 (#25113)

  • Add preliminary MiniMax-M3 support

Text-only port that re-uses existing components: MiniMax-M2 style GQA with per-head QK-norm and partial rotary, DeepSeek-V3 style leading-dense and routed/shared experts, and swigluoai activation. Sparse attention is not yet supported (dense fallback); vision tower and MTP heads are dropped.

  • MiniMax-M3 vision tower (mmproj + clip graph)

  • Delete m3_vision_ref.py

  • Update clip.cpp

  • MSA

  • Update constants.py

  • Update minimax.py

  • Cache creation. Working withotu flash attention

  • Added flash attention for sparse layers

  • Decomposed slow cpu OP into GPU + CPU ops. Massive speedup over long ctx

  • Rewrote indexer op to be cuda native. Modified flash attention to match per group block picking

  • Implement sparse attention calc out of stock ops.

  • Fix a cache allocation and cont issue

  • Fixed -fa auto crash, flagged debug spots

  • Delete vocab.json

  • Delete model.safetensors.index.json

  • Delete generation_config.json

  • Delete Minimax directory

  • Handled multi stream case to fall back on Dense Attention

  • Development scaffolding cleanup. No functional change to the decode or 4-way paths. Full debug harness remains at <8136a9c68ed7a5eb009aa67bba3fda8062f4648f> for reproducing the selection-parity validation.

  • Remove redundant comment from minimax-m3.cpp

  • Changed 3 Gelu Ops for vision into Gelu_erf ops

  • Assert that n_kv is multiple of 128

  • Rename MSA index tensors to indexer convention

Note: All GGUFs generated before this change will need to be regenerated.

  • Fix incorrect Assert

  • Review driven changes (#3)

  • Remove comment from conversion minimax.py

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Remove whitespaces from constants.py

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Tighten comment in minimax.py

Co-authored-by: Sigbjørn Skjæret [email protected]

  • inherit MiniMax-M3 from MiniMax-M2

  • drop dead text_config fallbacks

  • Add indexer writer methods

  • Reuse LLM_FFN_SWIGLU_OAI_MOE

  • Remove duplicate indexer setters, add only block_size/local_blocks, follow value naming convention

  • Fix conversion error /gguf_writer.py

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Update gguf-py/gguf/gguf_writer.py

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Update gguf-py/gguf/tensor_mapping.py

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Update conversion/minimax.py

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Update conversion/minimax.py

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Remove whitespace in src/llama-kv-cache.cpp

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Remove Whitespace in Update src/llama-model.h

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Remove whitespace in src/llama-hparams.h

Co-authored-by: Sigbjørn Skjæret [email protected]

  • Update minimax_m3.cpp

Rewrite code comment based on feedback and to better reflect the actual architecture, and reuse existing build_vit

  • Rename minimax_m3.cpp to minimax-m3.cpp

  • Update CMakeLists.txt

  • Remove debug code from clip.cpp

  • Update clip.cpp

  • Update comments in tools/mtmd/models/minimax-m3.cpp

  • Permute Q/K at conversion, drop precomputed sin/cos

  • Log cache size on launch, block ctx shift, support prompt caching

Log indexer cache size on launch

Disallow ctx shift

Support prompt caching

  • Update minimax-m3.cpp

  • Optimize implementation, add multi stream support.

Fully rewrote minimax-m3.cpp for speed and buffer size gains:

Unified the 4-way + decode, 1 FA call per layer instead of 4, with the groups mapped onto ne[3]

Custom CPU op now emits block-level mask, expanded on GPU, which causes CPU to GPU transfer to shrinks at prefill

Decode: ~25 nodes/layer vs ~50, no per-group concats/conts

Unified selection semantics, so both regimes rank bs + local bias (position-anchored local force), which means prefill/decode can no longer disagree on selection

can_reuse on the MSA bias input. Graph reuse at decode restored (was rebuilding the full graph every token)

In-place mask adds, shrinking compute buffer ~6.8 to ~4.2 GiB at ub2048/62k

Multi-stream: MSA now runs with -np N when kv_unified=false. Decode stays batched across streams (still 1 FA call), prefill loops per stream. dense fallback only for --kv-unified + multi-seq

Measured effect on expert offload bound setup: decode 6.2(4WAY)–7.15(MSA_decode) -> 7.7~7.8 t/s, flat from 5k to 60k+. prefill around 10% faster. buffer about 20% smaller, multi-user support.

  • set default cache type to F32

  • Fix potential DSA double indexer cache allocation bug, only allocate in-cache k_idx for archs that opt in

  • remove F16 downcasts in MSA attention, force F32 indexer score accum

  • Add Minimax eos to llama vocab

  • Guard edge case where idx cache can become stale after a tail trim

  • Update llama-kv-cache.h

  • Update llama-kv-cache.cpp

  • Update llama-kv-cache.cpp

  • Update llama-kv-cache.h

  • Change resize Pad to none, resize alg to Bicubic Pillow

  • Review driven changes

  • Update llama-kv-cache.cpp

  • rm unrotated pos_t

  • fused rope w + pad

  • rename merge --> merger for consistency

  • add review skill for mtmd

  • graph should use hparams n_merge

  • fix lint


Co-authored-by: Daniel Han [email protected] Co-authored-by: Sigbjørn Skjæret [email protected] Co-authored-by: Xuan Son Nguyen [email protected]

Website:

macOS/iOS:

Linux:

Android:

Windows:

openEuler:

  • DISABLED
  • openEuler x86 (310p)
  • openEuler x86 (910b, ACL Graph)
  • openEuler aarch64 (310p)
  • openEuler aarch64 (910b, ACL Graph)

UI: