Release Notes
mtmd: Add Vision Support for Minimax-M3 (#25113)
- Add preliminary MiniMax-M3 support
Text-only port that re-uses existing components: MiniMax-M2 style GQA with per-head QK-norm and partial rotary, DeepSeek-V3 style leading-dense and routed/shared experts, and swigluoai activation. Sparse attention is not yet supported (dense fallback); vision tower and MTP heads are dropped.
MiniMax-M3 vision tower (mmproj + clip graph)
Delete m3_vision_ref.py
Update clip.cpp
MSA
Update constants.py
Update minimax.py
Cache creation. Working withotu flash attention
Added flash attention for sparse layers
Decomposed slow cpu OP into GPU + CPU ops. Massive speedup over long ctx
Rewrote indexer op to be cuda native. Modified flash attention to match per group block picking
Implement sparse attention calc out of stock ops.
Fix a cache allocation and cont issue
Fixed -fa auto crash, flagged debug spots
Delete vocab.json
Delete model.safetensors.index.json
Delete generation_config.json
Delete Minimax directory
Handled multi stream case to fall back on Dense Attention
Development scaffolding cleanup. No functional change to the decode or 4-way paths. Full debug harness remains at <8136a9c68ed7a5eb009aa67bba3fda8062f4648f> for reproducing the selection-parity validation.
Remove redundant comment from minimax-m3.cpp
Changed 3 Gelu Ops for vision into Gelu_erf ops
Assert that n_kv is multiple of 128
Rename MSA index tensors to indexer convention
Note: All GGUFs generated before this change will need to be regenerated.
Fix incorrect Assert
Review driven changes (#3)
Remove comment from conversion minimax.py
Co-authored-by: Sigbjørn Skjæret [email protected]
- Remove whitespaces from constants.py
Co-authored-by: Sigbjørn Skjæret [email protected]
- Tighten comment in minimax.py
Co-authored-by: Sigbjørn Skjæret [email protected]
inherit MiniMax-M3 from MiniMax-M2
drop dead text_config fallbacks
Add indexer writer methods
Reuse LLM_FFN_SWIGLU_OAI_MOE
Remove duplicate indexer setters, add only block_size/local_blocks, follow value naming convention
Fix conversion error /gguf_writer.py
Co-authored-by: Sigbjørn Skjæret [email protected]
- Update gguf-py/gguf/gguf_writer.py
Co-authored-by: Sigbjørn Skjæret [email protected]
- Update gguf-py/gguf/tensor_mapping.py
Co-authored-by: Sigbjørn Skjæret [email protected]
- Update conversion/minimax.py
Co-authored-by: Sigbjørn Skjæret [email protected]
- Update conversion/minimax.py
Co-authored-by: Sigbjørn Skjæret [email protected]
- Remove whitespace in src/llama-kv-cache.cpp
Co-authored-by: Sigbjørn Skjæret [email protected]
- Remove Whitespace in Update src/llama-model.h
Co-authored-by: Sigbjørn Skjæret [email protected]
- Remove whitespace in src/llama-hparams.h
Co-authored-by: Sigbjørn Skjæret [email protected]
- Update minimax_m3.cpp
Rewrite code comment based on feedback and to better reflect the actual architecture, and reuse existing build_vit
Rename minimax_m3.cpp to minimax-m3.cpp
Update CMakeLists.txt
Remove debug code from clip.cpp
Update clip.cpp
Update comments in tools/mtmd/models/minimax-m3.cpp
Permute Q/K at conversion, drop precomputed sin/cos
Log cache size on launch, block ctx shift, support prompt caching
Log indexer cache size on launch
Disallow ctx shift
Support prompt caching
Update minimax-m3.cpp
Optimize implementation, add multi stream support.
Fully rewrote minimax-m3.cpp for speed and buffer size gains:
Unified the 4-way + decode, 1 FA call per layer instead of 4, with the groups mapped onto ne[3]
Custom CPU op now emits block-level mask, expanded on GPU, which causes CPU to GPU transfer to shrinks at prefill
Decode: ~25 nodes/layer vs ~50, no per-group concats/conts
Unified selection semantics, so both regimes rank bs + local bias (position-anchored local force), which means prefill/decode can no longer disagree on selection
can_reuse on the MSA bias input. Graph reuse at decode restored (was rebuilding the full graph every token)
In-place mask adds, shrinking compute buffer ~6.8 to ~4.2 GiB at ub2048/62k
Multi-stream: MSA now runs with -np N when kv_unified=false. Decode stays batched across streams (still 1 FA call), prefill loops per stream. dense fallback only for --kv-unified + multi-seq
Measured effect on expert offload bound setup: decode 6.2(4WAY)–7.15(MSA_decode) -> 7.7~7.8 t/s, flat from 5k to 60k+. prefill around 10% faster. buffer about 20% smaller, multi-user support.
set default cache type to F32
Fix potential DSA double indexer cache allocation bug, only allocate in-cache k_idx for archs that opt in
remove F16 downcasts in MSA attention, force F32 indexer score accum
Add Minimax eos to llama vocab
Guard edge case where idx cache can become stale after a tail trim
Update llama-kv-cache.h
Update llama-kv-cache.cpp
Update llama-kv-cache.cpp
Update llama-kv-cache.h
Change resize Pad to none, resize alg to Bicubic Pillow
Review driven changes
Update llama-kv-cache.cpp
rm unrotated pos_t
fused rope w + pad
rename merge --> merger for consistency
add review skill for mtmd
graph should use hparams n_merge
fix lint
Co-authored-by: Daniel Han [email protected] Co-authored-by: Sigbjørn Skjæret [email protected] Co-authored-by: Xuan Son Nguyen [email protected]
Website:
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (ROCm 7.2)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
Android:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.3 DLLs
- Windows x64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (HIP)
openEuler:
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI: