MiMo-V2.6-Flash-MOPD

Run with https://llama.app

llama serve -hf ggml-org/MiMo-V2.6-Flash-MOPD-GGUF

Source models

Notes

  • The MXFP4 output keeps the routed experts at their native MXFP4 precision.
  • The Q2_K output keeps the expert down projections at MXFP4, and quantizes the gate/up projections to Q2_K.
  • Includes MTP sidecars (Q4_0 and Q8_0) for speculative decoding (--mtp).
  • Includes a DFlash drafter sidecar (BF16 and Q8_0) for speculative decoding, converted from the dflash/ subdirectory of the source repo.
  • Includes a Q8_0 mmproj for the vision and audio encoders.
  • The Q2_K expert gate/up tensors are calibrated with the imatrix from https://huggingface.co/AesSedai/MiMo-V2.6-Flash-MOPD-GGUF

This model is automatically converted using https://github.com/ggml-org/convert

Downloads last month
4,759
GGUF
Model size
309B params
Architecture
mimo2
Hardware compatibility
Log In to add your hardware

2-bit

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ggml-org/MiMo-V2.6-Flash-MOPD-GGUF

Quantized
(23)
this model

Space using ggml-org/MiMo-V2.6-Flash-MOPD-GGUF 1