granite-4.1-8b โ€” Q4_K_M GGUF

Quantized version of ibm-granite/granite-4.1-8b
using llama.cpp with the Q4_K_M method (~4.5 bpw, mixed 4-bit K-quants).

How to run

llama-cli -m granite-4.1-8b-Q4_K_M.gguf -cnv -p "You are a helpful assistant"

Or load it directly in LM Studio or Ollama.

Quantization details

Method Bits/weight Size (approx)
Q4_K_M ~4.5 bpw ~5 GB
Downloads last month
40
GGUF
Model size
9B params
Architecture
granite
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for amidblue/granite4.1-8b-Q4_K_M

Quantized
(81)
this model