Instructions to use Hironabe333/gguf-kv-count-header-mismatch-poc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Hironabe333/gguf-kv-count-header-mismatch-poc with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Hironabe333/gguf-kv-count-header-mismatch-poc # Run inference directly in the terminal: llama cli -hf Hironabe333/gguf-kv-count-header-mismatch-poc
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Hironabe333/gguf-kv-count-header-mismatch-poc # Run inference directly in the terminal: llama cli -hf Hironabe333/gguf-kv-count-header-mismatch-poc
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Hironabe333/gguf-kv-count-header-mismatch-poc # Run inference directly in the terminal: ./llama-cli -hf Hironabe333/gguf-kv-count-header-mismatch-poc
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Hironabe333/gguf-kv-count-header-mismatch-poc # Run inference directly in the terminal: ./build/bin/llama-cli -hf Hironabe333/gguf-kv-count-header-mismatch-poc
Use Docker
docker model run hf.co/Hironabe333/gguf-kv-count-header-mismatch-poc
- LM Studio
- Jan
- Ollama
How to use Hironabe333/gguf-kv-count-header-mismatch-poc with Ollama:
ollama run hf.co/Hironabe333/gguf-kv-count-header-mismatch-poc
- Unsloth Desktop
- Docker Model Runner
How to use Hironabe333/gguf-kv-count-header-mismatch-poc with Docker Model Runner:
docker model run hf.co/Hironabe333/gguf-kv-count-header-mismatch-poc
- Lemonade
How to use Hironabe333/gguf-kv-count-header-mismatch-poc with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Hironabe333/gguf-kv-count-header-mismatch-poc
Run and chat with the model
lemonade run user.gguf-kv-count-header-mismatch-poc-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
GGUF kv_count Header Field Mismatch β PoC
Target: llama-cpp-python (llama.cpp C runtime)
Class: Parser/loader schema invariant β structural binary header field mismatch
Root: kv_count (uint64_le, offset 16 in GGUF header) trusted by C runtime without bounds validation
Impact: Deterministic SIGABRT crash (exit 134) on the tested runtime (llama-cpp-python 0.2.90, aarch64 Linux) when loading the provided PoC GGUF artifacts with mismatched kv_count
Version tested: llama-cpp-python 0.2.90 (aarch64 Linux). Findings scoped to tested version only.
Vulnerability
The GGUF binary format header contains a kv_count field at byte offset 16 (uint64_le) that specifies the number of key-value metadata entries in the file body. The llama.cpp C runtime (gguf_init_from_file) reads this field and iterates exactly kv_count times over the file body, without validating that kv_count matches the actual number of KV entries present.
When kv_count is set to a value larger than the actual number of entries (inflation), the runtime reads past the end of the KV section into tensor data, eventually calling ggml_calloc(0) β ggml_abort() (SIGABRT, ggml.c:366).
When kv_count is set to a value smaller than the actual number (deflation), the runtime stops reading KVs too early, then attempts to parse tensor info from a byte offset that is still inside the KV section. The resulting garbage n_dims value fails GGML_ASSERT(info->n_dims <= GGML_MAX_DIMS) (SIGABRT, ggml.c:21301).
Both paths produce a deterministic SIGABRT crash in the tested llama.cpp-based C runtime (llama-cpp-python 0.2.90) when loading the provided PoC artifacts.
Header Layout
Offset 0: magic (4 bytes, "GGUF")
Offset 4: version (uint32_le)
Offset 8: tensor_count (uint64_le)
Offset 16: kv_count (uint64_le) β VULNERABLE FIELD
Offset 24: KV metadata entries begin
Invariant Differential
The Python GGUFReader validator (gguf-py) reads kv_count and loops that many times in _build_fields(offs, kv_count) β but with no bounds check. When kv_count exceeds actual entries, GGUFReader raises a clean IndexError. When kv_count is below actual, it reads fewer entries and succeeds. No crash in the Python validator.
The C runtime (gguf_init_from_file) produces the same input β SIGABRT in both cases.
This is a parser/loader schema invariant: the same structurally mutated file produces divergent behavior β clean exception vs. process crash β between the Python and C components.
Files
| File | Description |
|---|---|
baseline.gguf |
Well-formed GGUF v3: kv_count=15 (correct) |
mut_kv_inflated.gguf |
kv_count patched 15β18 (inflation); causes SIGABRT |
mut_kv_deflated.gguf |
kv_count patched 15β10 (deflation); causes SIGABRT |
reproduce.py |
Full reproduction: build GGUF, patch, test both paths |
checker.py |
Python-only validator check (no llama-cpp-python needed) |
mutation_generator.py |
Standalone patcher: set kv_count to any value |
t0_evidence.json |
T0 test results (version 0.2.90, aarch64 Linux) |
prior_art_clearance.json |
Prior art check: all 13+ llama.cpp CVEs cleared |
distinctness_table.json |
Distinctness vs. ROUND_AI100, ROUND_AI154, TALOS-2024-1913 |
hash_matrix.json |
SHA256 hashes + expected behavior per file |
SHA256SUMS.txt |
Checksums for all artifact files |
Reproduction
Requirements
pip install llama-cpp-python==0.2.90 gguf
For aarch64 Linux (pre-built wheel tested):
pip install llama-cpp-python==0.2.90 --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cpu
pip install gguf
Run
python3 reproduce.py
Expected Output
--- baseline ---
GGUFReader: PASS (real_kvs=15, tensors=2)
llama_cpp : PYTHON_EXCEPTION (exit 1) β no model arch, controlled
--- kv_inflated ---
GGUFReader: FAIL β IndexError: index 0 is out of bounds ...
llama_cpp : SIGABRT_CRASH β exit 134
> ggml_calloc: failed to allocate 0.00 MB
> ggml.c:366: GGML_ASSERT ...
--- kv_deflated ---
GGUFReader: FAIL β IndexError: index 0 is out of bounds ...
llama_cpp : SIGABRT_CRASH β exit 134
> GGML_ASSERT(info->n_dims <= GGML_MAX_DIMS) failed
> ggml.c:21301
Python-only check (no llama-cpp-python)
pip install gguf
python3 checker.py
Mutate any GGUF file
# Inflate: set kv_count to 18 (any GGUF with fewer than 18 KVs)
python3 mutation_generator.py baseline.gguf my_inflated.gguf 18
# Deflate: set kv_count to 5 (any GGUF with more than 5 KVs)
python3 mutation_generator.py baseline.gguf my_deflated.gguf 5
Crash Details
Inflation path (kv_count 15β18)
The runtime iterates 18 times over KV entries. After reading all 15 real entries, the next 3 reads are from tensor info data. The tensor type bytes happen to produce a string type (0x08), causing the runtime to call ggml_calloc(0, string_length) where the length field overflows or is zero:
ggml_calloc: failed to allocate 0.00 MB
ggml.c:366: GGML_ASSERT(size >= 0) fatal error
Aborted (core dumped)
Deflation path (kv_count 15β10)
The runtime reads only 10 KV entries, then begins parsing tensor info from the byte position immediately after the 10th KV. This offset is still in the middle of the KV section. The 4 bytes at that position are interpreted as n_dims (uint32), yielding a garbage value far exceeding GGML_MAX_DIMS (4):
GGML_ASSERT: ggml.c:21301: info->n_dims <= GGML_MAX_DIMS
Aborted (core dumped)
Fix
gguf_init_from_file() should validate that kv_count does not exceed the remaining file bytes divided by minimum KV entry size before iterating, and should bounds-check each KV entry read against the file size. The GGUFReader Python implementation should similarly add bounds checking in _build_fields().
SHA256 Hashes
84d4c0cfe55b75e2b319d3f240641dc0569cba2e54acd4524475e20fd2b7873a baseline.gguf
cb674aaeab10dc3759edf38f991071615349b4814033f5661b78cac7d8819cb2 mut_kv_inflated.gguf
08fb22d09f65e5a944824f237657e6bbcdff1f5dfc7cacd0d69af5e945a488dc mut_kv_deflated.gguf
- Downloads last month
- -
We're not able to determine the quantization variants.