Downloads · 30 days
0
MBM7/gguf-tensor-offset-aliasing-poc
gguf-tensor-offset-aliasing-poc is a machine learning model from MBM7. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Status: Preparing for Huntr submission Package: gguf (PyPI) — official gguf-py from ggml-org/llama.cpp File / function: gguf/ggufreader.py, GGUFReader.buildtensors() Class: CWE-1284 / CWE-20 (missing validation of a s…
Downloads · 30 days
0
Access
Public
Updated Jul 20, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py9.2 KB · 59%
From the Hugging Face model README
Status: Preparing for Huntr submission
Package: gguf (PyPI) — official gguf-py from ggml-org/llama.cpp
File / function: gguf/gguf_reader.py, GGUFReader._build_tensors()
Class: CWE-1284 / CWE-20 (missing validation of a structural invariant)
Severity: Medium-High — not a crash/DoS, but a silent model-integrity violation: a GGUF file can declare tensors that don't actually contain the data their name/shape/dtype claim.
Each tensor's absolute position in a GGUF file is computed as:
data_offs = int(start_offs + offset_tensor[0])
where offset_tensor is a per-tensor uint64 read directly from the tensor-info section, with no validation that different tensors' offsets are unique or non-overlapping.
Tensor names are checked for duplicates (_build_tensors raises ValueError on a repeated name) — but nothing stops two tensors with different names, shapes, and/or dtypes from pointing at the exact same, or partially overlapping, bytes. The reader accepts this silently.
poc_gguf_tensor_offset_aliasing.py demonstrates two variants (it builds minimal, spec-correct GGUF files by hand, so it has no dependency on any particular writer library):
Two tensors, blk.0.attn_q.weight = [1,2,3,4] and blk.0.attn_k.weight = [9.9,9.9,9.9,9.9], both float32. Patching only the 8-byte offset field of attn_k.weight to alias attn_q.weight's position:
attn_q.weight: offset=192 data=[1. 2. 3. 4.]
attn_k.weight: offset=192 data=[1. 2. 3. 4.] <- should be [9.9, 9.9, 9.9, 9.9]
Both tensors report identical data. attn_k.weight's real declared data is silently unreachable — zero exception, zero warning.
blk.0.big_weight (float32, 8 elements) and blk.0.small_meta (int32, 2 elements). Patching small_meta's offset to alias the first 8 bytes of big_weight:
big_weight (F32): [1.1 2.2 3.3 4.4 5.5 6.6 7.7 8.8]
small_meta (I32): [1066192077 1074580685]
1066192077 and 1074580685 are the exact IEEE-754 bit patterns of 1.1 and 2.2 reinterpreted as raw int32. A tensor of any declared shape and dtype can be carved out of any byte range in the file, regardless of what other tensor(s) claim that same range.
pip install gguf numpy
python poc_gguf_tensor_offset_aliasing.py
A GGUF file can declare many distinct-looking weight tensors — correct names, correct shapes, passes casual inspection — that actually all alias onto a small set of real bytes. This can be used to:
_build_tensors() raises ValueError('Found duplicated tensor with name ...').offset_tensor is unsigned and added to start_offs (the data region's start); the minimum reachable position is start_offs itself.reshape() validation (individual out-of-bounds tensors are not the gap here — cross-tensor uniqueness is).Track allocated byte ranges while building the tensor list in _build_tensors(), and raise a clear error if any two tensors' [data_offs, data_offs + n_bytes) ranges overlap.
Distinct from the previously reported "KV Array Field Unbounded Length" finding in the same package (CWE-834/CWE-400, a resource-exhaustion issue) — this is a data-integrity issue (CWE-1284/CWE-20), closer in spirit to the numpy NPZ key-shadowing finding (CWE-706/CWE-345) reported separately, but via byte-offset overlap rather than string-key collision.
Please do not use this PoC against production systems you do not own or have explicit permission to test.