Downloads · 30 days
0
ericblackgachara/numpy-npy-poc
numpy-npy-poc is a machine learning model from ericblackgachara. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated May 31, 2026
Repo size
146 KB
Likes
1
Public
Click a slice to open those files.
.pdf146 KB · 91%
From the Hugging Face model README
Loading a crafted 1,214-byte NPY file causes np.load() to raise tokenize.TokenError
instead of the documented ValueError. The exception escapes uncaught from all numpy
handlers, crashing any application that catches only ValueError per numpy's documented API.
| Field | Value |
|---|---|
| Repository | numpy/numpy |
| Affected file | numpy/lib/_format_impl.py |
| Affected function | _read_array_header() line 661 |
| Confirmed version | 2.3.5 (installed), main branch commit b6e8ad8 |
| Python version | 3.12+ (C tokenizer enforces nesting limit) |
| NPY format | v1.0 and v2.0 (NOT v3.0) |
| Platform | huntr.com |
_read_array_header() calls _filter_header() when ast.literal_eval() raises SyntaxError.
For headers with 100+ nested parentheses, tokenize.generate_tokens() inside _filter_header()
raises tokenize.TokenError — which is not a subclass of ValueError and is never caught.
# numpy/lib/_format_impl.py line 661
except SyntaxError as e:
if version <= (2, 0):
header = _filter_header(header) # <-- raises TokenError, not caught!
crafted NPY v1.0 (100-level nested descr)
→ np.load()
→ format.read_array()
→ _read_array_header()
→ ast.literal_eval() raises SyntaxError
→ _filter_header()
→ tokenize.generate_tokens() raises TokenError
→ [UNCAUGHT] propagates out of np.load()
# 1. Install numpy (tested on 2.3.5 + Python 3.13)
pip install numpy
# 2. Run PoC
python3 poc_numpy.py
# Expected output:
# [+] CONFIRMED: tokenize.TokenError raised!
# [+] CONFIRMED: TokenError ESCAPED ValueError handler!
# depth=99: ValueError raised (safe)
# depth=100: [+] TokenError raised (VULNERABLE)
| Property | Value |
|---|---|
| Malicious file size | 1,214 bytes |
| Header length | 1,147 chars (limit: 10,000) |
| Nesting depth | 100 levels |
| User interaction required | Load the file (implicit in ML pipelines) |
# In _read_array_header(), wrap _filter_header() call:
from tokenize import TokenError as _TokenError
try:
header = _filter_header(header)
except _TokenError as te:
raise ValueError(
f"Cannot parse header (tokenizer error): {header!r}"
) from te
| File | Purpose |
|---|---|
poc_numpy.py | Working PoC — generates malicious NPY and demonstrates TokenError escape |
report.md | Full huntr-formatted report |
poc-evidence.html | Self-contained HTML evidence page with terminal output |
README.md | This file |