Downloads · 30 days
0
np-cr/tiny_prefill_decode_npir
tiny_prefill_decode_npir is a machine learning model from np-cr. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Tiny autoregressive prefill/decode NPIR fixture for multi-graph GQ tests (NPP02-7270). KV cache is graph I/O (not an internal stateful buffer), matching the USERINPUT/USEROUTPUT/USERINPUTMUTATION form measured on a re…
Downloads · 30 days
0
Access
Public
Updated Jul 22, 2026
Repo size
318 KB
Likes
0
Public
Click a slice to open those files.
.npir306 KB · 94%
From the Hugging Face model README
Tiny autoregressive prefill/decode NPIR fixture for multi-graph GQ tests (NPP02-7270). KV cache is graph I/O (not an internal stateful buffer), matching the USER_INPUT/USER_OUTPUT/USER_INPUT_MUTATION form measured on a real smollm2 torch.export graph_signature. Deterministic (seed 0), 2 layers, reduced dims.
Composition edges (one per K/V tensor, [k0, v0, k1, v1, ...] order):
prefill.index_copy_.output_0 -> decode.k0prefill.index_copy__1.output_0 -> decode.v0prefill.index_copy__2.output_0 -> decode.k1prefill.index_copy__3.output_0 -> decode.v1No decode -> decode edge is declared: the autoregressive feedback (decode's own K/V output feeding its next-step input) is a cycle and is correctly rejected by multi_graph.composition.topological_order (CyclicCompositionError) -- it is runtime/wiring's concern, not GQ's calibration DAG.
Calibration payload gotcha: decode's two non-edge inputs (input_ids, position_ids) need a PT-format payload (torch.saved dict[str, list[torch.Tensor]]) whose per-sample tensors already carry their own batch dimension (e.g. shape [1, 1], not [1]) -- unlike NPY, the PT loader does not add a batch dimension.