Downloads · 30 days
87
100% of all-time downloads
jsilvanus/aidos-echo-gguf
aidos-echo-gguf is a machine learning model from jsilvanus. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A byte-level echo (identity) "model" in GGUF. It is a smoke-test fixture, not a trained network — but it is a real llama-architecture transformer, so llama.cpp loads and runs it like any other model. The weights are h…
Downloads · 30 days
87
100% of all-time downloads
All-time downloads
87
Public
Repo size
2.4 MB
Likes
0
Public
Click a slice to open those files.
.gguf2.4 MB · 100%
From the Hugging Face model README
A byte-level echo (identity) "model" in GGUF. It is a smoke-test fixture, not a trained network — but it is a real llama-architecture transformer, so llama.cpp loads and runs it like any other model. The weights are hand-built, so greedy decoding emits exactly the last input token, unchanged, for every one of the 256 possible inputs. Any wrong byte is a runtime bug, never model drift.
Built for Aidos, a local-first AI agent project, as a fixture for testing its model-loading and inference runtime against a model whose correct output is known in advance for every input.
| File | echo.gguf (2.3 MB, GGUF v3, all F32) |
| Architecture | llama — 1 block, RMSNorm, RoPE, 1 head, SwiGLU FFN |
| Sizes | n_vocab 256, n_embd 256, n_ff 256, n_ctx 512 |
| Tokenizer | byte-level BPE, no merges — 1 token per byte, token id == byte |
| Contract | argmax at every position i is byte at position i, unchanged |
A smoke test needs a model whose correct output is known in advance, for every possible input. Real models do not offer that: their outputs shift with quantization, sampling, threading, and version bumps, so a test can only assert something vague and ends up passing on broken runtimes.
The identity function over bytes gives a total, exactly-specified function on all 256 inputs: the answer is always the input itself. This fixture computes it through genuine model machinery — float matmuls, an argmax, and a full RMSNorm/RoPE/attention/SwiGLU graph — so loading and running it exercises the same paths a real model would, while any wrong byte is unambiguously a runtime bug rather than model drift.
The model is not trained. The weights are constructed so the answer falls out
of the arithmetic; see the derivation below. This is the same construction as
aidos-rot13-gguf, with
the LM head set to the identity permutation instead of ROT13's.
Vocabulary and hidden size are both 256, one dimension per byte:
token_embd = I residual stream carries the one-hot e_t
attn_{q,k,v,output} = 0 the attention block contributes nothing
ffn_{gate,up,down} = 0 the FFN block contributes nothing
{attn,ffn,output}_norm = 1 RMSNorm scales but never mixes dimensions
output[t, t] = 1 the LM head is the identity permutation
The residual stream stays e_t through both blocks. The final RMSNorm turns
e_t into 16·e_t — the RMS of a one-hot vector in 256 dimensions is 1/16 —
so the logits are 16 at row t and 0 everywhere else. The argmax margin is
16.0: far wider than any fp16 or quantization error, so the model survives
conversion without changing its answer.
x, x, x, x, …. That means free-running generation does spell
out the right answer for a single-byte-repeated prompt, unlike ROT13.temp=0. Sampling at a non-zero temperature will
pick other tokens; that is expected, since all the losing logits are equal.Token 0 (NUL) is declared as BOS/EOS/UNK because a byte vocabulary has no better
sentinel, but the model only ever emits it when fed it. Generation is therefore
bounded by n_predict, not by EOS. add_bos_token is false, so tokenization
stays exactly one token per input byte.
Transform a whole string by reading the argmax at every position:
import llama_cpp, numpy as np
llm = llama_cpp.Llama(model_path="echo.gguf", n_ctx=512, logits_all=True, verbose=False)
tokens = llm.tokenize(b"Hello, World!", add_bos=False, special=False)
llm.eval(tokens)
print(bytes(int(np.asarray(llm.scores[i]).argmax()) for i in range(len(tokens))))
# b'Hello, World!'
Or check a single next-token prediction, which is all a minimal smoke test needs:
print(next(iter(llm.generate(tokens, temp=0.0)))) # 33 == ord('!')
With the llama.cpp CLI, free-running generation repeats the last prompt byte, since the identity function is a fixed point:
llama-cli -m echo.gguf -p "Hello" -n 4 --temp 0 # -> Hellooooo
The transduce.py script in the
Aidos repo
reads the argmax at every position in one forward pass, rather than sampling
from the end:
python3 transduce.py "Hello, World!" # Hello, World!
python3 transduce.py --mode stream "Hello, World!" # same, one token at a time
echo -n "Hello" | python3 transduce.py # Hello
Both modes are verified against real llama.cpp and must agree byte for byte:
This is the more useful shape for a runtime smoke test than generation is: it exercises prefill, per-position logits and the KV cache rather than a sampler, and it checks a whole string of known-correct bytes per pass instead of one.
Reading per-position logits requires the runtime to expose them. In
llama-cpp-python that means Llama(..., logits_all=True); without it, scores
is never populated at all, because sampling happens inside the sampler. A run
that silently returns zeros is that flag missing, not a broken model.
Source, build.py, and verify.py live in the
Aidos monorepo:
pip install gguf numpy # llama-cpp-python is optional, see below
python3 build.py # rewrites echo.gguf
python3 verify.py # exits non-zero on any mismatch
verify.py checks three independent layers:
gguf package, so a bad file is caught even if the writer was wrong.llama_cpp is importable, loads the file in actual
llama.cpp and checks tokenization and greedy decoding end to end. Skipped
with a note, not failed, when it is not installed.See also: aidos-echo-onxx,
the same fixture as a plain tensor-in/tensor-out ONNX graph, and
aidos-rot13-gguf, the
ROT13 sibling of this fixture.