Downloads · 30 days
7
16% of all-time downloads
ML-is-Fun/Fleck-M-500K-Base
Fleck-M-500K-Base is a machine learning model from ML-is-Fun. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Fleck-M-500K-Base — a compact decoder-only language model trained from scratch on 100M real tokenizer tokens.
Downloads · 30 days
7
16% of all-time downloads
All-time downloads
45
Public
Parameters
497K
997 KB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors997 KB · 93%
From the Hugging Face model README
Fleck-M-500K-Base — a compact decoder-only language model trained from scratch on 100M real tokenizer tokens.
| Architecture | Decoder-only Transformer |
| Parameters | 497,288 |
| Hidden size | 128 |
| FFN size | 336 |
| Physical blocks | 2 |
| Effective depth | 4 (A → B → A → B) |
| Attention | GQA — 4 query heads, 2 KV heads, head dimension 32 |
| Normalization | RMSNorm |
| Embedding | Factorized tied embedding, rank 64 |
| Vocabulary | 2,048 |
| Context length | 2,048 tokens |
| Canonical dtype | BF16 |
| Initialization | Fresh initialization |
| Dataset | FineWeb-Edu / FineWeb |
| Data mixture | 70% FineWeb-Edu + 30% FineWeb |
| Real tokenizer tokens | 100,000,000 exactly |
| Instruction tuning | No |
| Hardware | Apple M2 (10-core GPU) |
This is the immutable Base-100M parent used to produce the independent Fleck-M-500K instruction-tuned model.
The Base-100M checkpoint was evaluated with the corrected zero-shot aggregation protocol. This minimal public bundle does not include the evaluation artifact; the table below is a reference result for the released Base model.
Evaluation conditions: zero-shot, no chat template, FP32 evaluation, Apple Silicon MPS.
| Task | Metric | Shots | Base |
|---|---|---|---|
| HellaSwag | acc_norm | 0 | 25.04% |
| PIQA | acc_norm | 0 | 51.31% |
| ARC-Easy | acc_norm | 0 | 26.05% |
| ARC-Challenge | acc_norm | 0 | 24.15% |
| LAMBADA OpenAI | acc | 0 | 0.04% |
| WinoGrande | acc | 0 | 51.30% |
| BoolQ | acc | 0 | 37.83% |
| MMLU (57-subject macro) | acc | 0 | 23.14% |
| Eight-task mean | — | 0 | 29.86% |
Fleck-Tokenizer-2048| Token | ID | Role |
|---|---|---|
<bos> | 0 | sequence start |
<eos> | 1 | sequence end |
<pad> | 2 | padding |
<unk> | 3 | unknown token |
<|system|> | 4 | system turn |
<|user|> | 5 | user turn |
<|assistant|> | 6 | assistant turn |
<|eot|> | 7 | end of turn |
The bundle includes a self-contained inference.py; it does not import the Fleck-LM checkout. The accompanying config.json, generation_config.json, and tokenizer_config.json describe the custom architecture and generation/tokenizer defaults; standard transformers.AutoModel loading is not supported. Install the three runtime dependencies and run a single greedy continuation:
python -m pip install torch safetensors tokenizers
python inference.py \
--ckpt model.safetensors \
--tokenizer tokenizer.json \
--prompt "Hello, world" \
--max-tokens 32 \
--device cpu
The script reads and runs the BF16 checkpoint without an FP32 model copy, validates every SafeTensors key, shape, and dtype, and uses FP32 only for attention score/softmax and tied-logit accumulation. The model implements its factorized tied embedding/logits, effective-depth execution A → B → A → B, half-split RoPE, GQA, physical KV caches, RMSNorms, and greedy generation without repository-local imports.
This model is extremely small and is intended for research and local experimentation rather than general-purpose language generation. It may produce repetitions, malformed text, weak factual answers, or incoherent continuations. Benchmark scores should be interpreted in the context of the 497,288 parameter count.
MIT License.
The public bundle contains these files:
README.md — model card and usage documentationinference.py — standalone strict loader and greedy inference CLImodel.safetensors — BF16 model weightstokenizer.json — standalone tokenizerconfig.json — custom architecture configurationgeneration_config.json — greedy generation defaultstokenizer_config.json — tokenizer defaults and special-token mapping
No training data, optimizer state, or other training outputs are included.