Downloads · 30 days
135
23% of all-time downloads
Banaxi-Tech/BananaMind-2.1-NanoCoder
BananaMind-2.1-NanoCoder is a text generation model from Banaxi-Tech. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
BananaMind 2.1 NanoCoder is an under-10M-parameter code language model.
Downloads · 30 days
135
23% of all-time downloads
All-time downloads
580
Public
Parameters
12.3M
923 MB on disk
Likes
1
Public
Click a slice to open those files.
.pt105 MB · 67%
From the Hugging Face model README
BananaMind 2.1 NanoCoder is an under-10M-parameter code language model.
L1 → L2 → L3 → L4 → L5 → L3 → L4 → L5 → L6 → L7 → L8The n-gram module has independent 29,744-entry bigram and four-gram hash tables, each with 32-dimensional values. Their concatenated representation is projected to the 256-wide residual stream. It is injected through separate learned gates at the beginning of both middle-stack passes.
The exact 30B-token streamed mixture is:
| Source | Tokens | Share |
|---|---|---|
| The Stack v3 train | 22.5B | 75% |
| FineWeb-Edu | 7.5B | 25% |
Stack v3 is streamed as repository-ordered source files. Vendored files are skipped, while repository path, file path, and detected language are included in the training text. FineWeb-Edu supplies prose, naming, comments, and general language knowledge.
Checkpoints are uploaded every 5% with safetensors, tokenizer files, metrics, pinned dataset revisions, exact source-token accounting, and optimizer state.
./launch_training_hf_job.sh 4 fresh
./launch_training_hf_job.sh 4 resume
./launch_training_hf_job.sh 8 resume