Downloads · 30 days
0
KiranAN1988/NeemSutra-125M-Python-Instruct-Beta
NeemSutra-125M-Python-Instruct-Beta is a text generation model from KiranAN1988. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as apache-2.0.
NeemSutra-125M-Python-Instruct is a 125M-parameter Small Language Model (SLM) developed from first principles as an experimental study of small language model design, tokenization, domain adaptation, supervised fine-t…
Downloads · 30 days
0
Access
Public
Updated Sep 16, 2026
Repo size
1.5 GB
Likes
0
Public
Click a slice to open those files.
.pt1.5 GB · 100%
From the Hugging Face model README
NeemSutra-125M-Python-Instruct is a 125M-parameter Small Language Model (SLM) developed from first principles as an experimental study of small language model design, tokenization, domain adaptation, supervised fine-tuning, reproducible training, inference behavior, and stress testing.
The project focuses on understanding how a relatively small decoder-only Transformer behaves as training data, domain adaptation, and instruction tuning are progressively applied.
This release is intended primarily as a research and educational artifact.
The released model uses a decoder-only Transformer architecture with:
The released source contains the model implementation used for inference.
NeemSutra uses the custom Sutra BPE v2 tokenizer.
The tokenizer is part of the experimental design and is released alongside the model.
To evaluate memory efficiency during inference and sequence generation, KV-cache memory consumption was profiled across different batch sizes and context lengths.
Assumptions:
| Batch Size ($B$) | Context Length ($T$) | Pre-KV Cache Memory | Post-KV Cache Memory | KV Cache Overhead |
|---|---|---|---|---|
| 1 | 64 | ~12.0 MB | ~18.0 MB | +6.0 MB |
| 1 | 256 (Max) | ~12.0 MB | ~36.0 MB | +24.0 MB |
| 4 | 64 | ~12.0 MB | ~36.0 MB | +24.0 MB |
| 4 | 256 (Max) | ~12.0 MB | ~84.0 MB | +72.0 MB |
This analysis is included to characterize inference memory behavior as batch size and context length increase.
The model was developed through sequential training stages:
| Stage | Dataset / Corpus | Training Tokens |
|---|---|---|
| Base Pretraining | Common Pile v0.1-derived corpus | ~5.75B |
| Python DAPT | Lovett01/Python-Code-Large | ~7.4B |
| Python SFT | OLMo-Coding/starcoder-python-instruct | ~2.79B |
| Extended DAPT | togethercomputer/RedPajama-Data-1T | ~24B |
| Final SFT | nvidia/OpenCodeInstruct | ~2.0B |
The stages were performed sequentially as part of the experimental training program.
The final released checkpoint corresponds to the OpenCodeInstruct SFT stage.
The training trajectory is part of the research artifact and is intended to make the progression from base pretraining through domain adaptation and instruction tuning inspectable.
The base model was trained on a project-selected corpus derived from the Common Pile v0.1 collection.
The local corpus was organized into:
Primary references:
Common Pile v0.1 is an 8 TB collection of public-domain and openly licensed text from more than 30 sources. NeemSutra used a project-selected subset/organization of this broader collection.
Base training: ~5.75B tokens.
Dataset:
Training: ~7.4B tokens.
Dataset:
Training: ~2.79B tokens.
Dataset:
Training: ~24B tokens.
Dataset:
Training: ~2.0B tokens.
OpenCodeInstruct is published by NVIDIA and is licensed under CC BY 4.0 according to its dataset card. Users should review the applicable terms and attribution requirements of the upstream datasets before using or redistributing derived artifacts.
NeemSutra was developed through a sequential training workflow rather than starting from an existing pretrained language model.
The experimental workflow included:
The purpose of this workflow was to examine the effects of progressively changing the training distribution and training objective on a small language model.
The project included engineering work across the training and inference pipeline, including:
The release includes the scripts required to reproduce local inference with the published checkpoint.
The project focuses on:
The objective is to study how a small language model behaves as training data, domain adaptation, and instruction tuning are progressively applied.
This experimental model was developed using AI-assisted engineering workflows.
AI tools were used as engineering support during different stages of the project, including:
The model architecture, training experiments, dataset preparation, training execution, evaluation methodology, checkpoint management, and final release were conducted as part of the NeemSutra experimental workflow.
Evaluation was performed throughout development using intermediate and final checkpoints.
The evaluation suite includes:
The evaluation process was designed to examine multiple aspects of small-model behavior rather than relying on a single aggregate metric.
The released inference script also provides deterministic greedy decoding and configurable sampling for reproducible local tests.
The release includes a self-contained inference workflow supporting:
Inference supports deterministic greedy decoding as well as configurable sampling.
Refer to inference.py for the local inference entry point.
The released checkpoint includes SHA-256 verification information to allow users to verify the integrity of the downloaded model artifact.
Checkpoint:
neemsutra_125m_python_instruct.pt
SHA-256:
0f55fd31e4a703c5fe27029c4d00111671d65a09329acac0c64902133c5308d
NeemSutra-125M-Python-Instruct-Beta is intended for:
The model should be considered an experimental research artifact, not a production-ready general-purpose language model.
NeemSutra is a 125M-parameter experimental SLM with a 256-token context window.
Its behavior can vary significantly depending on:
The model may produce syntactically invalid, incomplete, repetitive, or incorrect code and should not be assumed to be reliable without validation.
The model and accompanying code in this repository are released under the Apache License 2.0, except where otherwise specified by upstream datasets or third-party components.
Users are responsible for reviewing the licenses and attribution requirements of upstream datasets before redistributing or using derived artifacts.
If you use NeemSutra in research, experimentation, or educational work, please cite this repository:
@misc{neemsutra2026,
title = {NeemSutra-125M-Python-Instruct-Beta},
author = {Kiran A N},
year = {2026},
publisher = {Hugging Face},
note = {Experimental 125M-parameter Small Language Model},
url = {https://huggingface.co/KiranAN1988/NeemSutra-125M-Python-Instruct-Beta}
}
NeemSutra-125M-Python-Instruct-Beta is a first-principles design and analysis study of a small language model, covering the complete experimental path from tokenizer construction and base pretraining through domain adaptation, supervised fine-tuning, inference optimization, evaluation, and stress testing.
The release is intended to make the architecture, training trajectory, engineering workflow, and observed SLM behavior inspectable and reproducible.