Downloads · 30 days
39
11% of all-time downloads
Smilyai-labs/CodVa-1-Small-IT
CodVa-1-Small-IT is a text generation model from Smilyai-labs. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
Downloads · 30 days
39
11% of all-time downloads
All-time downloads
344
Public
Repo size
372 GB
Likes
0
Public
Click a slice to open those files.
.json3.6 MB · 99%
From the Hugging Face model README
CodVa-1-Small-IT is an instruction-tuned Large Language Model developed by Smilyai-labs — a team of high-school students passionate about AI research and development. This model is the instruction-tuned (IT) variant of our base model, CodVa-1-Small, fine-tuned to follow instructions, answer questions, and assist with coding and mathematical reasoning tasks.
| Component | Configuration |
|---|---|
| Architecture | Custom Decoder-Only Transformer |
| Hidden Dimension | 1536 |
| Layers | 28 |
| Attention Heads | 24 (Query) / 6 (KV) |
| Attention Type | Grouped Query Attention (GQA) |
| Positional Encoding | RoPE (θ = 5,000,000) |
| FFN Type | Dense SwiGLU + Sparse MoE |
| MoE Experts | 16 routed + 2 shared |
| MoE Top-K | 2 routed experts per token |
| MoE Hidden Dim | 1024 |
| MoE Frequency | Every 2 layers |
| Context Length | 4096 tokens |
| Vocabulary Size | 32,000 (+ special tokens) |
| Normalisation | RMSNorm throughout |
| QK Norm | ✅ Enabled |
| Structural Bias | ✅ Enabled (4 relation types) |
| Precision | BFloat16 |
The base model was pre-trained from scratch on a large corpus of code and mathematical text, optimised for strong reasoning and programming capabilities.
CodVa-1-Small-IT was produced by supervised fine-tuning (SFT) of the base model on a curated instruction dataset.
| Setting | Value |
|---|---|
| Fine-tuning Method | Supervised Fine-Tuning (SFT) |
| Dataset | Bc-AI/codva-it-data-v5.1 |
| Sequence Length | 4096 tokens |
| Optimizer | AdamW (8-bit where available) |
| Learning Rate | 2e-5 |
| LR Schedule | Cosine decay with warmup |
| Weight Decay | 0.01 |
| Gradient Clipping | 1.0 |
| Precision | BFloat16 |
CodVa-1-Small-IT is intended for:
As with all language models — particularly smaller ones — there are important limitations to be aware of:
Smilyai-labs is a team of high-school students who are passionate about AI research. We design and train our own model architectures from scratch, rather than fine-tuning existing open-source models, with the goal of learning every part of the deep learning stack — from architecture design and custom CUDA kernels to dataset curation and training infrastructure.
CodVa is our flagship model series, focused on code and mathematical reasoning.
We are students building real models. Feedback, collaboration offers, and questions are very welcome.
If you use CodVa-1-Small-IT in your research or projects, please consider citing or crediting the Smilyai-labs team:
@misc{codva1small,
author = {Smilyai-labs},
title = {CodVa-1-Small-IT: An Instruction-Tuned Code and Math Language Model},
year = {2025},
howpublished = {\url{https://huggingface.co/Smilyai-labs/CodVa-1-Small-IT}},
}
Please refer to the repository's license file for terms of use. If you intend to use this model for commercial purposes, please contact the Smilyai-labs team directly.