Downloads · 30 days
41
6% of all-time downloads
spitfire4794/Alo-70m-Base
Alo-70m-Base is a machine learning model from spitfire4794. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Alo-70M-Base is a compact, 69-million parameter foundational Large Language Model (LLM) built exclusively for the Bengali (Bangla) language. Trained entirely from scratch, it aims to lower the compute barrier for Beng…
Downloads · 30 days
41
6% of all-time downloads
All-time downloads
673
Public
Parameters
69M
276 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors276 MB · 98%
From the Hugging Face model README
Alo-70M-Base is a compact, 69-million parameter foundational Large Language Model (LLM) built exclusively for the Bengali (Bangla) language. Trained entirely from scratch, it aims to lower the compute barrier for Bengali NLP research and provide a viable path for deploying localized AI assistants on standard CPUs and resource-constrained edge devices.
Despite its ultra-lightweight footprint, Alo-70M-Base achieves competitive parity on localized reasoning tasks against significantly larger cross-lingual models (such as 270M and 1B parameter baselines).
This model is the base foundational model. We also release the instruction-tuned version, the synthetic dataset used for alignment, and the standalone tokenizer:
You can load and generate text with Alo-70M-Base using the transformers library.
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "spitfire4794/Alo-70M-Base"
tokenizer_id = "spitfire4794/beng_bpe"
# Load the custom Bengali BPE tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(tokenizer_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
prompt = "বাংলাদেশের রাজধানী ঢাকা"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
# Generate text
outputs = model.generate(
**inputs,
max_new_tokens=50,
repetition_penalty=1.1,
do_sample=True,
temperature=0.7
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Alo-70M-Base is based on a scaled-down LLaMA transformer architecture optimized for hardware alignment and parameter efficiency.
silu)tie_word_embeddings = False). The static lookup tables account for 48.6% (33.5M) of the total parameter budget.The model was pre-trained on a curated, strictly Bengali corpus containing approximately 19.25 billion tokens. The dataset is a blend of:
Alo-70M-Base was initialized with completely randomized weights (no transfer learning) and trained from scratch.
adam_torch_xla) with beta_1 = 0.9, beta_2 = 0.95, epsilon = 10^-8bfloat16 (bf16) precision with gradient checkpointing enabled to maximize throughput.The model was evaluated in a zero-shot setting across various Bengali reasoning and knowledge benchmarks using a continuation-based evaluation methodology (calculating conditional log-probabilities directly over raw Bengali text to avoid "pointer-reasoning" bottlenecks).
| Benchmark | Alo-70M-Base | Gemma-3-270M-IT | TigerLLM-1B-IT |
|---|---|---|---|
| bangla_mmlu_bn | 26.31% | 26.81% | 27.66% |
| bangla_commonsenseqa_bn | 28.42% | 22.77% | 25.14% |
| indicbench_arc_bn_challenge | 22.70% | 25.34% | 27.13% |
| boolqa_bn | 48.42% | 51.30% | 52.40% |
| openbookqa_bn | 31.39% | 31.99% | 34.21% |
| piqa_bn | 50.49% | 49.51% | 49.51% |
| hellaswag_bn | 27.27% | 27.85% | 31.01% |
Note: Alo-70M-Base matches or outperforms both the 270M and 1B baselines on CommonsenseQA and PIQA despite having 4x to 14x fewer parameters.
Technical paper out soon.