Downloads · 30 days
138
5% of all-time downloads
spitfire4794/Alo-70m
Alo-70m is a text generation model from spitfire4794. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Alo-70M is the instruction-tuned version of the ultra-lightweight 69-million parameter Bengali language model, Alo-70M-Base. Built on a scaled-down LLaMA architecture, it is designed to act as a highly efficient, edge…
Downloads · 30 days
138
5% of all-time downloads
All-time downloads
2.6K
Public
Parameters
69M
276 MB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors276 MB · 98%
From the Hugging Face model README
Alo-70M is the instruction-tuned version of the ultra-lightweight 69-million parameter Bengali language model, Alo-70M-Base. Built on a scaled-down LLaMA architecture, it is designed to act as a highly efficient, edge-deployable localized AI assistant.
Fine-tuned on a curated dataset of instruction-response pairs using the ChatML format, Alo-70M is aligned for tasks such as summarization, entity extraction, text editing, and question answering in native Bengali. Despite its compact footprint, it offers a viable path for edge AI deployment on standard CPUs and mobile hardware.
Alo-70M was trained using the ChatML template. The chat template is built directly into the Jinja template of the tokenizer (spitfire4794/beng_bpe). You can leverage it using Hugging Face's apply_chat_template interface:
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "spitfire4794/Alo-70M"
# Load the custom Bengali BPE tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
# Define the instruction in ChatML format
messages = [
{"role": "user", "content": "নিচের অনুচ্ছেদটি সংক্ষেপে সারসংক্ষেপ করুন: [এখানে আপনার টেক্সট লিখুন]"}
]
# Apply the pre-configured ChatML template
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt"
).to(model.device)
# Generate text
outputs = model.generate(
**inputs,
max_new_tokens=150,
repetition_penalty=1.1,
do_sample=True,
temperature=0.6,
top_p=0.9
)
# Decode response (omitting user prompt)
response = outputs[0][inputs.shape[-1]:]
print(tokenizer.decode(response, skip_special_tokens=True))
Alo-70M was aligned using a curated subset of the Bangla-SFT-50k dataset formatted using ChatML.
adamw_torch_fused) with beta_1 = 0.9, beta_2 = 0.999, epsilon = 10^-8Like its base model, Alo-70M utilizes a parameter-efficient architecture:
The model was evaluated zero-shot across Bengali reasoning and knowledge benchmarks (continuation-based log-probability evaluation):
| Benchmark | Alo-70M (SFT) | Alo-70M-Base | Gemma-3-270M-IT | TigerLLM-1B-IT |
|---|---|---|---|---|
| bangla_mmlu_bn | 26.29% | 26.31% | 26.81% | 27.66% |
| bangla_commonsenseqa_bn | 25.88% | 28.42% | 22.77% | 25.14% |
| indicbench_arc_bn_challenge | 24.15% | 22.70% | 25.34% | 27.13% |
| boolqa_bn | 48.70% | 48.42% | 51.30% | 52.40% |
| openbookqa_bn | 30.58% | 31.39% | 31.99% | 34.21% |
| piqa_bn | 50.05% | 50.49% | 49.51% | 49.51% |
| hellaswag_bn | 26.89% | 27.27% | 27.85% | 31.01% |
Note: The 69M instruction-tuned model outperforms the larger Gemma-3-270M-IT baseline on tasks like CommonsenseQA and PIQA.
Technical paper out soon.