Downloads · 30 days
0
Ailaysa-AI/asai-tokenizer-model
asai-tokenizer-model is a machine learning model from Ailaysa-AI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
<p align="center" <img src="https://ailaysa.com/logo512.png" alt="Ailaysa" width="200"/ </p
Downloads · 30 days
0
Access
Public
Updated Mar 26, 2026
Repo size
1.1 MB
Likes
6
Public
Click a slice to open those files.
.model1.1 MB · 54%
From the Hugging Face model README
Asai (அசை) — the fundamental unit of rhythm in Tamil prosody (யாப்பிலக்கணம்). In classical Tamil literature, Asai represents the cadence formed by letters, classified into:
Just as Asai forms the building blocks of Tamil verse, this tokenizer provides the foundational building blocks for Tamil language AI.
> "To build AI that understands Indic languages, one must first understand their soul."
| Feature | Benefit |
|---|---|
| Morphological Awareness | Preserves Tamil suffix chains and grammatical markers |
| Semantic Density | Each token carries more linguistic meaning |
| Unicode Normalization | Handles inconsistencies across input systems |
| Production-Ready | Fast, efficient, easy integration |
Comparison on Tamil sentence: "தமிழை உலகமெங்கும் கொண்டு சேர்ப்போம்."
| Tokenizer | Tokens | Efficiency |
|---|---|---|
| Asai | 8 | 100% |
| GPT-4.x & Legacy | 51 | 15.7% |
| LLaMA-3 | 54 | 14.8% |
| Mistral | 48 | 16.7% |
| Qwen | 42 | 19% |
Key Metrics:
pip install ailaysa
from ailaysa import tokenizer
# Load tokenizer
tok = tokenizer.load("asai-v1")
# Input text
text = "தமிழை உலகமெங்கும் கொண்டு சேர்ப்போம்."
# Encode
encoded = tok.encode(text)
print(encoded.ids) # Token IDs
print(encoded.tokens) # Token strings
print(encoded.length) # Number of tokens
Asai introduces a Linguistic Semantic Layer that operates before token segmentation:
| Traditional Tokenizers | Asai Approach |
|---|---|
| Optimize for compression | Optimize for linguistic fidelity |
| Statistical frequency | Morphological structure |
| Language-agnostic | Tamil-first, extensible to Indic |
asai-v1/
├── Linguistic Semantic Layer # Pre-segmentation analysis
├── Morphological Analyzer # Root + suffix identification
├── Subword Segmenter # Optimized token generation
└── Unicode Normalizer # Input standardization
Built by a growing community of AI engineers, researchers, linguists, and open-source contributors.
@software{ailaysa2026,
title = {Ailaysa: Indic Language NLP Toolkit},
author = {Mukesh Anand G and Ailaysa Technologies},
year = {2026},
url = {https://github.com/Ailaysa-Technologies/Asai-Tokenizer}
}
MIT License — Open for research, commercial, and personal use.
<p align="center"> <b>Built with precision. Inspired by heritage. Open for the future.</b> </p>