Downloads · 30 days
0
Supernova11c/Supernova-Nepali-Tokenizer
Supernova-Nepali-Tokenizer is a machine learning model from Supernova11c. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
A high-performance, production-ready Byte-Level BPE tokenizer accuracy engineered for the Nepali language and Devanagari script. Developed as part of the Supernova project to enable efficient and accurate Nepali LLM p…
Downloads · 30 days
0
Access
Public
Updated Sep 16, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.json176 KB · 96%
From the Hugging Face model README
A high-performance, production-ready Byte-Level BPE tokenizer accuracy engineered for the Nepali language and Devanagari script. Developed as part of the Supernova project to enable efficient and accurate Nepali LLM processing.
Tested on the Supernova-teraillm dataset:
| Tokenizer | Tokens per Word | Efficiency |
|---|---|---|
| Supernova-Nepali (Ultra) | 3.79 | 2.20x Better |
| GPT-2 (Standard) | 8.21 | Baseline |
import time
from transformers import AutoTokenizer
# Load the dedicated Nepali tokenizer (Pure Tokenizer Repository)
model_id = "Supernova11c/Supernova-Nepali-Tokenizer"
print(f"Loading tokenizer for: {model_id}")
# Use clean_up_tokenization_spaces=False for BPE tokenizers to prevent warnings/corruption
tokenizer = AutoTokenizer.from_pretrained(model_id, clean_up_tokenization_spaces=False)
def stress_test_tokenizer(tokenizer):
print(f"\n--- Running Tokenizer Stress Test ---")
print(f"Tokenizer Class: {type(tokenizer).__name__}\n")
# 1. Edge Cases & Special Characters Test
edge_cases = [
"Hello, world! 🌍🚀", # Emojis & punctuation
" Multiple spaces and\nnewlines\t", # Whitespace handling
"The quick brown fox jumps over the lazy dog." * 50, # Repetition
"1234567890 -+*/=<>@#$%^&*()_[]{}|\\:;\"'.,?", # Symbols & Numbers
"नमस्ते संसार 🌟 नेपाल 🌍", # Nepali / Multi-lingual
"", # Empty string
]
print("1. Edge Case Testing:")
for i, text in enumerate(edge_cases):
try:
encoded = tokenizer.encode(text)
decoded = tokenizer.decode(encoded, skip_special_tokens=True)
match = "✓" if (text.strip() == decoded.strip() or not text) else "⚠️ (Whitespace diff)"
print(f" Test {i+1}: {match} | Length: {len(text)} chars -> {len(encoded)} tokens")
except Exception as e:
print(f" Test {i+1}: ❌ FAILED with error: {e}")
# 2. Throughput / Speed Test
print("\n2. Throughput Performance Test:")
sample_text = (
"नेपाल एक सुन्दर देश हो। यहाँ विभिन्न जातजाति र भाषाभाषीका मानिसहरू बसोबास गर्छन्। "
) * 500 # ~35,000 characters
num_iterations = 100
# Warmup
_ = tokenizer.encode(sample_text)
start_time = time.time()
for _ in range(num_iterations):
_ = tokenizer.encode(sample_text)
end_time = time.time()
total_time = end_time - start_time
total_chars = len(sample_text) * num_iterations
total_tokens = len(tokenizer.encode(sample_text)) * num_iterations
print(f" Processed {total_chars:,} characters in {total_time:.4f} seconds.")
print(f" Speed: {total_chars / total_time:,.2f} chars/sec")
print(f" Speed: {total_tokens / total_time:,.2f} tokens/sec")
# 3. Vocabulary & Configuration Check
print("\n3. Vocabulary & Configuration Check:")
print(f" Vocabulary Size: {len(tokenizer):,}")
print(f" Model Max Length: {getattr(tokenizer, 'model_max_length', 'N/A')}")
print(f" Pad Token: {tokenizer.pad_token} (ID: {tokenizer.pad_token_id})")
print(f" EOS Token: {tokenizer.eos_token} (ID: {tokenizer.eos_token_id})")
print("\n--- Stress Test Complete ---")
# Execute the test
stress_test_tokenizer(tokenizer)
[PAD], [UNK], [BOS], [EOS]Supernova text processing architecture is engineered for extreme, zero-overhead systems efficiency. Running entirely on standard CPU hardware without any GPU acceleration or heavy vector models, it delivers elite-tier throughput: