Downloads · 30 days
13
22% of all-time downloads
Supernova11c/Supernova-Nepali-Normalizer-V2
Supernova-Nepali-Normalizer-V2 is a machine learning model from Supernova11c. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
13
22% of all-time downloads
All-time downloads
59
Public
Repo size
—
Likes
1
Public
Click a slice to open those files.
.md4.4 KB · 61%
From the Hugging Face model README

Supernova Nepali Normalizer V2 is an independent, lightweight, rule-based text normalization system designed for Nepali and Nepali-English mixed text.
The normalizer was validated on 9,112 text samples containing approximately 936K characters.
Validation results:
V2 is a deterministic normalization system.
It does not use neural model weights and does not perform context-based spelling correction.
V2 can be used as:
V2 is an independent model/component.
It is not the same model as Supernova Nepali Normalizer V3.
V2 can optionally be used before or after V3 as a deterministic normalization and safety layer.
Apache-2.0 How to Use
##Supernova Nepali Normalizer V2 is a lightweight deterministic Nepali text normalization tool.
It is designed to clean common Unicode and punctuation inconsistencies in Nepali text.
Installation
Clone the repository:
git clone https://huggingface.co/Supernova11c/Supernova-Nepali-Normalizer-V2 cd Supernova-Nepali-Normalizer-V2
Python Usage
from normalizer import SupernovaNepaliNormalizer
normalizer = SupernovaNepaliNormalizer()
text = "कृपया कृपय\u200cा मलाई गणितमा कमजोर छु।"
result = normalizer.normalize(text)
print(result)
Example:
कृपया कृपया मलाई गणितमा कमजोर छु।
What V2 Does
Supernova Nepali Normalizer V2 currently provides deterministic normalization for:
For example:
\u200c \u200d
can be removed when they occur in unwanted positions.
Common punctuation variants are also normalized:
— → - – → - “ → " ” → " ‘ → ' ’ → '
No Neural Model Required
V2 does not require:
It runs locally using deterministic rules.
Example
from normalizer import SupernovaNepaliNormalizer
normalizer = SupernovaNepaliNormalizer()
examples = [ "मलाई नेपाली राम्रोसँग लेख्न सिक्नुछ।", "कृपया कृपय\u200cा मलाई गणितमा कमजोर छु।", "यो एउटा—परीक्षण वाक्य हो।" ]
for text in examples: print("Input :", text) print("Output:", normalizer.normalize(text)) print()
Design Philosophy
Supernova V2 is intentionally simple and predictable.
The normalizer performs only the transformations explicitly defined by its rules. It does not generate new text or make neural predictions.
This makes V2 suitable as a preprocessing component before other NLP systems.
License
See the repository license and project files for licensing information.
Supernova text processing architecture is engineered for extreme, zero-overhead systems efficiency. Running entirely on standard CPU hardware without any GPU acceleration or heavy vector models, it delivers elite-tier throughput:
Language Detection & Processing: 1,237,070,359+ characters/sec
Hardware Requirement: Standard CPU (Zero GPU dependency, ultra-low memory footprint)
Architecture: Modular, deterministic, and hallucination-free text pipeline.
Test Environment: Google Colab Free Tier (Standard Shared CPU Runtime)