Downloads · 30 days
38
100% of all-time downloads
Henil1/rumik-oss-1-FP8-Dynamic
rumik-oss-1-FP8-Dynamic is a text-to-speech model from Henil1. Use it when you need text read aloud. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
An FP8 quantization of rumik-ai/rumik-oss-1, Rumik's 3B multilingual text to speech model for 22 Indic languages and English, with expressive delivery control, inline vocalizations and 24 kHz audio output.
Downloads · 30 days
38
100% of all-time downloads
All-time downloads
38
Public
Parameters
3.4B
4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4 GB · 99%
How the weights are stored.
F8_E4M32.8B · 83%
From the Hugging Face model README
An FP8 quantization of rumik-ai/rumik-oss-1, Rumik's 3B multilingual text to speech model for 22 Indic languages and English, with expressive delivery control, inline vocalizations and 24 kHz audio output.
This checkpoint stores the weights of the transformer linear layers in 8 bit floating point (FP8) and quantizes activations dynamically at runtime. It needs no calibration data and keeps the output heads in bf16. The goal is a smaller checkpoint and lower memory use with the same behavior as the original model.
This is an independent quantization. It is not an official Rumik release. All credit for the model, training and data goes to the Rumik team.
| Base model | rumik-ai/rumik-oss-1 |
| Method | llm-compressor, FP8_DYNAMIC scheme (data free) |
| Weights | FP8 (E4M3), per channel scales |
| Activations | FP8 (E4M3), dynamic per token scales |
| Kept in bf16 | lm_head, stop_predictor, embeddings, norms |
| Format | compressed-tensors |
| Size on disk | 4.0 GB (bf16 original: 7.1 GB) |
Native FP8 compute needs an NVIDIA GPU with compute capability 8.9 or higher (Ada Lovelace, Hopper or Blackwell). On GPUs without FP8 support, use the original bf16 model instead.
pip install -U transformers accelerate compressed-tensors
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "Henil1/rumik-oss-1-FP8-Dynamic"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo, device_map="auto", trust_remote_code=True
)
The model uses custom code, so trust_remote_code=True is required. Speech generation, the voice list and the audio codec setup are unchanged from the original model, so follow the usage section of the original model card for generating and decoding audio. 100 audio tokens make one second of speech, and the Mimi codec decodes them to 24 kHz audio.
| Control | Values |
|---|---|
| Voice | Ira, Aisha, Siya, Zoya |
| Tone | happy, sad, angry, excited, professional |
| Accent | Hindi, Telugu, Tamil, Kannada, Bengali, Punjabi, Indian English |
| Pace | slow, fast, steady |
| Inline sounds | <laugh>, <chuckle>, <sigh> |
The delivery description can also be written inline, for example <description="excited, Hindi accent, fast pace"> your text here <laugh>.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
SRC = "rumik-ai/rumik-oss-1"
model = AutoModelForCausalLM.from_pretrained(
SRC, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True
)
tok = AutoTokenizer.from_pretrained(SRC, trust_remote_code=True)
recipe = QuantizationModifier(
targets="Linear",
scheme="FP8_DYNAMIC",
ignore=["lm_head", "re:.*stop_predictor.*"],
)
oneshot(model=model, recipe=recipe)
model.save_pretrained("rumik-oss-1-FP8-Dynamic", save_compressed=True)
tok.save_pretrained("rumik-oss-1-FP8-Dynamic")
Research and non commercial use only, under the CC BY NC 4.0 license with the acceptable use addendum inherited from Tiny Aya Fire, unchanged from the original model. The original LICENSE and NOTICE files are included with this repository, and this repository adds FP8 weight quantization as a modification. The Mimi codec is licensed under CC BY 4.0.
@unpublished{govindu2026rumikoss1,
title = {{rumik-oss 1 technical report}},
author = {Govindu Pranav and Anant Shukla and Suryansh Shakya and Aman Anand and Vatsal Bharti},
year = {2026},
note = {In preparation}
}