Downloads · 30 days
6
5% of all-time downloads
SpiceeChat/Genre-Classifier-1-20M-BASE-BF16
Genre-Classifier-1-20M-BASE-BF16 is a text classification model from SpiceeChat. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
<p align="center" <img src="https://huggingface.co/spaces/SpiceeChat/README/resolve/main/Spiceechat.png" alt="SpiceeChat" width="1400" </p
Downloads · 30 days
6
5% of all-time downloads
All-time downloads
131
Public
Parameters
19.7M
39.4 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors39.4 MB · 93%
From the Hugging Face model README
A lightweight 20M-parameter CausalLM built for first-name gender classification. This is a base model, designed to be fine-tuned on downstream tasks rather than used directly.
| Property | Value |
|---|---|
| Architecture | FirstNameGenderForCausalLM |
| Parameters | ~20M |
| Context Length | 20 tokens |
| Layers | 4 |
| Attention Heads | 4 |
| Hidden Size | 384 |
| Vocab Size | 32,768 |
| Tensor Type | BF16 |
| License | Apache 2.0 |
| Token | ID |
|---|---|
F_ID | 42 |
M_ID | 49 |
PAD_ID | 0 |
The model uses a causal language modeling objective with weight tying between the input embedding and output head (head.weight = tok_emb.weight).
A lightweight GPT-style decoder with:
sageattention is not installed)from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained(
"SpiceeChat/Genre-Classifier-1-20M-BASE-BF16",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(
"SpiceeChat/Genre-Classifier-1-20M-BASE-BF16",
trust_remote_code=True
)
The model provides a dedicated predict_gender() method:
inputs = tokenizer("Arjun", return_tensors="pt")
pred_idx, probs = model.predict_gender(inputs.input_ids)
gender = "M" if pred_idx.item() == 1 else "F"
print(gender) # M
This base model was pre-trained on a large-scale first-name dataset. It is not fine-tuned for any specific downstream task — it's meant to be used as a starting point.
To fine-tune this model on your own dataset:
trust_remote_code=TrueNote: The model expects input sequences of length ≤ 20 tokens. Longer names will be truncated.
| Package | Version |
|---|---|
transformers | >= 4.30.0 |
torch | >= 2.0.0 |
sageattention | optional, for faster attention |
Built by PhysiQuanty for SpiceeChat.