Downloads · 30 days
27
22% of all-time downloads
Elldreth/t5_base_prompt_translator
t5_base_prompt_translator is a text generation model from Elldreth. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Transform natural language descriptions into optimized WD14 tags for Stable Diffusion!
Downloads · 30 days
27
22% of all-time downloads
All-time downloads
125
Public
Parameters
223M
892 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors892 MB · 100%
From the Hugging Face model README
Transform natural language descriptions into optimized WD14 tags for Stable Diffusion!
This model translates creative natural language prompts into standardized WD14-format tags, trained on 95,000 high-quality prompts from Arcenciel.io.
from transformers import T5Tokenizer, T5ForConditionalGeneration
# Load model and tokenizer
tokenizer = T5Tokenizer.from_pretrained("Elldreth/t5_base_prompt_translator")
model = T5ForConditionalGeneration.from_pretrained("Elldreth/t5_base_prompt_translator")
# Translate a prompt
prompt = "translate prompt to tags: magical girl with blue hair in a garden"
inputs = tokenizer(prompt, return_tensors="pt", max_length=160, truncation=True)
outputs = model.generate(**inputs, max_length=256, num_beams=4)
tags = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(tags)
# Output: 1girl, blue hair, garden, outdoors, solo, long hair, dress, flowers, standing, day, smile, magical girl
Name: t5_base_prompt_translator
Base Model: T5-Base (Google)
Parameters: 220 million
Training Data: 95,000 high-quality prompts from Arcenciel.io
Training Duration: ~10 hours on RTX 4090
Model Size: ~850 MB
Accuracy: 85-90% tag matching
Final Loss: ~1.2-1.3
Inference Performance (RTX 4090):
VRAM Usage:
Note: The model performs exceptionally well even at high beam counts (32-64) on RTX 4090, making it practical to use maximum quality settings for production work.
Input Format:
translate prompt to tags: [natural language description]
Output Format:
tag1, tag2, tag3, tag4, ...
Tag Format:
tag \(descriptor\)shrug \(clothing\), blue eyes, long hairNote: Quality filtering was intentionally avoided to prevent limiting the training data diversity. Engagement metrics (hearts, likes) are not consistently used across the site, so filtering by them would have reduced dataset quality rather than improved it.
config.json - Model configurationmodel.safetensors - Model weights (safetensors format)tokenizer_config.json - Tokenizer configurationspiece.model - SentencePiece tokenizer modelspecial_tokens_map.json - Special tokens mappingadded_tokens.json - Additional tokensgeneration_config.json - Generation defaultstraining_args.bin - Training arguments (metadata)This model is based on T5-Base by Google, which is licensed under Apache 2.0.
Model License: Apache 2.0
Training Data: Arcenciel.io (public API)
Usage: Free for commercial and non-commercial use
If you use this model in your work, please cite:
T5X Prompt Translator Base 95K
Trained on Arcenciel.io dataset using WD14 v1.4 MOAT tagger
Base model: T5-Base (Google)
Version 1.0 (Current)
This model is designed to work with the ComfyUI-T5X-Prompt-Translator custom node:
See the ComfyUI custom node repository for installation instructions.
Primary Use Case: Converting creative natural language descriptions into optimized WD14-format tags for Stable Diffusion image generation.
Example Applications:
"translate prompt to tags: [your description]"Note: Quality filtering was intentionally avoided to maximize training data diversity. Engagement metrics (hearts, likes) are not consistently used across the source platform.
@misc{t5-base-prompt-translator,
title={T5 Base Prompt Translator: Natural Language to WD14 Tags},
author={Elldreth},
year={2024},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/Elldreth/t5_base_prompt_translator}},
}