Downloads · 30 days
54
40% of all-time downloads
SulphurAI/sulphur-caption
sulphur-caption is a video-text-to-text model from SulphurAI. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-4.0.
This is a captioning model I made purely to make things easier and cheaper for myself. If you don't like how the model works or it doesn't work on your favorite runtime, I'm sorry, but I'm not going to put any further…
Downloads · 30 days
54
40% of all-time downloads
All-time downloads
135
Public
Repo size
66.1 GB
Likes
10
Public
Click a slice to open those files.
.safetensors65.8 GB · 99%
From the Hugging Face model README
This is a captioning model I made purely to make things easier and cheaper for myself. If you don't like how the model works or it doesn't work on your favorite runtime, I'm sorry, but I'm not going to put any further effort into this, as it works how I want it to. If someone wants to figure out gguf or anything like that, you're welcome to try, you can dm me on discord under the username kiwv. Also it's worth noting this model will not be used in sulphur 3's dataset creation.
You're going to need ~70gb vram to run this model. The fp8 variant will require ~38gb.
Note the fp8 variant is ~1.6x faster than the bf16 variant.
Also this model isn't SUPER great, it's just solid. I wouldn't rely on this model solely, if you're using it to produce captions for a lora, look at the captions, or pass in tags, or something.
I was planning to release this model at a better time, but it's been sitting here so long I just decided to release it. I would probably change some stuff if I redid this model now but whatever.
You need ffmpeg, ffprobe, and unzip available on PATH.
cd /wherever/you/downloaded/this/thing
unzip -q vllm.zip
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip setuptools wheel
pip install -r requirements.txt
Or run the helper:
cd /wherever/you/downloaded/this/thing
./setup_vllm.sh
source .venv/bin/activate
Edit VIDEO_PATH in example_run.py:
VIDEO_PATH = "/path/to/video.mp4"
Then run:
cd /path/to/this/repo
source .venv/bin/activate
python example_run.py
The script prints the raw model output.
Edit these variables in caption_folder.py:
INPUT_FOLDER = "/path/to/videos"
OUTPUT_FOLDER = "/path/to/captions"
Then run:
cd /path/to/this/repo
source .venv/bin/activate
python caption_folder.py
The script mirrors the input folder structure into OUTPUT_FOLDER and writes one .txt file per video. It loads the model once and batches videos using BATCH_SIZE.
The model was trained with some settings, grouped under Prompt Settings. Here are the recommended values:
HAS_THINKING | TrueCAPTION_LENGTH | "very large"INCLUDE_WATERMARK_INFO | TrueVULGARITY | "low"UNCERTAINTY | "low"CHARACTER_NAMES | "none"FLUFF | "none"HAS_REPETITION | FalseSPECULATION | "low"TEMPORAL_DETAIL | "medium"VISUAL_SPECIFICITY | "moderate"CAMERA_DETAIL | "medium"CAPTION_STYLE | "plain"HAS_THINKING | True, FalseCAPTION_LENGTH | "very small", "small", "medium", "large", "very large"INCLUDE_WATERMARK_INFO | True, FalseVULGARITY | "none", "low", "medium", "high"UNCERTAINTY | "none", "low", "medium", "high"CHARACTER_NAMES | "none", "ambiguous", "single", "multiple"FLUFF | "none", "low", "medium", "high"HAS_REPETITION | True, FalseSPECULATION | "none", "low", "medium", "high"TEMPORAL_DETAIL | "static", "low", "medium", "high"VISUAL_SPECIFICITY | "generic", "moderate", "detailed", "excessive"CAMERA_DETAIL | "none", "low", "medium", "high"CAPTION_STYLE | "plain", "verbose", "ornate", "robotic"Generation settings are grouped under vLLM / Generation Hyperparameters:
NUM_FRAMES | (was trained between 4 and 20)MAX_MODEL_LENMAX_NUM_SEQSBATCH_SIZE in caption_folder.pyGPU_MEMORY_UTILIZATIONMAX_TOKENSTEMPERATURE | (recommend a lower value)TOP_PREPETITION_PENALTYThis graph shows measured vLLM decode throughput on an RTX PRO 6000 using 12 video frames, max_tokens = 96, max_model_len = 4096, prefix cache disabled, and unique requests:

