Downloads · 30 days
0
Jiaha0Hu4ng/OneEmo-Base
OneEmo-Base is a video-text-to-text model from Jiaha0Hu4ng. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
<p align="center" <img src="https://cdn-uploads.huggingface.co/production/uploads/66953e2f417bbfcd51933cd0/0XVrXOZhxIMCYTwecNa.png" width="280"/ </p OneEmo-Base is a unified multimodal reasoning model for emotion perc…
Downloads · 30 days
0
Access
Public
Updated Aug 7, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md9.5 KB · 86%
From the Hugging Face model README
This is a SFT stage 1 training version of OneEmo which using Curriculum SFT strategy.
The model was trained and evaluated in the following environment:
torch: 2.10.0+cu128torchaudio: 2.10.0+cu128torchvision: 0.25.0+cu128transformers: 5.2.0vllm: 0.19.0ms-swift: 4.1.3 (optional, for alternative inference engine)oneemo)Use the code below to get started with the model.
vllm serve Jiaha0Hu4ng/OneEmo-Base \
--host 0.0.0.0 \
--port 8000 \
--served-model-name OneEmo-Base \
--max-model-len 32768 \
--mm-encoder-tp-mode data \
--media-io-kwargs '{"video": {"num_frames": 16}}' \
--default-chat-template-kwargs '{"enable_thinking": true}' \
--reasoning-parser qwen3
import torch
from swift import get_model_processor, get_template
from swift.infer_engine import TransformersEngine, InferRequest, RequestConfig
model_path = "Jiaha0Hu4ng/OneEmo-Base"
model, processor = get_model_processor(
model_path,
model_type="qwen3_5",
torch_dtype=torch.bfloat16,
attn_impl="flash_attn",
)
template = get_template(processor, enable_thinking=True)
engine = TransformersEngine(model, template=template)
request_config = RequestConfig(max_tokens=2048, temperature=0.7)
infer_request = InferRequest(
messages=[{
"role": "user",
"content": (
"Based on the video, describe the range of emotions the character "
"may be feeling."
)
}],
videos=["path/to/your/video.mp4"]
)
resp_list = engine.infer([infer_request], request_config=request_config)
print(resp_list[0].choices[0].message.content)
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
from qwen_vl_utils import process_vision_info
model_path = "Jiaha0Hu4ng/OneEmo-Base"
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2", # fallback to "sdpa" if flash-attn is not installed
device_map="auto",
)
processor = AutoProcessor.from_pretrained(model_path)
# Example: open-vocabulary multimodal emotion recognition (OVMER)
messages = [
{
"role": "user",
"content": [
{
"type": "video",
"video": "path/to/your/video.mp4",
"max_pixels": 360 * 420,
"fps": 2.0,
},
{
"type": "text",
"text": "Based on the video, describe the range of emotions the character may be feeling.",
},
],
}
]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
).to(model.device)
with torch.no_grad():
generated_ids = model.generate(
**inputs,
max_new_tokens=2048,
temperature=0.7,
top_p=0.9,
top_k=50,
)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text[0])
For other tasks, replace the prompt text with the corresponding template (see the full model card on Hugging Face for the complete prompt table). The model also supports deployment with vLLM (OpenAI-compatible API) and inference via ms-swift; refer to the repository for those examples.
We use Emo-World-130K to train OneEmo.
[To Be continued]
[To Be continued]
[To Be continued]
[To Be continued]
[To Be continued]
[To Be continued]
<!-- **APA:** Huang, J. et al. (2026). *OneEmo: A Unified Multimodal Model for Emotion, Intention, and Mental-Health Understanding in Videos*. Hugging Face. https://huggingface.co/Jiaha0Hu4ng/OneEmo -->OneEmo can be used directly for video-based emotion analysis, intention recognition, sentiment classification, humour/sarcasm detection, and for generating empathetic or supportive textual responses. It is suitable for researchers and developers building affective computing applications, mental-health chatbots, or social signal processing tools.
The model can be fine-tuned on domain-specific datasets (e.g., clinical interviews, customer service videos) or integrated into larger dialogue systems for emotionally aware conversation. When fine-tuning, users should follow the same prompt format and task prefix convention.
Note: The model is not intended for:
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. We recommend: