Downloads · 30 days
68
26% of all-time downloads
Jiaha0Hu4ng/OneEmo
OneEmo is a video-text-to-text model from Jiaha0Hu4ng. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
<p align="center" <img src="https://cdn-uploads.huggingface.co/production/uploads/66953e2f417bbfcd51933cd0/0XVrXOZhxIMCYTwecNa.png" width="280"/ </p <p align="center" <a href="https://arxiv.org/abs/2608.06013"<img src…
Downloads · 30 days
68
26% of all-time downloads
All-time downloads
266
Public
Parameters
4.5B
9.1 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors9.1 GB · 100%
From the Hugging Face model README
The model was trained and evaluated in the following environment:
torch: 2.10.0+cu128torchaudio: 2.10.0+cu128torchvision: 0.25.0+cu128transformers: 5.2.0vllm: 0.19.0ms-swift: 4.1.3 (optional, for alternative inference engine)oneemo)Use the code below to get started with the model.
vllm serve Jiaha0Hu4ng/OneEmo \
--host 0.0.0.0 \
--port 8000 \
--served-model-name OneEmo \
--max-model-len 32768 \
--mm-encoder-tp-mode data \
--media-io-kwargs '{"video": {"num_frames": 16}}' \
--default-chat-template-kwargs '{"enable_thinking": true}' \
--reasoning-parser qwen3
import torch
from swift import get_model_processor, get_template
from swift.infer_engine import TransformersEngine, InferRequest, RequestConfig
model_path = "Jiaha0Hu4ng/OneEmo"
model, processor = get_model_processor(
model_path,
model_type="qwen3_5",
torch_dtype=torch.bfloat16,
attn_impl="flash_attn",
)
template = get_template(processor, enable_thinking=True)
engine = TransformersEngine(model, template=template)
request_config = RequestConfig(max_tokens=2048, temperature=0.7)
infer_request = InferRequest(
messages=[{
"role": "user",
"content": (
"Based on the video, describe the range of emotions the character "
"may be feeling."
)
}],
videos=["path/to/your/video.mp4"]
)
resp_list = engine.infer([infer_request], request_config=request_config)
print(resp_list[0].choices[0].message.content)
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
from qwen_vl_utils import process_vision_info
model_path = "Jiaha0Hu4ng/OneEmo"
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2", # fallback to "sdpa" if flash-attn is not installed
device_map="auto",
)
processor = AutoProcessor.from_pretrained(model_path)
# Example: open-vocabulary multimodal emotion recognition (OVMER)
messages = [
{
"role": "user",
"content": [
{
"type": "video",
"video": "path/to/your/video.mp4",
"max_pixels": 360 * 420,
"fps": 2.0,
},
{
"type": "text",
"text": "Based on the video, describe the range of emotions the character may be feeling.",
},
],
}
]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
text=[text],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
).to(model.device)
with torch.no_grad():
generated_ids = model.generate(
**inputs,
max_new_tokens=2048,
temperature=0.7,
top_p=0.9,
top_k=50,
)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text[0])
For other tasks, replace the prompt text with the corresponding template (see the full model card on Hugging Face for the complete prompt table). The model also supports deployment with vLLM (OpenAI-compatible API) and inference via ms-swift; refer to the repository for those examples.
We use Emo-World-130K and Emo-Chord strategy to train OneEmo-Base and OneEmo.
BibTeX:
@article{huang2026oneemo,
title={OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction},
author={Jiahao Huang and Zheng Lian and Jingyi Zhang and Zhide Chen and Xiaojiang Peng and Shaonan Wang},
year={2026},
journal={arXiv preprint arXiv:2608.06013},
}
OneEmo can be used directly for video-based emotion analysis, intention recognition, sentiment classification, humour/sarcasm detection, and for generating empathetic or supportive textual responses. It is suitable for researchers and developers building affective computing applications, mental-health chatbots, or social signal processing tools.
The model can be fine-tuned on domain-specific datasets (e.g., clinical interviews, customer service videos) or integrated into larger dialogue systems for emotionally aware conversation. When fine-tuning, users should follow the same prompt format and task prefix convention.
Note: The model is not intended for:
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. We recommend: