Downloads · 30 days
720
49% of all-time downloads
prithivMLmods/Video-ORA-4B-GGUF
Video-ORA-4B-GGUF is a video-text-to-text model from prithivMLmods. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Video-ORA-4B is a 4-billion-parameter unified video understanding model built on Qwen3.5-4B and post-trained with OraRL (Annotations as Rollouts) — an annotation-augmented, on-policy reinforcement learning method that…
Downloads · 30 days
720
49% of all-time downloads
All-time downloads
1.5K
Public
Repo size
46.9 GB
Likes
2
Public
Click a slice to open those files.
.gguf47.6 GB · 100%
From the Hugging Face model README
Video-ORA-4B is a 4-billion-parameter unified video understanding model built on Qwen3.5-4B and post-trained with OraRL (Annotations as Rollouts) — an annotation-augmented, on-policy reinforcement learning method that equips a single compact model to handle seven task families with direct, task-native answers and no chain-of-thought decoding: temporal grounding, visual tracking, image/video segmentation, spatial grounding, spatial-temporal grounding, video question answering, and spatial intelligence. It retains a 262,144-token native context window from its base model and delivers strong compact-model results in matched seven-family benchmark comparisons against multimodal baselines despite skipping reasoning traces at inference, with evaluations run using direct-answer prompts (
enable_thinking=False) to match its reported protocol. The model is served via vLLM (with aqwen3reasoning parser and configurable video frame sampling) or a Transformers serving endpoint, occupies roughly 8.6 GiB in BF16 weight loading, and is intended for research on structured video/spatial perception, benchmark evaluation, and task-specific adaptation — explicitly out of scope for safety-critical decisions, identity inference, or surveillance deployment. Trained on public dataset splits with evaluation identities and media excluded from the training mixture, it is released under the Apache License 2.0, consistent with its Qwen3.5-4B base, and serves as the smaller, more deployment-friendly sibling to Video-ORA-9B in the OraRL model family.
[!NOTE] Model: https://huggingface.co/OraRL/Video-ORA-4B
| File Name | Quant Type | File Size | File Link |
|---|---|---|---|
| Video-ORA-4B.BF16.gguf | BF16 | 9.7 GB | Download |
| Video-ORA-4B.F16.gguf | F16 | 9.7 GB | Download |
| Video-ORA-4B.Q3_K_L.gguf | Q3_K_L | 2.69 GB | Download |
| Video-ORA-4B.Q3_K_M.gguf | Q3_K_M | 2.54 GB | Download |
| Video-ORA-4B.Q3_K_S.gguf | Q3_K_S | 2.34 GB | Download |
| Video-ORA-4B.Q4_0.gguf | Q4_0 | 2.9 GB | Download |
| Video-ORA-4B.Q4_K_M.gguf | Q4_K_M | 3.07 GB | Download |
| Video-ORA-4B.Q4_K_S.gguf | Q4_K_S | 2.92 GB | Download |
| Video-ORA-4B.Q5_0.gguf | Q5_0 | 3.43 GB | Download |
| Video-ORA-4B.Q5_K_M.gguf | Q5_K_M | 3.51 GB | Download |
| Video-ORA-4B.Q5_K_S.gguf | Q5_K_S | 3.43 GB | Download |
| Video-ORA-4B.mmproj-bf16.gguf | mmproj-bf16 | 676 MB | Download |
| Video-ORA-4B.mmproj-f16.gguf | mmproj-f16 | 676 MB | Download |
LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp