Downloads · 30 days
11
24% of all-time downloads
viavicdev/Tarsier2-Recap-7b-MLX-4bit
Tarsier2-Recap-7b-MLX-4bit is a video-text-to-text model from viavicdev. Use it for the video-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
A 4-bit MLX conversion of omni-research/Tarsier2-Recap-7b, prepared for native inference on Apple Silicon.
Downloads · 30 days
11
24% of all-time downloads
All-time downloads
46
Public
Parameters
8.3B
5.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.6 GB · 100%
How the weights are stored.
U327.6B · 92%
From the Hugging Face model README
A 4-bit MLX conversion of omni-research/Tarsier2-Recap-7b, prepared for native inference on Apple Silicon.
Tarsier2-Recap-7b produces unusually specific, grounded video descriptions — it tends to name concrete details rather than summarise a scene generically. No official MLX build exists. This repository provides one: the same weights, quantized to 4-bit and run on the Mac GPU through MLX.
The underlying model is in the Qwen2-VL 7B class. Hugging Face may display a lower automatic parameter count (~1.9 B) because MLX 4-bit weights are packed into 32-bit integers.
Verified with mlx-vlm 0.6.8 on macOS
15.7.2, Python 3.12, mlx 0.32.0.
pip install mlx-vlm==0.6.8
brew install ffmpeg # for extracting frames
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
model, processor = load("viavicdev/Tarsier2-Recap-7b-MLX-4bit")
# Provide the video as a list of extracted frame images (see note 1)
frames = ["frame_01.jpg", "frame_02.jpg", "frame_03.jpg", "frame_04.jpg"]
prompt = apply_chat_template(
processor, model.config,
"Describe this video in detail.",
num_images=len(frames), # must match len(frames)
)
result = generate(
model, processor, prompt,
image=frames, # parameter is image= (not images=)
max_tokens=300,
temperature=0.3,
repetition_penalty=1.3, # see note 3
repetition_context_size=40,
)
print(result.text) # .text — generate() returns a GenerationResult
Extract frames with ffmpeg, for example eight evenly spaced frames from a clip
of known duration D:
ffmpeg -i clip.mp4 -vf "fps=8/D,scale=-2:448" -frames:v 8 -q:v 2 frame_%02d.jpg
mlx-vlm video decoder path
is unreliable for this model and can produce output unrelated to the clip.
Extracting frames with ffmpeg and passing them as a list (image=[...] with
a matching num_images) delivers the actual frames to the model.chat_template.json
and chat_template.jinja). Some MLX repackings omit it, which makes
apply_chat_template fail.repetition_penalty=1.3, repetition_context_size=40 and
temperature=0.3 kept output stable in our testing; in a small internal set
of 22 short clips, one degenerated without a penalty and none did with it.A single measured run, not a benchmark.
mlx 0.32.0, mlx-vlm 0.6.8max_tokens=300, as in the usage example aboveOutput (run 2 of 3, verbatim):
A model train, consisting of a green engine and several red freight cars, moves along the tracks from left to right. The background features a large crowd of people gathered in an open area with numerous colorful cars parked and displayed. The train passes by the car display area multiple times, moving smoothly along the tracks.
For comparison, 8 frames at 252×448 (a vertical clip) on the same machine ran in a median of 5.5 s. Frame resolution dominates runtime — budget accordingly.
The original checkpoint could not be converted directly with the standard
mlx-vlm workflow, because its published architecture metadata does not fully
match the underlying multimodal model structure.
For this release, the checkpoint was adapted to a compatible native MLX layout before 4-bit quantization. This required resolving differences in the vision-language architecture and in the positional encoding configuration.
The resulting checkpoint runs through the standard native inference path in
mlx-vlm. The conversion changes the storage and runtime format only; the model
was not retrained or fine-tuned.
The complete conversion procedure is not included in this repository.
This conversion inherits the Apache-2.0 license of the base model.
All credit for the original model, training and weights belongs to the Tarsier2 authors (omni-research/Tarsier2-Recap-7b).
This repository provides the MLX format conversion, 4-bit quantization, packaging, Apple Silicon compatibility testing and usage documentation. The model was not retrained or fine-tuned.
When citing Tarsier2, cite the original authors.