Downloads · 30 days
15
18% of all-time downloads
cabbagel/caT-MDC
caT-MDC is a image-text-to-text model from cabbagel. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
caT-MDC is the model submitted by team caT to the MDC track of the MARS2 2026 Challenge.
Downloads · 30 days
15
18% of all-time downloads
All-time downloads
85
Public
Parameters
9.4B
18.8 GB on disk
Likes
1
Trending 1
Click a slice to open those files.
.safetensors18.8 GB · 100%
From the Hugging Face model README
caT-MDC is the model submitted by team caT to the MDC track of the
MARS2 2026 Challenge.
The model uses the Qwen3.5-9B multimodal architecture and was post-trained with the team's cold-start and group-based reinforcement-learning pipeline. This repository contains the complete merged model in Hugging Face Transformers format rather than a LoRA adapter.
| Item | Description |
|---|---|
| Team | caT |
| Challenge | MARS2 2026 |
| Track | MDC |
| Backbone | Qwen3.5-9B |
| Architecture | Qwen3_5ForConditionalGeneration |
| Model type | Multimodal vision-language model |
| Weight format | Safetensors |
| Precision | BFloat16 |
| Training stage | Cold-start post-training followed by GSPO-stage reinforcement learning |
| Release format | Complete merged model |
The released checkpoint is the selected MDC submission model. According to the
archived configuration in args.json, its reinforcement-learning stage used:
1e-514288048The competition training dataset is not redistributed in this model repository.
config.json: model architecture and configurationgeneration_config.json: default generation configurationmodel-*.safetensors: sharded model weightsmodel.safetensors.index.json: weight indexpreprocessor_config.json: multimodal preprocessing configurationprocessor_config.json: processor configurationtokenizer.json: tokenizertokenizer_config.json: tokenizer configurationchat_template.jinja: conversation templateargs.json: archived training argumentspip install -U transformers accelerate pillow
Qwen3.5 requires a recent Transformers version. Refer to the official Qwen3.5-9B model card for current compatibility guidance.
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
model_id = "cabbagel/caT-MDC"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
model_id,
dtype="auto",
device_map="auto",
)
print(model.__class__.__name__)
Expected model class:
Qwen3_5ForConditionalGeneration
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
model_id = "cabbagel/caT-MDC"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
model_id,
dtype="auto",
device_map="auto",
)
messages = [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Briefly describe your multimodal reasoning capabilities.",
}
],
}
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=256)
output_ids = generated_ids[:, inputs["input_ids"].shape[1]:]
response = processor.batch_decode(
output_ids,
skip_special_tokens=True,
)[0]
print(response)
For image and video inputs, follow the multimodal message format documented in the official Qwen3.5 model card.
This model is released for:
This work builds on Qwen3.5-9B. We thank the Qwen team and the MARS2 2026 organizers.