Downloads · 30 days
6
17% of all-time downloads
furproxy/9b-107
9b-107 is a image-text-to-text model from furproxy. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as other.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
6
17% of all-time downloads
All-time downloads
35
Public
Parameters
9.4B
20.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors20.9 GB · 100%
How the weights are stored.
BF168.4B · 89%
From the Hugging Face model README
This model is a fine-tuned version of /workspace/models/Qwen3.5-9B on the my_caption dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
family_to_muon_lr = { "language": _fallback(getattr(training_args, "language_muon_lr", 1e-1), language_lr), "vision": _fallback(getattr(training_args, "vision_muon_lr", 2e-5), vision_lr), "merger": _fallback(getattr(training_args, "merger_muon_lr", 2e-4), merger_lr), }
family_to_adamw_lr = { "language": _fallback(getattr(training_args, "language_adamw_lr", 8e-6), language_lr), "vision": _fallback(getattr(training_args, "vision_adamw_lr", 5e-6), vision_lr), "merger": _fallback(getattr(training_args, "merger_adamw_lr", 1e-5), merger_lr), }
train_batch_size: 3
eval_batch_size: 8
seed: 42
distributed_type: multi-GPU
gradient_accumulation_steps: 10
total_train_batch_size: 30
optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
lr_scheduler_type: cosine_with_min_lr
lr_scheduler_warmup_steps: 0.05
num_epochs: 4