Downloads · 30 days
0
Obscure-Entropy/MangaliCa
MangaliCa is a image-to-text model from Obscure-Entropy. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as cc-by-nc-sa-4.0.
<div align="center" <img src="./assets/mangalica.jpg" width="50%" / </div
Downloads · 30 days
0
Access
Public
Updated Jan 12, 2026
Repo size
648 GB
Likes
0
Public
Click a slice to open those files.
.safetensors599 GB · 100%
From the Hugging Face model README
MangaliCa is the first publicly available Hungarian–English bilingual vision–language model designed for image captioning and image–text retrieval.
The model is built on the CoCa (Contrastive Captioner) framework and jointly optimizes contrastive alignment and autoregressive caption generation across two languages.
MangaliCa integrates:
[!CAUTION] The model was trained on a newly constructed 70M-sample Hungarian–English bilingual image–caption dataset, the largest multimodal dataset involving Hungarian to date.
Total parameters: ~1.8B
Trainable parameters (LoRA): ~15M
hu), English (en)MangaliCa was evaluated on multiple benchmarks with Hungarian translations:
| Dataset | R@1 | R@3 | R@5 | R@25 | R@100 | NDCG@1 | NDCG@10 | NDCG@100 | MRR |
|---|---|---|---|---|---|---|---|---|---|
| GBC-10M | 35.6% | 60.0% | 70.0% | 91.0% | 98.6% | 35.6% | 57.5% | 61.4% | 0.51 |
| MS-COCO | 6.05% | 12.2% | 17.3% | 43.5% | 69.3% | 6.05% | 14.4% | 23.3% | 0.13 |
| text-to-image-2M | 41.5% | 62.7% | 72.6% | 91.7% | 98.7% | 41.5% | 61.0% | 64.6% | 0.55 |
| XM3600 | 11.3% | 22.5% | 28.9% | 53.8% | 76.9% | 11.3% | 23.4% | 31.4% | 0.20 |