Downloads · 30 days
16
5% of all-time downloads
digitranslab/Megamind-v2-VL-high
Megamind-v2-VL-high is a image-text-to-text model from digitranslab. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Megamind-v2-VL is an 8B-parameter vision–language model for long-horizon, multi-step tasks in real software environments (e.g., browsers and desktop apps). It combines language reasoning with visual perception to foll…
Downloads · 30 days
16
5% of all-time downloads
All-time downloads
323
Public
Parameters
8.8B
17.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors17.5 GB · 100%
From the Hugging Face model README
Megamind-v2-VL is an 8B-parameter vision–language model for long-horizon, multi-step tasks in real software environments (e.g., browsers and desktop apps). It combines language reasoning with visual perception to follow complex instructions, maintain intermediate state, and recover from minor execution errors.
We recognize the importance of long-horizon execution for real-world tasks, where small per-step gains compound into much longer successful chains—so Megamind-v2-VL is built for stable, many-step execution. For evaluation, we use The Illusion of Diminishing Returns: Measuring Long-Horizon Execution in LLMs, which measures execution length. This benchmark aligns with public consensus on what makes a strong coding model—steady, low-drift step execution—suggesting that robust long-horizon ability closely tracks better user experience.
Variants
Tasks where the plan and/or knowledge can be provided up front, and success hinges on stable, many-step execution with minimal drift:
Compared with its base (Qwen-3-VL-8B-Thinking), Megamind-v2-VL shows no degradation on standard text-only and vision tasks—and is slightly better on several—while delivering stronger long-horizon execution on the Illusion of Diminishing Returns benchmark.
Megamind-v2-VL is optimized for direct integration with the Megamind. Simply select the model from the Megamind interface for immediate access to its full capabilities.
Using vLLM:
vllm serve digitranslab/Megamind-v2-VL-high \
--host 0.0.0.0 \
--port 1234 \
--enable-auto-tool-choice \
--tool-call-parser hermes \
--reasoning-parser qwen3
Using llama.cpp:
llama-server --model Megamind-v2-VL-high-Q8_0.gguf \
--vision-model-path mmproj-Megamind-v2-VL-high.gguf \
--host 0.0.0.0 \
--port 1234 \
--jinja \
--no-context-shift
For optimal performance in agentic and general tasks, we recommend the following inference parameters:
temperature: 1.0
top_p: 0.95
top_k: 20
repetition_penalty: 1.0
presence_penalty: 1.5
Updated Soon