Downloads · 30 days
49
5% of all-time downloads
Backup-bdg/Xoron-Dev-MultiMoe
Xoron-Dev-MultiMoe is a any-to-any model from Backup-bdg. Use it for the any-to-any task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
49
5% of all-time downloads
All-time downloads
989
Public
Parameters
4.6B
237 GB on disk
Likes
8
Public
Click a slice to open those files.
.safetensors9.2 GB · 64%
From the Hugging Face model README
Xoron-Dev
Xoron-Dev is the definitive open-source architecture for Omni-Modal Artificial Intelligence. Unlike legacy models that treat vision and audio as plugins, Xoron-Dev is designed for native, high-fidelity perception across every major sensory dimension.
Xoron-Dev represents a massive leap in multimodal reasoning, combining cutting-edge Sparse MoE architecture with a refined sensory stack.
Xoron-Dev exclusively uses SigLIP-2 for superior zero-shot performance and semantic alignment.
Our custom VidTok encoder uses 3D Volumetric Compression to ingest up to 32 frames of high-definition video natively. Xoron doesn't just see a sequence of images—it understands motion, causality, and temporal context.
Xoron-Dev processes Raw 16kHz PCM Audio directly. No Mel Spectrograms, no lossy Fourier transforms.
A sophisticated Mixture of Experts (MoE) backbone that dynamically routes the logic of every token through specialized hardware-aware sub-networks.
Unlike standard MoE models with uniform experts, Xoron-Dev implements a specialized Deep Expert system.
Xoron-Dev features a dedicated FastPonderBlock for near-instant latent deliberation.
halt_head monitors latent entropy. Once the model reaches a decision (entropy threshold < 0.2), it breaks the ponder loop and returns to token decoding, reducing unnecessary FLOPs by up to 90%.Using Ring Attention, Xoron-Dev can analyze books, hour-long videos, or massive codebases with native 128K context window support.
The easiest way to experience Xoron-Dev is via the xorfice engine—the SOTA orchestrator for multimodal deployment.
pip install xorfice
from xorfice import XoronEngine
# The engine automatically handles weights and optimizations
# Correct model slug: Backup-bdg/Xoron-Dev-MultiMoe
engine = XoronEngine(model_path="Backup-bdg/Xoron-Dev-MultiMoe")
# Start an omni-modal conversation
response = engine.generate(
prompt="Who is this person and what are they doing?",
images="https://example.com/interview.jpg",
videos="https://example.com/interview.mp4"
)
print(response["text"])
| Feature | Xoron-Dev |
|---|---|
| Vision Backbone | SigLIP-2 |
| Video Compression | VidTok 3D |
| Audio Ingestion | Raw PCM |
| Inference Efficiency | Sparse MoE (5B) |
| Context Window | 128K (Ring) |
Fully integrated with MobileDiffusion, Xoron-Dev doesn't just understand—it creates.
Xoron-Dev is more than a model—it's a vision for the future of AI. Build your own multimodal agent today.
Powered by Xoron-Dev Team