Downloads · 30 days
15
17% of all-time downloads
Jetlink/JetLLMPremium-v1.0-397B-A17B
JetLLMPremium-v1.0-397B-A17B is a image-text-to-text model from Jetlink. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
JetLLMPremium-3.5 is a large-scale multimodal Mixture-of-Experts model published by Jetlink.
Downloads · 30 days
15
17% of all-time downloads
All-time downloads
90
Public
Parameters
403B
807 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors807 GB · 100%
How the weights are stored.
BF16403B · 100%
From the Hugging Face model README
JetLLMPremium-3.5 is a large-scale multimodal Mixture-of-Experts model published by Jetlink.
It is intended for teams that want to manage deployment, access, and internal distribution from their own namespace while preserving compatibility with the original upstream model ecosystem.
JetLLMPremium-3.5 is a post-trained multimodal autoregressive model with:
This model is suitable for advanced workloads such as:
The upstream model card states compatibility with:
This model is not intended for lightweight local deployment in its original form.
The official upstream documentation does not define a single universal minimum hardware requirement, because actual requirements vary depending on:
The upstream Qwen model card provides official serving examples using tensor parallelism across 8 GPUs for both the standard and FP8 variants.
Recommended reference configurations for self-hosted deployment:
--language-model-onlyNote: hardware requirements vary significantly based on precision, context length, KV cache settings, batch size, and whether multimodal inputs are enabled. The configurations above should be treated as practical deployment references rather than universal minimum requirements.
For the original model weights, this model should be treated as a high-end multi-GPU server model.
A practical baseline is:
The upstream model card provides serving examples using tensor parallelism across 8 GPUs for both SGLang and vLLM.
That makes 8-GPU deployment the safest reference point to document for standard, non-quantized serving.
The upstream vLLM example also documents a --language-model-only option, which skips the vision encoder and multimodal profiling to free more memory for KV cache. This can be useful when your workload is purely text-based.
For most production teams:
Recommended environment:
Common dependencies may include:
torchtransformerstorchvisionpillowAdditional runtime-specific packages may be required depending on your serving framework.
Install Transformers:
pip install "transformers[serving] @ git+https://github.com/huggingface/transformers.git@main"
Basic loading example:
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "Jetlink/JetLLMPremium-3.5"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
trust_remote_code=True,
)
vllm serve Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3
vllm serve Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
vllm serve Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--language-model-only
python -m sglang.launch_server \
--model-path Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tp-size 8 \
--mem-fraction-static 0.8 \
--context-length 262144 \
--reasoning-parser qwen3
python -m sglang.launch_server \
--model-path Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tp-size 8 \
--mem-fraction-static 0.8 \
--context-length 262144 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder
python -m sglang.launch_server \
--model-path Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tp-size 8 \
--mem-fraction-static 0.8 \
--context-length 262144 \
--reasoning-parser qwen3 \
--speculative-algo NEXTN \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4
JetLLMPremium-3.5 natively supports 262,144 tokens.
For tasks that exceed this window, the upstream documentation recommends using long-context scaling techniques such as YaRN, which are supported in several frameworks including Transformers, vLLM, KTransformers, and SGLang.
As with other frontier-scale multimodal language models, outputs should be reviewed before use in:
Human review, policy controls, and tool-level validation are strongly recommended.
This repository follows the same license as the upstream release.
If you redistribute, fine-tune, quantize, or otherwise modify this model, make sure your usage remains compliant with the upstream license and attribution requirements.
Original model and research release by the Qwen team.
Upstream model:
Qwen/Qwen3.5-397B-A17BThis repository is an organization-managed copy and is not the original upstream source.
Please cite the original Qwen release when using this model in research, evaluation, or production documentation.
@misc{qwen3.5,
title = {Qwen3.5 Technical Report},
author = {Qwen Team},
year = {2026},
publisher = {Alibaba Cloud},
howpublished = {\url{https://huggingface.co/Qwen/Qwen3.5-397B-A17B}}
}
JetLLMPremium-3.5, Jetlink tarafından yayınlanan büyük ölçekli bir multimodal Mixture-of-Experts modelidir.
Bu depo; modeli kendi namespace'i altında yönetmek, erişimi kontrol etmek ve dağıtımı kolaylaştırmak isteyen ekipler için hazırlanmıştır. Amaç, upstream model ekosistemiyle uyumluluğu koruyarak kurumsal kullanım sağlamaktır.
JetLLMPremium-3.5, aşağıdaki özelliklere sahip, post-train edilmiş çok kipli (multimodal) otoregresif bir modeldir:
Bu model aşağıdaki gelişmiş kullanım senaryoları için uygundur:
Upstream model kartına göre model şu ekosistemlerle uyumludur:
Bu model, orijinal haliyle hafif yerel kullanım için tasarlanmamıştır.
Resmi upstream dokümantasyonu tek ve evrensel bir minimum donanım gereksinimi belirtmez. Çünkü gerçek ihtiyaçlar şunlara bağlı olarak değişir:
Upstream Qwen model kartı, hem standart hem de FP8 varyantları için 8 GPU üzerinde tensor parallelism kullanan resmi serving örnekleri sunmaktadır.
Self-hosted deployment için önerilen referans konfigürasyonlar:
--language-model-only ile devre dışı bırakılarak bellek baskısı azaltılabilirNot: donanım gereksinimleri; precision, bağlam uzunluğu, KV cache ayarları, batch size ve multimodal girişlerin açık olup olmamasına göre ciddi şekilde değişir. Bu nedenle yukarıdaki konfigürasyonlar evrensel minimum gereksinim olarak değil, pratik dağıtım referansı olarak değerlendirilmelidir.
Orijinal model ağırlıklarıyla kullanımda bu model, yüksek seviye çoklu GPU sunucu modeli olarak düşünülmelidir.
Pratik bir başlangıç seviyesi şunları içerir:
Upstream model kartında hem SGLang hem de vLLM için 8 GPU üzerinde tensor parallelism kullanan serving örnekleri bulunmaktadır.
Bu nedenle standart, quantize edilmemiş serving için dokümante edilebilecek en güvenli referans noktası 8 GPU dağıtımıdır.
Upstream vLLM örneğinde ayrıca --language-model-only seçeneği de yer alır. Bu seçenek vision encoder'ı ve multimodal profiling'i devre dışı bırakarak KV cache için daha fazla bellek açar. Sadece metin tabanlı iş yüklerinde faydalı olabilir.
Çoğu production ekip için en mantıklı yaklaşım:
Önerilen ortam:
Yaygın bağımlılıklar şunları içerebilir:
torchtransformerstorchvisionpillowKullandığınız serving framework'üne göre ek bağımlılıklar gerekebilir.
Transformers kurulumu:
pip install "transformers[serving] @ git+https://github.com/huggingface/transformers.git@main"
Temel yükleme örneği:
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "Jetlink/JetLLMPremium-3.5"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
trust_remote_code=True,
)
vllm serve Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3
vllm serve Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
vllm serve Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--language-model-only
python -m sglang.launch_server \
--model-path Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tp-size 8 \
--mem-fraction-static 0.8 \
--context-length 262144 \
--reasoning-parser qwen3
python -m sglang.launch_server \
--model-path Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tp-size 8 \
--mem-fraction-static 0.8 \
--context-length 262144 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder
python -m sglang.launch_server \
--model-path Jetlink/JetLLMPremium-3.5 \
--port 8000 \
--tp-size 8 \
--mem-fraction-static 0.8 \
--context-length 262144 \
--reasoning-parser qwen3 \
--speculative-algo NEXTN \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4
JetLLMPremium-3.5 yerel olarak 262.144 token destekler.
Bu pencereyi aşan görevlerde upstream dokümantasyonu; Transformers, vLLM, KTransformers ve SGLang gibi framework'ler tarafından desteklenen YaRN benzeri uzun bağlam ölçekleme tekniklerini önermektedir.
Diğer frontier-scale multimodal language model'lerde olduğu gibi, model çıktıları şu alanlarda insan denetimi olmadan kullanılmamalıdır:
İnsan incelemesi, politika kontrolleri ve tool seviyesinde doğrulama güçlü şekilde önerilir.
Bu depo, upstream sürümle aynı lisansı takip eder.
Modeli yeniden dağıtıyor, fine-tune ediyor, quantize ediyor veya başka şekilde değiştiriyorsan; kullanımının upstream lisans ve attribution gereklilikleriyle uyumlu olduğundan emin olmalısın.
Orijinal model ve araştırma yayını Qwen ekibine aittir.
Upstream model:
Qwen/Qwen3.5-397B-A17BBu depo, kurum tarafından yönetilen bir kopyadır ve orijinal upstream kaynak değildir.
Bu modeli araştırma, değerlendirme veya production dokümantasyonunda kullanıyorsan, lütfen orijinal Qwen sürümüne atıf yap.
@misc{qwen3.5,
title = {Qwen3.5 Technical Report},
author = {Qwen Team},
year = {2026},
publisher = {Alibaba Cloud},
howpublished = {\url{https://huggingface.co/Qwen/Qwen3.5-397B-A17B}}
}