Downloads · 30 days
18
30% of all-time downloads
VextLabsinc/juwel-beryl
juwel-beryl is a image-text-to-text model from VextLabsinc. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
80-layer vision-language base. The required base for all 16 JUWEL GEM adapters.
Downloads · 30 days
18
30% of all-time downloads
All-time downloads
61
Public
Parameters
41.2B
82.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors82.3 GB · 100%
From the Hugging Face model README
80-layer vision-language base. The required base for all 16 JUWEL GEM adapters.
Read this first: safety refusals have been removed
Beryl is abliterated — refusal directions were ablated, so it will not reliably decline harmful requests. It ships with no safety layer. If you deploy it, you are responsible for adding your own filtering, policy, and human oversight. Do not treat this as an aligned assistant model. See
USE_POLICY.md.
| Architecture | Qwen3VLForConditionalGeneration (image-text-to-text) |
| Layers | 80 (parent had 64; 16 added by CIP, parent's 64 frozen) |
| Hidden size | 5120 |
| Attention heads | 64 (head dim 128) · 8 KV heads |
| Intermediate size | 25600 |
| Context window | 262144 tokens (256K), rope_type: default — no extended-rope variant |
| Vision tower | 27 layers, inherited from the parent |
| Vocab | 151936 — Qwen's tokenizer, shipped unmodified |
| Precision | BF16 (verified in the shard headers, not just declared) |
| Size | 76.7 GiB, 19 shards, 1234 tensors |
| Parent | Qwen3-VL-32B-Instruct (Apache-2.0) |
| License | Apache-2.0 — see LICENSE + NOTICE. AS IS, no warranty. |
Internal lineage: theron-base-v9 plus one capability-injection rung. Layers 0–63 are
the parent's frozen weights; layers 64–79 were added and trained.
from transformers import AutoModelForImageTextToText, AutoTokenizer
model = AutoModelForImageTextToText.from_pretrained(
"VextLabsinc/juwel-beryl", torch_dtype="auto", device_map="auto")
tok = AutoTokenizer.from_pretrained("VextLabsinc/juwel-beryl")
Use AutoModelForImageTextToText, not AutoModelForCausalLM — this is a
vision-language architecture and the causal-LM class will not load it.
Footprint: ~82 GB in BF16 (2x 48 GB GPUs or better), ~41 GB at FP8, ~21 GB at NF4.
Beryl is the base for 16 domain LoRA adapters (VextLabsinc/gem-*: ruby/code,
aquamarine/math, spinel/reasoning, lapis/language, bloodstone/medical,
sardonyx/legal, tanzanite/finance, peridot/science, topaz/engineering,
citrine/business, opal/education, jade/humanities, pearl/reconciler, plus
alexandrite, tourmaline, turquoise). Each is native to this exact 80-layer geometry
(layers_to_transform 0–79, r=64) and is not interchangeable with other JUWEL
bases.
from peft import PeftModel
model = PeftModel.from_pretrained(model, "VextLabsinc/gem-ruby")
Status: PENDING. No benchmark numbers are published for this model yet, and that is a real absence rather than a hidden result.
Until 2026-07-28 the GEM adapters named a 144-layer base in their cards while being native to 80 layers, so nothing in the release was loadable as trained and no valid benchmark could be produced. Publishing Beryl is what makes the fleet runnable. Structural compatibility is verified — all 560 target modules per adapter resolve against this model's tensor index — but output-quality evaluation requires GPU time and has not been completed.
When numbers arrive they will be BF16, measured on our own weights on our own hardware, with per-sample audit trails. We do not publish a score we cannot reproduce, and we will not present quantized screening runs as headline results.
Research, evaluation, and authorized professional workflows where a human remains responsible for outcomes, and where the deployer supplies their own safety layer.
See USE_POLICY.md. In short: no unauthorized system access, no malware or fraud,
no CSAM, no weapons development, nothing violating applicable law. Publishing weights
is not permission to break the law, and Vext Labs does not endorse misuse.
Provided AS IS under LICENSE. To the maximum extent permitted by law,
Vext Labs, Inc. disclaims all warranties and is not liable for damages arising from
use or misuse. You are solely responsible for lawful, authorized use.