Downloads · 30 days
32
53% of all-time downloads
Weidows/WeMM-Embedding-2B-FP8
WeMM-Embedding-2B-FP8 is a image-text-to-text model from Weidows. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
FP8 (8-bit float, E4M3) quantization of tencent/WeMM-Embedding-2B, intended for vLLM / SGLang-class backends with FP8 support (RTX 4090 / Ada and newer have native FP8 tensor cores).
Downloads · 30 days
32
53% of all-time downloads
All-time downloads
60
Public
Repo size
3.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors3.2 GB · 99%
From the Hugging Face model README
FP8 (8-bit float, E4M3) quantization of tencent/WeMM-Embedding-2B, intended for vLLM / SGLang-class backends with FP8 support (RTX 4090 / Ada and newer have native FP8 tensor cores).
This is a separate repo from the GGUF build — GGUF targets llama.cpp; this FP8 build targets GPU inference servers.
model.fp8.safetensors — weights stored as float8_e4m3fn (per-tensor scale).fp8_scales.json — per-layer dequant scale (layer name -> scalar).WeMMEmbedding modeling, chat templates) are mirrored from the base model so AutoModel can load it.Note: the saved weights are raw fp8 + scale. To run, dequantize at load time (fp8 -> bf16) or serve through a backend that natively consumes fp8. See Usage below.
All numbers use the same engine (transformers / torch AutoModel) for both the BF16 baseline and the FP8 model, so Δρ is pure FP8 rounding error.
| Model | Bits/Weight | STS-B Spearman ρ | Δρ vs BF16 | Emb Cosine vs BF16 |
|---|---|---|---|---|
| BF16 | 16.00 | 0.8124 | — | — |
| FP8 (manual per-tensor E4M3) | 8.00 | 0.8114 | +0.12% | 0.9987 |
Metrics:
Conclusion: FP8 causes negligible quality loss (Δρ = +0.12%, Emb Cosine = 0.999) on STS-B. This is the recommended format when serving on FP8-capable GPUs.
import torch, json, safetensors.torch as st
from transformers import AutoModel, AutoProcessor
sd = st.load_file("model.fp8.safetensors")
scales = json.load(open("fp8_scales.json"))
for k, s in scales.items(): # k like 'model.xxx.weight'
sd[k] = (sd[k].to(torch.float32) * s).to(torch.bfloat16) # dequant
# save a runnable bf16 copy, or load directly:
model = AutoModel.from_pretrained(".", trust_remote_code=True, dtype=torch.bfloat16)
These backends expect a compressed-tensors / native fp8 checkpoint. The raw fp8+scale layout here is not yet wrapped for direct vLLM loading; repackaging into the backend's fp8 format (or quantizing the base model with the backend's own fp8 path) is required before serving. The quality numbers above already demonstrate the FP8 format itself is near-lossless.
llmcompressor oneshot did not actually quantize the qwen3_5 custom Linear layers, so the manual path is used for the reported numbers.