Downloads · 30 days
439
35% of all-time downloads
TuTuCSF/WeMM-Embedding-2B-GGUF
WeMM-Embedding-2B-GGUF is a feature extraction model from TuTuCSF. Use it when you need embeddings to search or compare text. The card lists the license as apache-2.0.
Downloads · 30 days
439
35% of all-time downloads
All-time downloads
1.3K
Public
Repo size
7.4 GB
Likes
0
Public
Click a slice to open those files.
.gguf7.4 GB · 100%
From the Hugging Face model README
156 MMEB-v3 tasks · ~54 hours local evaluation · Q4_K_M + Q8_0 mmproj
<a id="english"></a>
GGUF conversion of tencent/WeMM-Embedding-2B, focused on local multimodal embedding inference with llama.cpp.
Recommended validated pair
WeMM-Embedding-2B-Q4_K_M.ggufmmproj-WeMM-Embedding-2B-Q8_0.gguf64, 128, 256, 512, 1024, 2048This repository is not only a conversion. The Q4_K_M + Q8_0 projector pair was evaluated end-to-end on 156 MMEB-v3 tasks, requiring approximately 54 hours of local evaluation.
Evaluation date: 2026-09-04
| Benchmark group | Metric | Tasks | Tencent official WeMM-2B | This GGUF | Δ | Score retained |
|---|---|---|---|---|---|---|
| MMEB-v2 Image | Hit@1 | 36 | 79.60 | 79.18 | -0.42 | 99.47% |
| MMEB-v2 VisDoc | NDCG@5 | 27 | 80.70 | 78.07 | -2.63 | 96.75% |
| MMEB-v3 Text | NDCG@5 | 53 | 45.30 | 43.65 | -1.65 | 96.37% |
| MMEB-v3 MCMR | Hit@1 | 1 | 42.50 | 41.01 | -1.49 | 96.49% |
The published Tencent scores above are the native WeMM-Embedding-2B aggregate results from the WeMM technical report. They are not a BF16 rerun on the same RTX 5080 machine, so the differences should be read as practical agreement with the published native baseline rather than a perfectly controlled quantization-only ablation.
The important result is that the Q4_K_M main model with a Q8_0 visual projector preserves roughly 96–99% of the published native score across the four directly comparable benchmark groups.
This run covers 156 / 190 MMEB-v3 tasks:
| Modality / group | Tasks | Metric | Local result |
|---|---|---|---|
| Image | 36 | Hit@1 | 79.18 |
| VisDoc | 27 | NDCG@5 | 78.07 |
| Text | 53 | NDCG@5 | 43.65 |
| Tool | 35 | Hit@1 | 49.00 |
| Memory | 4 | Hit@1 | 35.67 |
| MCMR | 1 | Hit@1 | 41.01 |
| Tool + Memory, no-GUI Agent subset | 39 | Hit@1 | 47.63 |
Not evaluated in this run:
The original WeMM model supports video, but video was intentionally excluded from this GGUF benchmark. Audio is not supported by WeMM-Embedding. Because the 8 GUI tasks were excluded, the local 39-task Tool + Memory result must not be compared directly with Tencent's published 47-task Agent aggregate.
The tables below separate published native-model results from this repository's local GGUF result. This matters because comparing a local Q4_K_M run directly with a native BF16/FP16 leaderboard submission is useful as a practical reference, but it is not a controlled quantization-only ablation.
Tencent reports the following results on the 78-task MMEB-v2 benchmark.
| Model | Size | AVG | Image Hit@1 | Video Hit@1 | VisDoc NDCG@5 |
|---|---|---|---|---|---|
| Qwen3-VL-Embedding | 2B | 73.2 | 75.0 | 61.9 | 79.2 |
| WeMM-Embedding | 2B | 77.9 | 79.6 | 70.8 | 80.7 |
| WeMM-Embedding | 4B | 79.2 | 80.8 | 72.1 | 82.0 |
| Qwen3-VL-Embedding | 8B | 77.8 | 80.1 | 67.1 | 82.4 |
| WeMM-Embedding | 9B | 80.6 | 81.9 | 74.3 | 83.3 |
At roughly the same 2B scale, official WeMM-Embedding-2B leads Qwen3-VL-Embedding-2B by:
| MMEB-v2 metric | WeMM-2B advantage |
|---|---|
| AVG | +4.7 |
| Image | +4.6 |
| Video | +8.9 |
| VisDoc | +1.5 |
At the larger scale, WeMM-Embedding-9B compared with Qwen3-VL-Embedding-8B:
| MMEB-v2 metric | WeMM-9B advantage |
|---|---|
| AVG | +2.8 |
| Image | +1.8 |
| Video | +7.2 |
| VisDoc | +0.9 |
The largest consistent WeMM advantage on MMEB-v2 is video retrieval, while the gap on visual-document retrieval is much smaller.
Tencent reports the following results on all 190 MMEB-v3 tasks.
| Model | Size | V3-All | Text NDCG@5 | Agent Hit@1 | MCMR Hit@1 | Audio |
|---|---|---|---|---|---|---|
| Qwen3-VL-Embedding | 2B | 50.9 | 39.2 | 39.3 | 42.0 | 0.0 |
| WeMM-Embedding | 2B | 56.0 | 45.3 | 45.1 | 42.5 | 0.0 |
| WeMM-Embedding | 4B | 58.2 | 47.9 | 49.0 | 41.9 | 0.0 |
| Qwen3-VL-Embedding | 8B | 53.5 | 42.5 | 38.4 | 38.0 | 0.0 |
| WeMM-Embedding | 9B | 59.5 | 48.8 | 51.0 | 49.3 | 0.0 |
At the 2B scale, official WeMM-Embedding-2B leads Qwen3-VL-Embedding-2B by:
| MMEB-v3 metric | WeMM-2B advantage |
|---|---|
| V3-All | +5.1 |
| Text | +6.1 |
| Agent | +5.8 |
| MCMR | +0.5 |
At the larger scale, WeMM-Embedding-9B compared with Qwen3-VL-Embedding-8B:
| MMEB-v3 metric | WeMM-9B advantage |
|---|---|
| V3-All | +6.0 |
| Text | +6.3 |
| Agent | +12.6 |
| MCMR | +11.3 |
This is the most relevant table for evaluating the quality of the conversion itself.
| Benchmark group | Official WeMM-2B | This 2B Q4_K_M GGUF | Δ | Score retained |
|---|---|---|---|---|
| Image | 79.60 | 79.18 | -0.42 | 99.47% |
| VisDoc | 80.70 | 78.07 | -2.63 | 96.75% |
| Text | 45.30 | 43.65 | -1.65 | 96.37% |
| MCMR | 42.50 | 41.01 | -1.49 | 96.49% |
Across these four directly comparable published benchmark groups, the tested Q4_K_M + Q8_0 projector pair retains approximately 96–99% of the published native WeMM-Embedding-2B score.
This comparison answers a practical deployment question: how does the locally quantized WeMM-2B compare with the published native Qwen3-VL-Embedding-2B baseline?
| Benchmark group | Qwen3-VL-Embedding-2B native | WeMM-2B Q4_K_M GGUF | GGUF difference |
|---|---|---|---|
| Image | 75.00 | 79.18 | +4.18 |
| VisDoc | 79.20 | 78.07 | -1.13 |
| Text | 39.20 | 43.65 | +4.45 |
| MCMR | 42.00 | 41.01 | -0.99 |
Even after Q4_K_M quantization of the main model, this WeMM-2B GGUF remains clearly ahead of the published Qwen3-VL-Embedding-2B native baseline on Image and Text, while trailing by about one point on VisDoc and MCMR.
This is not an apples-to-apples comparison, but it is useful for understanding the deployment efficiency of the 2B GGUF.
| Benchmark group | Qwen3-VL-Embedding-8B native | WeMM-2B Q4_K_M GGUF | GGUF difference |
|---|---|---|---|
| Image | 80.10 | 79.18 | -0.92 |
| VisDoc | 82.40 | 78.07 | -4.33 |
| Text | 42.50 | 43.65 | +1.15 |
| MCMR | 38.00 | 41.01 | +3.01 |
Notably, the tested 2B Q4_K_M WeMM GGUF exceeds the published Qwen3-VL-Embedding-8B score on Text and MCMR, while coming within one point on Image. Qwen3-VL-Embedding-8B remains substantially stronger on VisDoc.
This local run intentionally excludes Video and GUI, so several official aggregate columns cannot be reconstructed faithfully:
In short, use Image / VisDoc / Text / MCMR for direct published-score comparisons, and treat Tool / Memory as additional local evidence.
| Item | Configuration |
|---|---|
| GPU | NVIDIA GeForce RTX 5080 16 GB |
| CPU | AMD Ryzen 7 9800X3D |
| System RAM | 32 GB |
| OS | Windows 11 |
| Inference backend | llama.cpp llama-server |
| Main GGUF | WeMM-Embedding-2B-Q4_K_M.gguf |
| mmproj | mmproj-WeMM-Embedding-2B-Q8_0.gguf |
| Image / VisDoc / Tool / Memory context | 32768, KV F16 |
| Text context | 262144, KV Q8_0 |
| Pooling | last |
| Normalization | L2 |
| Parallel slots | -np 1 |
| Approx. evaluation time | 54 hours |
The Text benchmark intentionally uses 262K context. This matters for LongEmbed tasks and avoids turning long-document truncation into a fake "quantization loss".
The final result file records the two context profiles explicitly:
image/visdoc/tool/memory = ctx32768 kvf16text = ctx262144 kvq8_0The benchmark follows Tencent's released mmeb_v3_eval pipeline, which is based on the official VLM2Vec MMEB-v3 evaluator.
Pinned references used by this evaluation work:
9ed7e2d7914cd67a031c3e2a4fef3faaac3147212638a8413fda4b98668a29ea763b4898814bfea74a5560b2b64384204b6fea8a82ea986eba51f5aaThe GGUF backend replaces only model inference. Dataset parsing, instructions, query/candidate construction and RankingMetrics remain aligned with the released evaluator.
The tested modalities are:
Dataset-defined sample caps are preserved where present. Other tasks use their full query/candidate sets.
WeMM embedding extraction uses the final <embedding> token with last-token pooling.
For text input, the validated raw prompt surface is:
<|im_start|>user
YOUR_TEXT<|im_end|><embedding>
There is no newline between <|im_end|> and <embedding>.
The tokenizer must leave <embedding> as the unique final token. When using llama-server, disable automatic EOS insertion:
--override-kv tokenizer.ggml.add_eos_token=bool:false
The benchmark performs a /tokenize audit before evaluation and rejects a configuration where <embedding> is not the unique final token.
Validated 32K multimodal / general retrieval configuration:
llama-server \
-m WeMM-Embedding-2B-Q4_K_M.gguf \
--mmproj mmproj-WeMM-Embedding-2B-Q8_0.gguf \
--embedding \
--pooling last \
--embd-normalize 2 \
--override-kv tokenizer.ggml.add_eos_token=bool:false \
-ngl 99 \
-c 32768 \
-np 1 \
-b 512 \
-ub 512 \
--cache-type-k f16 \
--cache-type-v f16 \
--flash-attn on \
--image-min-tokens 64 \
--image-max-tokens 8192 \
--no-cache-prompt \
--no-webui \
--host 127.0.0.1 \
--port 8080
For the 262K Text benchmark profile:
llama-server \
-m WeMM-Embedding-2B-Q4_K_M.gguf \
--embedding \
--pooling last \
--embd-normalize 2 \
--override-kv tokenizer.ggml.add_eos_token=bool:false \
-ngl 99 \
-c 262144 \
-np 1 \
-b 512 \
-ub 512 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--flash-attn on \
--no-cache-prompt \
--no-webui \
--host 127.0.0.1 \
--port 8080
With the server running:
curl http://127.0.0.1:8080/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"input":"<|im_start|>user\nRepresent the meaning of this sentence.<|im_end|><embedding>","encoding_format":"float"}'
The returned vector should have 2048 dimensions and be L2 normalized.
For image / visual-document embeddings, use the paired:
mmproj-WeMM-Embedding-2B-Q8_0.gguf
The benchmark uses WeMM-compatible image/text ordering, an image budget of 64–8192 visual tokens, last-token pooling and L2 normalization.
Raw multimodal HTTP request schemas can change between llama.cpp builds. If reproducing the benchmark, keep the exact prompt ordering and verify that <embedding> remains the final token after tokenization.
For reproducibility, this repository should include the final benchmark outputs:
benchmarks/
├── 2B_scores_summary.csv
└── 2B_scores_detail.txt
2B_scores_summary.csv contains one row per evaluated dataset.
2B_scores_detail.txt contains all reported ranking metrics for all 156 datasets.
The 2B conversion set uses these filenames:
WeMM-Embedding-2B-Q4_K_M.gguf
WeMM-Embedding-2B-BF16.gguf
mmproj-WeMM-Embedding-2B-Q8_0.gguf
mmproj-WeMM-Embedding-2B-BF16.gguf
For local deployment, the benchmarked pair is:
WeMM-Embedding-2B-Q4_K_M.gguf
+
mmproj-WeMM-Embedding-2B-Q8_0.gguf
llama.cpp multimodal APIs evolve. Revalidate tokenizer ending, image handling and embedding dimensions when changing builds.The practical conclusion from this run:
WeMM-Embedding-2B-Q4_K_M.gguf+mmproj-WeMM-Embedding-2B-Q8_0.ggufis a high-fidelity local quantization of WeMM-Embedding-2B for the evaluated image, visual-document, text, tool and memory retrieval workloads.
Across the four directly comparable published benchmark groups, the GGUF pair retains approximately 96–99% of the published native WeMM-Embedding-2B score, while reducing the main model to Q4_K_M for local deployment.
tencent/WeMM-Embedding-2BTencent/WeMM-Embedding2608.24053The upstream WeMM-Embedding model is released under the Apache License 2.0. This GGUF repository follows the upstream licensing terms. Third-party evaluation datasets and components retain their original licenses.
<a id="chinese"></a>
tencent/WeMM-Embedding-2B 的 GGUF 转换版本,面向使用 llama.cpp 的本地多模态 Embedding 推理。
推荐且已经完整验证的组合
WeMM-Embedding-2B-Q4_K_M.ggufmmproj-WeMM-Embedding-2B-Q8_0.gguf64, 128, 256, 512, 1024, 2048这个仓库不只是一次 GGUF 转换。
Q4_K_M + Q8_0 mmproj组合已经完成 156 个 MMEB-v3 任务的端到端实测,本地累计评测时间约 54 小时。
评测日期:2026-09-04
| Benchmark 组 | 指标 | 任务数 | 腾讯官方 WeMM-2B | 本仓库 GGUF | Δ | 分数保留率 |
|---|---|---|---|---|---|---|
| MMEB-v2 Image | Hit@1 | 36 | 79.60 | 79.18 | -0.42 | 99.47% |
| MMEB-v2 VisDoc | NDCG@5 | 27 | 80.70 | 78.07 | -2.63 | 96.75% |
| MMEB-v3 Text | NDCG@5 | 53 | 45.30 | 43.65 | -1.65 | 96.37% |
| MMEB-v3 MCMR | Hit@1 | 1 | 42.50 | 41.01 | -1.49 | 96.49% |
上表中的腾讯分数来自 WeMM 技术报告公布的原生模型聚合成绩。它们不是在同一张 RTX 5080 上重新跑出来的 BF16 对照,因此这些差值应该理解为“本地 GGUF 结果与官方原生基线的一致程度”,而不是严格控制变量下的纯量化损失。
最重要的结论是:
Q4_K_M主模型配合Q8_0视觉投影器,在四个可直接比较的公开 benchmark 组上保留了大约 96%–99% 的官方原生分数。
本次评测覆盖 156 / 190 个 MMEB-v3 任务:
| 模态 / 分组 | 任务数 | 指标 | 本地结果 |
|---|---|---|---|
| Image | 36 | Hit@1 | 79.18 |
| VisDoc | 27 | NDCG@5 | 78.07 |
| Text | 53 | NDCG@5 | 43.65 |
| Tool | 35 | Hit@1 | 49.00 |
| Memory | 4 | Hit@1 | 35.67 |
| MCMR | 1 | Hit@1 | 41.01 |
| Tool + Memory(不含 GUI 的 Agent 子集) | 39 | Hit@1 | 47.63 |
本轮未测试:
原始 WeMM 模型支持视频,但本次 GGUF benchmark 有意排除了 Video。WeMM-Embedding 不支持 Audio。
由于排除了 8 个 GUI 任务,本地 39 项 Tool + Memory 结果不能直接与腾讯官方公布的 47 项 Agent 聚合分数比较。
下面将官方原生模型结果与本仓库的本地 GGUF 实测结果分开列出。这样做很重要,因为拿本地 Q4_K_M 直接和原生 BF16/FP16 榜单提交比较,作为部署参考很有价值,但并不是严格意义上的纯量化消融实验。
腾讯在 78 项 MMEB-v2 上公布的结果如下:
| 模型 | 尺寸 | AVG | Image Hit@1 | Video Hit@1 | VisDoc NDCG@5 |
|---|---|---|---|---|---|
| Qwen3-VL-Embedding | 2B | 73.2 | 75.0 | 61.9 | 79.2 |
| WeMM-Embedding | 2B | 77.9 | 79.6 | 70.8 | 80.7 |
| WeMM-Embedding | 4B | 79.2 | 80.8 | 72.1 | 82.0 |
| Qwen3-VL-Embedding | 8B | 77.8 | 80.1 | 67.1 | 82.4 |
| WeMM-Embedding | 9B | 80.6 | 81.9 | 74.3 | 83.3 |
同为约 2B 规模时,官方 WeMM-Embedding-2B 相比 Qwen3-VL-Embedding-2B:
| MMEB-v2 指标 | WeMM-2B 优势 |
|---|---|
| AVG | +4.7 |
| Image | +4.6 |
| Video | +8.9 |
| VisDoc | +1.5 |
大模型尺寸下,WeMM-Embedding-9B 相比 Qwen3-VL-Embedding-8B:
| MMEB-v2 指标 | WeMM-9B 优势 |
|---|---|
| AVG | +2.8 |
| Image | +1.8 |
| Video | +7.2 |
| VisDoc | +0.9 |
MMEB-v2 中 WeMM 最稳定、最明显的优势出现在视频检索;VisDoc 上的差距则小得多。
腾讯在完整 190 项 MMEB-v3 上公布的结果:
| 模型 | 尺寸 | V3-All | Text NDCG@5 | Agent Hit@1 | MCMR Hit@1 | Audio |
|---|---|---|---|---|---|---|
| Qwen3-VL-Embedding | 2B | 50.9 | 39.2 | 39.3 | 42.0 | 0.0 |
| WeMM-Embedding | 2B | 56.0 | 45.3 | 45.1 | 42.5 | 0.0 |
| WeMM-Embedding | 4B | 58.2 | 47.9 | 49.0 | 41.9 | 0.0 |
| Qwen3-VL-Embedding | 8B | 53.5 | 42.5 | 38.4 | 38.0 | 0.0 |
| WeMM-Embedding | 9B | 59.5 | 48.8 | 51.0 | 49.3 | 0.0 |
同为 2B 规模时,官方 WeMM-Embedding-2B 相比 Qwen3-VL-Embedding-2B:
| MMEB-v3 指标 | WeMM-2B 优势 |
|---|---|
| V3-All | +5.1 |
| Text | +6.1 |
| Agent | +5.8 |
| MCMR | +0.5 |
大模型尺寸下,WeMM-Embedding-9B 相比 Qwen3-VL-Embedding-8B:
| MMEB-v3 指标 | WeMM-9B 优势 |
|---|---|
| V3-All | +6.0 |
| Text | +6.3 |
| Agent | +12.6 |
| MCMR | +11.3 |
这是判断本次 GGUF 转换质量最重要的一张表。
| Benchmark 组 | 官方 WeMM-2B | 本仓库 2B Q4_K_M GGUF | Δ | 分数保留率 |
|---|---|---|---|---|
| Image | 79.60 | 79.18 | -0.42 | 99.47% |
| VisDoc | 80.70 | 78.07 | -2.63 | 96.75% |
| Text | 45.30 | 43.65 | -1.65 | 96.37% |
| MCMR | 42.50 | 41.01 | -1.49 | 96.49% |
在四个可以直接和官方公开成绩比较的 benchmark 组上,Q4_K_M + Q8_0 mmproj 组合保留了约 96%–99% 的官方原生能力。
这张表更接近实际部署问题:已经量化到 Q4 的 WeMM-2B,与 Qwen3-VL-Embedding-2B 官方原生模型相比如何?
| Benchmark 组 | Qwen3-VL-Embedding-2B 原生 | WeMM-2B Q4_K_M GGUF | GGUF 差值 |
|---|---|---|---|
| Image | 75.00 | 79.18 | +4.18 |
| VisDoc | 79.20 | 78.07 | -1.13 |
| Text | 39.20 | 43.65 | +4.45 |
| MCMR | 42.00 | 41.01 | -0.99 |
即使主模型已经量化为 Q4_K_M,本仓库的 WeMM-2B 在 Image 和 Text 上仍明显高于官方 Qwen3-VL-Embedding-2B 原生基线;VisDoc 和 MCMR 则落后约 1 分。
这不是严格同口径比较,但很适合衡量 2B GGUF 的部署效率。
| Benchmark 组 | Qwen3-VL-Embedding-8B 原生 | WeMM-2B Q4_K_M GGUF | GGUF 差值 |
|---|---|---|---|
| Image | 80.10 | 79.18 | -0.92 |
| VisDoc | 82.40 | 78.07 | -4.33 |
| Text | 42.50 | 43.65 | +1.15 |
| MCMR | 38.00 | 41.01 | +3.01 |
值得注意的是:
实测的 2B Q4_K_M WeMM GGUF 在 Text 和 MCMR 上超过了官方 Qwen3-VL-Embedding-8B,Image 也只差 0.92 分;Qwen3-VL-Embedding-8B 在 VisDoc 上则明显更强。
本地评测有意排除了 Video 与 GUI,因此无法严谨重建部分官方聚合指标:
Tool + Memory = 47.63 Hit@1 / 39 tasks 可以作为 no-GUI Agent 子集参考,但不能写成官方 47-task Agent 成绩。因此,推荐使用 Image / VisDoc / Text / MCMR 做直接官方成绩对比,把 Tool / Memory 视为额外的本地验证证据。
| 项目 | 配置 |
|---|---|
| GPU | NVIDIA GeForce RTX 5080 16 GB |
| CPU | AMD Ryzen 7 9800X3D |
| 系统内存 | 32 GB |
| 操作系统 | Windows 11 |
| 推理后端 | llama.cpp llama-server |
| 主 GGUF | WeMM-Embedding-2B-Q4_K_M.gguf |
| mmproj | mmproj-WeMM-Embedding-2B-Q8_0.gguf |
| Image / VisDoc / Tool / Memory context | 32768,KV F16 |
| Text context | 262144,KV Q8_0 |
| Pooling | last |
| Normalization | L2 |
| 并行槽位 | -np 1 |
| 累计评测时间 | 约 54 小时 |
Text benchmark 有意使用 262K context。这是 LongEmbed 等超长文本任务能够公平评测的关键,避免把长文截断造成的损失误判为量化损失。
最终结果文件明确记录了两套 context profile:
image/visdoc/tool/memory = ctx32768 kvf16text = ctx262144 kvq8_0评测沿用腾讯发布的 mmeb_v3_eval pipeline;它基于官方 VLM2Vec MMEB-v3 evaluator。
本轮评测使用的固定版本:
9ed7e2d7914cd67a031c3e2a4fef3faaac3147212638a8413fda4b98668a29ea763b4898814bfea74a5560b2b64384204b6fea8a82ea986eba51f5aaGGUF backend 只替换模型推理部分。数据集解析、instructions、query/candidate 构造以及 RankingMetrics 均继续与腾讯发布的 evaluator 对齐。
已测试:
数据集本身定义了 sample cap 的任务继续沿用该 cap;其他任务使用完整 query / candidate 集合。
WeMM 的 embedding 从最终 <embedding> token 提取,并使用 last-token pooling。
纯文本的验证 prompt 形式:
<|im_start|>user
YOUR_TEXT<|im_end|><embedding>
<|im_end|> 与 <embedding> 之间不能插入换行。
tokenizer 必须确保 <embedding> 是唯一的最后一个 token。使用 llama-server 时,需要关闭自动 EOS:
--override-kv tokenizer.ggml.add_eos_token=bool:false
benchmark 在正式评测前通过 /tokenize 检查这一点;如果 <embedding> 不是唯一的最终 token,评测会直接拒绝继续。
已经验证的 32K 多模态 / 通用检索配置:
llama-server \
-m WeMM-Embedding-2B-Q4_K_M.gguf \
--mmproj mmproj-WeMM-Embedding-2B-Q8_0.gguf \
--embedding \
--pooling last \
--embd-normalize 2 \
--override-kv tokenizer.ggml.add_eos_token=bool:false \
-ngl 99 \
-c 32768 \
-np 1 \
-b 512 \
-ub 512 \
--cache-type-k f16 \
--cache-type-v f16 \
--flash-attn on \
--image-min-tokens 64 \
--image-max-tokens 8192 \
--no-cache-prompt \
--no-webui \
--host 127.0.0.1 \
--port 8080
用于 262K Text benchmark 的配置:
llama-server \
-m WeMM-Embedding-2B-Q4_K_M.gguf \
--embedding \
--pooling last \
--embd-normalize 2 \
--override-kv tokenizer.ggml.add_eos_token=bool:false \
-ngl 99 \
-c 262144 \
-np 1 \
-b 512 \
-ub 512 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--flash-attn on \
--no-cache-prompt \
--no-webui \
--host 127.0.0.1 \
--port 8080
服务器启动后:
curl http://127.0.0.1:8080/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"input":"<|im_start|>user\nRepresent the meaning of this sentence.<|im_end|><embedding>","encoding_format":"float"}'
返回向量应为 2048 维,并已经完成 L2 normalization。
Image / VisDoc embedding 需要配套:
mmproj-WeMM-Embedding-2B-Q8_0.gguf
本次 benchmark 使用与 WeMM 对齐的图像 / 文本顺序、64–8192 visual tokens 图像预算、last-token pooling 与 L2 normalization。
llama.cpp 的多模态 HTTP schema 可能随版本变化。复现 benchmark 时,应优先保证 prompt 顺序完全一致,并重新确认 <embedding> 在 tokenization 后仍是最后一个 token。
建议仓库同时保留最终结果文件:
benchmarks/
├── 2B_scores_summary.csv
└── 2B_scores_detail.txt
2B_scores_summary.csv:每个评测数据集一行。
2B_scores_detail.txt:包含全部 156 个数据集的所有 ranking metrics。
2B 转换文件使用:
WeMM-Embedding-2B-Q4_K_M.gguf
WeMM-Embedding-2B-BF16.gguf
mmproj-WeMM-Embedding-2B-Q8_0.gguf
mmproj-WeMM-Embedding-2B-BF16.gguf
本轮完整 benchmark 使用的是:
WeMM-Embedding-2B-Q4_K_M.gguf
+
mmproj-WeMM-Embedding-2B-Q8_0.gguf
llama.cpp 多模态 API 会持续变化。更换 build 后,应重新验证 tokenizer 尾部、图像处理和 embedding 维度。本轮测试支持以下实际结论:
WeMM-Embedding-2B-Q4_K_M.gguf+mmproj-WeMM-Embedding-2B-Q8_0.gguf是 WeMM-Embedding-2B 在已测试 Image、VisDoc、Text、Tool、Memory retrieval 工作负载上的高保真本地量化版本。
在四个能够与腾讯官方公开成绩直接比较的 benchmark 组上,这套 GGUF 组合保留了大约 96%–99% 的官方原生分数,同时把主模型压缩到了 Q4_K_M,更适合本地部署。
tencent/WeMM-Embedding-2BTencent/WeMM-Embedding2608.24053上游 WeMM-Embedding 模型采用 Apache License 2.0。本 GGUF 仓库遵循上游许可条款。第三方评测数据集和组件继续遵循各自原始 License。