Downloads · 30 days
44
100% of all-time downloads
fcmeyer/F2LLM-v2-0.6B-bf16-mlx
F2LLM-v2-0.6B-bf16-mlx is a feature extraction model from fcmeyer. Use it when you need embeddings to search or compare text. It is set up for mlx-embeddings. The card lists the license as apache-2.0.
MLX-native port of codefuse-ai/F2LLM-v2-0.6B (a 0.6B Qwen3-based multilingual embedding model, 1024-dim, last-token pooling, L2-normalized, MRL-trained) for Apple Silicon via mlx-embeddings.
Downloads · 30 days
44
100% of all-time downloads
All-time downloads
44
Public
Parameters
596M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 99%
From the Hugging Face model README
MLX-native port of codefuse-ai/F2LLM-v2-0.6B (a 0.6B Qwen3-based multilingual embedding model, 1024-dim, last-token pooling, L2-normalized, MRL-trained) for Apple Silicon via mlx-embeddings.
python -m mlx_embeddings.convert --hf-path codefuse-ai/F2LLM-v2-0.6B --mlx-path ./F2LLM-v2-0.6B-bf16-mlx --dtype bfloat16 (mlx-embeddings 0.1.0).model. added per mlx-embeddings convention; no
quantization). 1.1 GB.model.safetensors (+ index), config.json, tokenizer files,
modules.json, config_sentence_transformers.json, 1_Pooling/config.json
(last-token pooling, include_prompt=true).from mlx_embeddings.utils import load
model, tokenizer = load("fcmeyer/F2LLM-v2-0.6B-bf16-mlx")
query_prompt = "Instruct: Given a question, retrieve passages that can help answer the question.\nQuery: "
texts = [
query_prompt + "What is F2LLM used for?",
"We present F2LLM, a family of fully open embedding LLMs.",
"F2LLM 是 CodeFuse 开源的系列嵌入模型。",
]
inputs = tokenizer.batch_encode_plus(
texts, return_tensors="mlx", padding=True, truncation=True, max_length=4096,
)
outputs = model(inputs["input_ids"], attention_mask=inputs["attention_mask"])
embeddings = outputs.text_embeds # pooled + normalized, (3, 1024)
similarity = embeddings[0:1] @ embeddings[1:].T
Notes:
e = e[..., :128]; e = e / max(norm(e), 1e-9).truncation=True keeps the head + appends EOS (verified: first 511 tokens + EOS
at max_length=512).mlx-embeddings Qwen3 fast-SDPA masking behavior (reproduces on the
float32 conversion too, smaller at ~0.14 hidden-state diff vs ~2.4 in bf16), not
a weight error. For max fidelity, pad to similar lengths or encode
length-mismatched inputs separately.10-text suite (query + EN/ZH/RU docs + unrelated + short/long/code/mixed-language):
@misc{f2llm-v2,
title={F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World},
author={Ziyin Zhang and Zihan Liao and Hang Yu and Peng Di and Rui Wang},
year={2026},
eprint={2603.19223},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.19223},
}
Source model card and license (Apache-2.0): https://huggingface.co/codefuse-ai/F2LLM-v2-0.6B