Downloads · 30 days
4.1K
40% of all-time downloads
software-mansion/react-native-executorch-paraphrase-multilingual-MiniLM-L12-v2
react-native-executorch-paraphrase-multilingual-MiniLM-L12-v2 is a sentence similarity model from software-mansion. Use it when you need a score for how close two texts are. It is set up for executorch. The card lists the license as apache-2.0.
This repository hosts the paraphrase-multilingual-MiniLM-L12-v2 models exported for the React Native ExecuTorch library as ExecuTorch .pte programs, ready to run on device.
Downloads · 30 days
4.1K
40% of all-time downloads
All-time downloads
10.3K
Public
Repo size
5.6 GB
Likes
0
Public
Click a slice to open those files.
.pte1.5 GB · 99%
From the Hugging Face model README
This repository hosts the paraphrase-multilingual-MiniLM-L12-v2 models exported for the
React Native ExecuTorch
library as ExecuTorch .pte programs, ready to run on device.
Upstream model: paraphrase-multilingual-MiniLM-L12-v2
| Path | Backend | Precision |
|---|---|---|
coreml/paraphrase_multilingual_minilm_l12_v2_coreml_fp16.pte | coreml | fp16 |
vulkan/paraphrase_multilingual_minilm_l12_v2_vulkan_fp16.pte | vulkan | fp16 |
xnnpack/paraphrase_multilingual_minilm_l12_v2_xnnpack_fp32.pte | xnnpack | fp32 |
xnnpack/paraphrase_multilingual_minilm_l12_v2_xnnpack_8da4w.pte | xnnpack | 8da4w |
config.json 59 B
coreml/config.json 971 B
coreml/paraphrase_multilingual_minilm_l12_v2_coreml_fp16.pte 225 MB
tokenizer.json 16.3 MB
tokenizer_config.json 526 B
vulkan/config.json 971 B
vulkan/paraphrase_multilingual_minilm_l12_v2_vulkan_fp16.pte 224 MB
xnnpack/config.json 1.6 kB
xnnpack/paraphrase_multilingual_minilm_l12_v2_xnnpack_8da4w.pte 379 MB
xnnpack/paraphrase_multilingual_minilm_l12_v2_xnnpack_fp32.pte 448 MB
These files are published for the ExecuTorch v1.4.1 runtime. ExecuTorch gives no forward compatibility guarantee, so an older runtime may fail to load them.
To use them in React Native ExecuTorch, pass the model constant shipped in the library's model registry to the corresponding task pipeline. See the documentation.
To load these files in your own ExecuTorch runtime, read the compatibility note first.
xlm-roberta-base) + mean pooling + L2 norm. No additional dense projection head — the model output dim equals the encoder hidden size.<s> / </s> wrapping; the exporter concatenates these XLM-R-style start/end tokens at id 0 / 2 inside the program).The exporter wraps the HuggingFace transformer with the standard sentence-transformers contract: token IDs go in, the program prepends <s> and appends </s>, mean pooling is applied to the last hidden state weighted by the attention mask, and the output is L2-normalized to a 384-d vector.
Unsupported combinations (rejected by the exporter, documented for reference):
model.to(torch.float16) causes softmax / LayerNorm overflow and the runtime output is NaN. XNNPACK's size wins come from quantization, not fp16.coremltools has no MIL mapping for the torch.int8 tensors torchao emits (KeyError: torch.int8). The CoreML-native way to shrink further is ct.optimize.coreml palette/linear quantization, not torchao source transforms.