Downloads · 30 days
16
41% of all-time downloads
mlboydaisuke/all-mpnet-base-v2-ExecuTorch
all-mpnet-base-v2-ExecuTorch is a feature extraction model from mlboydaisuke. Use it when you need embeddings to search or compare text. The card lists the license as apache-2.0.
The most downloaded sentence-transformer there is. Text in, one 768-dimensional vector out, for search and retrieval that never leaves the device.
Downloads · 30 days
16
41% of all-time downloads
All-time downloads
39
Public
Repo size
874 MB
Likes
0
Public
Click a slice to open those files.
.pte874 MB · 100%
From the Hugging Face model README
The most downloaded sentence-transformer there is. Text in, one 768-dimensional vector out, for search and retrieval that never leaves the device.
input_ids and attention_mask, both [1, 256] int64[1, 768], mean-pooled and L2-normalised inside the graphsentence-transformers stores it per model, and the shelf's seven embedding models do
not agree. This one pools mean and
normalises, read from
1_Pooling/config.json and modules.json rather than inferred from the family name.
Getting it wrong does not throw; it returns vectors that look fine and rank wrong.
| build | file | size (MB) | Mac ms* | backend takes | worst cosine vs eager | retrieval budget |
|---|---|---|---|---|---|---|
| fp32 | embed_all_mpnet_xnnpack_fp32.pte | 435.8 | 35.2 | 75.5% | 1.000000 | 0% |
| fp16 | embed_all_mpnet_xnnpack_fp16.pte | 218.1 | 56.2 | 65.7% | 1.000000 | 11% |
| Core ML (fp16, iOS) | embed_all_mpnet_coreml_all.pte | 220.2 | 6.2 | 100.0% | 0.999993 | 32% |
*Mac arm64, one 256-token sequence, fastest of five medians of ten — a reference point for relative cost, not a device number. The host shares its cores with other work, and a single median does not survive that: the same eager model here measured 19.6 ms and 182.8 ms twenty minutes apart. Contention only ever adds time, so the fastest repetition is the one that means something. Torch eager fp32, measured the same way, is 37.0 ms.
Cosine is measured against the model run in eager through its own pooling, over eight sentences. The last column is the one that decides: rank those eight against each other, and ask whether this build's score error is smaller than the gap between the document a query retrieves and the runner-up. Every shipped build keeps all eight top-1 results.
embed_all_mpnet_xnnpack_int8.pte is 181.3 MB — smaller than fp16's 218.1 MB, because the token embedding
table is only 94 MB of the 435.8 MB model (22%), leaving most of the
weight in linears for int8 to shrink.
It is withheld on the number that decides. Ranking the eight test sentences against each other, this build moves a pair score by at most 0.0081 while the closest fp32 decision — the gap between the document a query retrieves and the runner-up — is 0.0026. That is 316% of the room available, against a bar of 50%.
Correlation reads 0.998871 for this build, which no correlation gate would stop.
torch.export -> to_edge_transform_and_lower(partitioner) -> .pte (conversion scripts: executorch-models)