Downloads · 30 days
15
31% of all-time downloads
hishammoizuddin/vericite-embedder
vericite-embedder is a sentence similarity model from hishammoizuddin. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
This is the sentence-embedding model used by VeriCite, a free plagiarism/paraphrase checker that cites its sources.
Downloads · 30 days
15
31% of all-time downloads
All-time downloads
49
Public
Parameters
22.7M
90.9 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors90.9 MB · 99%
How the weights are stored.
F3222.7M · 100%
From the Hugging Face model README
This is the sentence-embedding model used by VeriCite, a free plagiarism/paraphrase checker that cites its sources.
These are the unmodified weights of
sentence-transformers/all-MiniLM-L6-v2,
mirrored here so VeriCite can pin a stable, versioned copy rather than
depending on the upstream repo's main branch at request time. All credit
for the model itself belongs to the original authors (Nils Reimers and the
Sentence-Transformers / UKP Lab team) — see their
Sentence-BERT paper and the
original model card
for training data, evaluation results, and the full technical writeup. No
fine-tuning, distillation, or architecture change has been applied here.
VeriCite's matching pipeline needs a small, CPU-fast embedding model to
compute semantic similarity between a document passage and candidate
source text, as one signal alongside exact-match fuzzy alignment and a
lexical-overlap guard (semantic similarity alone is not sufficient to
flag a passage — see the VeriCite Space
"How it works" tab for the full method). all-MiniLM-L6-v2 fits that
constraint well: 384-dim output, ~90MB, and fast enough to run comfortably
on a free-tier 2-vCPU CPU box with no GPU.
Identical to the upstream model — either via sentence-transformers:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("hishammoizuddin/vericite-embedder")
embeddings = model.encode(["Some sentence to embed.", "Another one."], normalize_embeddings=True)
or via transformers directly (mean-pool the token embeddings yourself,
since this checkpoint is a base encoder, not a SentenceTransformer-native
format at the transformers layer):
from transformers import AutoModel, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("hishammoizuddin/vericite-embedder")
model = AutoModel.from_pretrained("hishammoizuddin/vericite-embedder")
Apache 2.0, unchanged from the upstream model.