Downloads · 30 days
81
24% of all-time downloads
NorskHelsenett/SPLADE-v1
SPLADE-v1 is a feature extraction model from NorskHelsenett. Use it when you need embeddings to search or compare text. It is set up for sentence-transformers. The card lists the license as apache-2.0.
A SPLADE learned-sparse retriever for Norwegian-language health, welfare and assistive-technology content. It turns text into a sparse vector over the vocabulary: most dimensions are zero, and the non-zero ones are we…
Downloads · 30 days
81
24% of all-time downloads
All-time downloads
344
Public
Parameters
167M
670 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors670 MB · 99%
From the Hugging Face model README
A SPLADE learned-sparse retriever for Norwegian-language health, welfare and assistive-technology content. It turns text into a sparse vector over the vocabulary: most dimensions are zero, and the non-zero ones are weighted terms. Alongside the words that actually appear, the model expands each text with closely related terms it has learned — so everyday, lay phrasing is matched to the official terminology used in source documents, and vice-versa.
Because the output is a weighted bag of terms, it drops straight into an inverted index
(e.g. Elasticsearch / OpenSearch rank_features) and works well as the sparse leg of a
hybrid search stack.
| Model type | SparseEncoder (SPLADE), sentence-transformers |
| Base architecture | multilingual BERT (BertForMaskedLM) |
| Language | Norwegian bokmål (multilingual base) |
| Max sequence length | 512 tokens |
| Vocabulary size | ~106k (multilingual) |
| Output | sparse vector in vocabulary space (weighted terms) |
from sentence_transformers import SparseEncoder
model = SparseEncoder("NorskHelsenett/SPLADE-v1")
query = "Hvem har rett til hjelpemidler?"
documents = [
"Personer med varig nedsatt funksjonsevne kan søke om hjelpemidler fra NAV.",
"Åpningstidene for legekontoret er mandag til fredag 08–15.",
]
q_emb = model.encode_query([query])
d_emb = model.encode_document(documents)
scores = model.similarity(q_emb, d_emb)
print(scores)
# Inspect which terms the query expanded to (term -> weight)
print(model.decode(q_emb[0], top_k=15))
model.decode(...) gives you (term, weight) pairs. Index the document term weights
as feature weights (e.g. Elasticsearch rank_features), and at query time score documents
by the dot product between the query terms and the stored document terms. This is the
inference-free / index-time-expansion setup SPLADE is designed for.
Fine-tuned by knowledge distillation from a stronger teacher ranker on Norwegian health and welfare query–document data, with sparsity regularization so the output vectors stay compact and index-efficient. Training and evaluation used curated Norwegian questions paired with human-labelled source documents.
Evaluated on held-out Norwegian health/welfare questions with human-labelled ground-truth source documents. The model provides strong document-level retrieval while keeping the sparse vectors compact, and it recovers relevant documents that pure keyword matching misses by expanding lay phrasing to the formal terms used in the sources.
SPLADE runs the input through a masked-language-model head and, for every token position,
looks at the distribution over the whole vocabulary. Those distributions are pooled and
passed through log(1 + ReLU(·)), producing a single sparse vector where each non-zero
entry is a vocabulary term with a learned importance weight — combining exact-match
retrieval with learned term expansion, all computable at index time.