Downloads · 30 days
12
39% of all-time downloads
dschulmeist/TiME-da-s
TiME-da-s is a feature extraction model from dschulmeist. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
Monolingual BERT-style encoder that outputs embeddings for Danish. Distilled from FacebookAI/xlm-roberta-large.
Downloads · 30 days
12
39% of all-time downloads
All-time downloads
31
Public
Repo size
878 MB
Likes
0
Public
Click a slice to open those files.
.bin428 MB · 95%
From the Hugging Face model README
Monolingual BERT-style encoder that outputs embeddings for Danish. Distilled from FacebookAI/xlm-roberta-large.
from transformers import AutoTokenizer, AutoModel
import torch
repo = "dschulmeist/TiME-da-s"
tok = AutoTokenizer.from_pretrained(repo)
mdl = AutoModel.from_pretrained(repo)
def mean_pool(last_hidden_state, attention_mask):
mask = attention_mask.unsqueeze(-1).type_as(last_hidden_state)
return (last_hidden_state * mask).sum(1) / mask.sum(1).clamp(min=1e-9)
inputs = tok(["example sentence"], padding=True, truncation=True, return_tensors="pt")
outputs = mdl(**inputs)
emb = mean_pool(outputs.last_hidden_state, inputs['attention_mask'])
print(emb.shape)