Downloads · 30 days
133
6% of all-time downloads
fgaim/tiroberta-bi-encoder
tiroberta-bi-encoder is a sentence similarity model from fgaim. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
This model is a bi-encoder model for the Tigrinya language based on TiRoBERTa-base. The model maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like text embedding, clusteri…
Downloads · 30 days
133
6% of all-time downloads
All-time downloads
2.1K
Public
Parameters
125M
499 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors499 MB · 99%
From the Hugging Face model README
This model is a bi-encoder model for the Tigrinya language based on TiRoBERTa-base. The model maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like text embedding, clustering, or semantic search.
This is part of a work that introduces monolingual bi-encoder language models for Tigrinya. For a smaller and lightweight model look at TiELECTRA-bi-encoder. The models are based on the sentence-transformers architecture and are trained on Tigrinya question-answering and information retrieval datasets. The models are designed to support semantic search tasks, such as information retrieval, text representation, and question answering.
Using this model becomes easy when you have sentence-transformers installed:
pip install -U sentence-transformers
Then use the model as follows:
from sentence_transformers import SentenceTransformer
sentences = ["ሓደ ሰብኣይ ፈረስ ይጋልብ ኣሎ።", "ሓንቲ ጓል ክራር ትጻወት ኣላ።"]
model = SentenceTransformer('fgaim/tiroberta-bi-encoder')
embeddings = model.encode(sentences)
print(embeddings)
Use the transformers library as follows: Pass the input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings.
import torch
from transformers import AutoModel, AutoTokenizer
# Mean Pooling - Take attention mask into account for correct averaging
def mean_pooling(model_output, attention_mask):
token_embeddings = model_output[0] # First element of model_output contains all token embeddings
input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(input_mask_expanded.sum(1), min=1e-9)
# Sentences we want sentence embeddings for
sentences = ["ሓደ ሰብኣይ ፈረስ ይጋልብ ኣሎ።", "ሓንቲ ጓል ክራር ትጻወት ኣላ።"]
# Load model from HuggingFace Hub
tokenizer = AutoTokenizer.from_pretrained("fgaim/tiroberta-bi-encoder")
model = AutoModel.from_pretrained("fgaim/tiroberta-bi-encoder")
# Tokenize sentences
encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors="pt")
# Compute token embeddings
with torch.no_grad():
model_output = model(**encoded_input)
# Perform pooling. In this case, mean pooling.
sentence_embeddings = mean_pooling(model_output, encoded_input["attention_mask"])
print("Sentence embeddings:", sentence_embeddings)
The model properties:
| Model Size | Layers | Attn. Heads | Hidden Size | FFN | Parameters | Max. Seq |
|---|---|---|---|---|---|---|
| BASE | 12 | 12 | 768 | 3072 | 125M | 512 |
512768SentenceTransformer(
Transformer(
{
'max_seq_length': 512,
'do_lower_case': False
}
) # with Transformer model: RobertaModel
Pooling(
{
'word_embedding_dimension': 768,
'pooling_mode_cls_token': False,
'pooling_mode_mean_tokens': True,
'pooling_mode_max_tokens': False,
'pooling_mode_mean_sqrt_len_tokens': False,
'pooling_mode_weightedmean_tokens': False,
'pooling_mode_lasttoken': False,
'include_prompt': True,
}
)
)
If you use this model in your product or research, you can cite it as follows:
@misc{gaim-2024-semantic-search,
title = {{Semantic Search Models for Tigrinya}},
author = {Fitsum Gaim},
month = {January},
year = {2024},
publisher = {Hugging Face Hub},
doi = {10.57967/hf/6068},
url = {https://huggingface.co/fgaim/tiroberta-bi-encoder}
}