Downloads · 30 days
66
3% of all-time downloads
fgaim/tielectra-bi-encoder
tielectra-bi-encoder is a sentence similarity model from fgaim. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
This model is a bi-encoder model for the Tigrinya language based on TiELECTRA-small. The model maps sentences & paragraphs to a 256 dimensional dense vector space and can be used for tasks like text embedding, cluster…
Downloads · 30 days
66
3% of all-time downloads
All-time downloads
1.9K
Public
Parameters
13.5M
54 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors54 MB · 98%
From the Hugging Face model README
This model is a bi-encoder model for the Tigrinya language based on TiELECTRA-small. The model maps sentences & paragraphs to a 256 dimensional dense vector space and can be used for tasks like text embedding, clustering, or semantic search.
This is part of a work that introduces monolingual bi-encoder language models for Tigrinya. For a larger and more powerful model look at TiRoBERTa-bi-encoder. The models are based on the sentence-transformers architecture and are trained on Tigrinya question-answering and information retrieval datasets. The models are designed to support semantic search tasks, such as information retrieval, text representation, and question answering.
Using this model becomes easy when you have sentence-transformers installed:
pip install -U sentence-transformers
Then use the model as follows:
from sentence_transformers import SentenceTransformer
sentences = ["ሓደ ሰብኣይ ፈረስ ይጋልብ ኣሎ።", "ሓንቲ ጓል ክራር ትጻወት ኣላ።"]
model = SentenceTransformer('fgaim/tielectra-bi-encoder')
embeddings = model.encode(sentences)
print(embeddings)
Use the transformers library as follows: Pass the input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings.
import torch
from transformers import AutoModel, AutoTokenizer
# Mean Pooling - Take attention mask into account for correct averaging
def mean_pooling(model_output, attention_mask):
token_embeddings = model_output[0] # First element of model_output contains all token embeddings
input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(input_mask_expanded.sum(1), min=1e-9)
# Sentences we want sentence embeddings for
sentences = ["ሓደ ሰብኣይ ፈረስ ይጋልብ ኣሎ።", "ሓንቲ ጓል ክራር ትጻወት ኣላ።"]
# Load model from HuggingFace Hub
tokenizer = AutoTokenizer.from_pretrained("fgaim/tielectra-bi-encoder")
model = AutoModel.from_pretrained("fgaim/tielectra-bi-encoder")
# Tokenize sentences
encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors="pt")
# Compute token embeddings
with torch.no_grad():
model_output = model(**encoded_input)
# Perform pooling. In this case, mean pooling.
sentence_embeddings = mean_pooling(model_output, encoded_input["attention_mask"])
print("Sentence embeddings:", sentence_embeddings)
The model properties:
| Model Size | Layers | Attn. Heads | Hidden Size | FFN | Parameters | Max. Seq |
|---|---|---|---|---|---|---|
| SMALL | 12 | 4 | 256 | 1024 | 14M | 512 |
512256SentenceTransformer(
(0): Transformer({'max_seq_length': 512, 'do_lower_case': False}) with Transformer model: ElectraModel
(1): Pooling({'word_embedding_dimension': 256, 'pooling_mode_cls_token': False, 'pooling_mode_mean_tokens': True, 'pooling_mode_max_tokens': False, 'pooling_mode_mean_sqrt_len_tokens': False, 'pooling_mode_weightedmean_tokens': False, 'pooling_mode_lasttoken': False, 'include_prompt': True})
)
If you use this model in your product or research, you can cite it as follows:
@misc{gaim-2024-semantic-search,
title = {{Semantic Search Models for Tigrinya}},
author = {Fitsum Gaim},
month = {January},
year = {2024},
publisher = {Hugging Face Hub},
doi = {10.57967/hf/6068},
url = {https://huggingface.co/fgaim/tiroberta-bi-encoder}
}