Downloads · 30 days
8
0% of all-time downloads
tutur90/SONAR-Text-to-Text
SONAR-Text-to-Text is a machine learning model from tutur90. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Original weights of SONAR converted for hugging face model.
Downloads · 30 days
8
0% of all-time downloads
All-time downloads
17.1K
Public
Parameters
1.6B
6.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6.5 GB · 99%
From the Hugging Face model README
Original weights of SONAR converted for hugging face model.
This is a part of Open SONAR project, an open training pipeline for SONAR. This pipeline could be use to finetune or train from scratch a sonar model.
Examples are avalaible here.
SONARForSpeech2Text.from_pretrained("tutur90/SONAR-Text-to-Text")
NllbTokenizer.from_pretrained("tutur90/SONAR-Text-to-Text")
The code of SONARForSpeech2Text avalaible in Open SONAR - Model and NllbTokenizer Open SONAR - Tokenizer
inputs = tokenizer(sentence, langs=src_lang, return_tensors="pt")
generated = sonar.generate(
**inputs,
target_lang_ids=[tokenizer.convert_tokens_to_ids(tgt_lang)],
max_length=128,
num_beams=1,
do_sample=False,
)
decoded = tokenizer.batch_decode(generated, skip_special_tokens=True)[0]
print(f"SONAR {src_lang} -> {tgt_lang}: {decoded}")
encoder = SONARForText2Text.from_pretrained("tutur90/SONAR-Text-to-Text")
encoder.set_encoder_only() # Delete decoder to save memory, this options is not needed
inputs = tokenizer(sentence, langs=src_lang, return_tensors="pt")
embeddings = encoder.encode(**inputs)
decoder = SONARForText2Text.from_pretrained("tutur90/SONAR-Text-to-Text")
decoder.set_decoder_only() # Same
decoded = decoder.decode(
encoder_outputs,
target_lang_ids=[tokenizer.convert_tokens_to_ids(tgt_lang)],
max_length=128,
num_beams=1,
do_sample=False,
)
decoded = tokenizer.batch_decode(decoded, skip_special_tokens=True)[0]