Downloads · 30 days
0
kaist-ai/langbridge_encoder_tokenizer
langbridge_encoder_tokenizer is a machine learning model from kaist-ai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
- Repository: https://github.com/kaistAI/LangBridge - Paper: LangBridge: Multilingual Reasoning Without Multilingual Supervision - Point of Contact: [email protected] 🤔LMs good at reasoning are mostly English-centri…
Downloads · 30 days
0
Access
Public
Updated Jan 29, 2024
Repo size
20.6 MB
Likes
3
Public
Click a slice to open those files.
.json16.3 MB · 79%
From the Hugging Face model README
🤔LMs good at reasoning are mostly English-centric (MetaMath, Orca 2, etc).
😃Let’s adapt them to solve multilingual tasks. BUT without using multilingual data!
LangBridge “bridges” mT5 encoder and the target LM together while utilizing only English data. In test time, LangBridge models can solve multilingual reasoning tasks effectively.

This is the tokenizer used for the encoder models of LangBridge. LangBridge models require two tokenizers, one for the encoder model and one for the language model. To the best of our knowledge there is no way of uploading two tokenizers for a model. So this seperate repository was created.
Please refer to the Github repository for detailed usage examples.
Check out other LangBridge models.
We have:
If you find the following model helpful, please consider citing our paper!
BibTeX:
@misc{yoon2024langbridge,
title={LangBridge: Multilingual Reasoning Without Multilingual Supervision},
author={Dongkeun Yoon and Joel Jang and Sungdong Kim and Seungone Kim and Sheikh Shafayat and Minjoon Seo},
year={2024},
eprint={2401.10695},
archivePrefix={arXiv},
primaryClass={cs.CL}
}