Downloads · 30 days
28.5K
6% of all-time downloads
Twitter/twhin-bert-base
twhin-bert-base is a fill-mask model from Twitter. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
[](http://makeapullrequest.com) [](https://arxiv.org/abs/2209.07562)
Downloads · 30 days
28.5K
6% of all-time downloads
All-time downloads
516K
Public
Parameters
279M
2.2 GB on disk
Likes
46
Public
Click a slice to open those files.
.bin1.1 GB · 50%
How the weights are stored.
F32279M · 100%
From the Hugging Face model README
This repo contains models, code and pointers to datasets from our paper: TwHIN-BERT: A Socially-Enriched Pre-trained Language Model for Multilingual Tweet Representations. [PDF] [HuggingFace Models]
TwHIN-BERT is a new multi-lingual Tweet language model that is trained on 7 billion Tweets from over 100 distinct languages. TwHIN-BERT differs from prior pre-trained language models as it is trained with not only text-based self-supervision (e.g., MLM), but also with a social objective based on the rich social engagements within a Twitter Heterogeneous Information Network (TwHIN).
TwHIN-BERT can be used as a drop-in replacement for BERT in a variety of NLP and recommendation tasks. It not only outperforms similar models semantic understanding tasks such text classification), but also social recommendation tasks such as predicting user to Tweet engagement.
We initially release two pretrained TwHIN-BERT models (base and large) that are compatible wit the HuggingFace BERT models.
| Model | Size | Download Link (🤗 HuggingFace) |
|---|---|---|
| TwHIN-BERT-base | 280M parameters | Twitter/TwHIN-BERT-base |
| TwHIN-BERT-large | 550M parameters | Twitter/TwHIN-BERT-large |
To use these models in 🤗 Transformers:
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained('Twitter/twhin-bert-base')
model = AutoModel.from_pretrained('Twitter/twhin-bert-base')
inputs = tokenizer("I'm using TwHIN-BERT! #TwHIN-BERT #NLP", return_tensors="pt")
outputs = model(**inputs)
<!-- ## 2. Set up environment and data
### Environment
TBD
## 3. Fine-tune TwHIN-BERT
TBD -->
If you use TwHIN-BERT or out datasets in your work, please cite the following:
@article{zhang2022twhin,
title={TwHIN-BERT: A Socially-Enriched Pre-trained Language Model for Multilingual Tweet Representations},
author={Zhang, Xinyang and Malkov, Yury and Florez, Omar and Park, Serim and McWilliams, Brian and Han, Jiawei and El-Kishky, Ahmed},
journal={arXiv preprint arXiv:2209.07562},
year={2022}
}