Downloads · 30 days
24
6% of all-time downloads
BEE-spoke-data/mega-encoder-small-16k-v1
mega-encoder-small-16k-v1 is a fill-mask model from BEE-spoke-data. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as artistic-2.0.
This is a "huggingface-native" pretrained encoder-only model with 16384 context length. The model architecture is MEGA.
Downloads · 30 days
24
6% of all-time downloads
All-time downloads
420
Public
Parameters
122M
488 MB on disk
Likes
4
Public
Click a slice to open those files.
.safetensors488 MB · 99%
From the Hugging Face model README
This is a "huggingface-native" pretrained encoder-only model with 16384 context length. The model architecture is MEGA.
Despite being a long-context model evaluated on a short-context benchmark, MEGA holds up decently:
| Model | Size | CTX | Avg |
|---|---|---|---|
| mega-encoder-small-16k-v1 | 122M | 16384 | 0.777 |
| bert-base-uncased | 110M | 512 | 0.7905 |
| roberta-base | 125M | 514 | 0.86 |
| bert-plus-L8-4096-v1.0 | 88.1M | 4096 | 0.8278 |
| mega-wikitext103 | 7.0M | 10000 | 0.48 |
| Model | Size | CTX | Avg | CoLA | SST2 | MRPC | STSB | QQP | MNLI | QNLI | RTE |
|---|---|---|---|---|---|---|---|---|---|---|---|
| mega-encoder-small-16k-v1 | 122M | 16384 | 0.777 | 0.454 | 0.914 | 0.8404 | 0.906 | 0.894 | 0.806 | 0.842 | 0.556 |
| bert-base-uncased | 110M | 512 | 0.7905 | 0.521 | 0.935 | 0.889 | 0.858 | 0.712 | 0.84 | 0.905 | 0.664 |
| roberta-base | 125M | 514 | 0.86 | 0.64 | 0.95 | 0.9 | 0.91 | 0.92 | 0.88 | 0.93 | 0.79 |
| bert-plus-L8-4096-v1.0 | 88.1M | 4096 | 0.8278 | 0.6272 | 0.906 | 0.8659 | 0.9207 | 0.906 | 0.832 | 0.9 | 0.6643 |
| mega-wikitext103 | 7M | 10000 | 0.480 | 0.00 | 0.732 | 0.748 | -0.087 | 0.701 | 0.54 | 0.598 | 0.513 |
The evals for MEGA/bert-plus can be found in this open wandb project and are taken as the max observed values on the validation sets. The values for other models are taken as reported in their papers.
</details>This encoder model has 8 layers, hidden size 768, and a feedforward ratio of 3x. The resulting total size is 122M params.
<details> <summary><strong>Architecture Details</strong></summary>Details:
"simple" relative positional embeddings instead of the rotary embeddings touted in the paper.
facebook/bart-large
This model was trained with the transformers package. You can find (mostly unorganized) training runs on wandb here.
<details> <summary><strong>Training Details</strong></summary>This is a pretrained model intended to be fine-tuned on various encoder-compatible tasks. However, if you are interested in testing inference with this model or have a deep passion for predicting mask tokens, you can use the following code:
import json
from transformers import pipeline
pipe = pipeline("fill-mask", model="BEE-spoke-data/mega-encoder-small-16k-v1")
text = "I love to <mask> memes."
result = pipe(text)
print(json.dumps(result, indent=2))
If fine-tuning this model on <task>, using gradient checkpointing makes training at 16384 context quite feasible. By installing the transformers fork below and passing gradient_checkpointing=True in the training args, you should be able to finetune at batch size 1 with VRAM to spare on a single 3090/4090.
pip uninstall -y transformers
pip install -U git+https://github.com/pszemraj/transformers.git@mega-gradient-checkpointing
pip install -U huggingface-hub
if there is sufficient interest, we can look at making a PR into the official repo.
if you find this useful, please consider citing this DOI, it would make us happy.
@misc{beespoke_data_2024,
author = {Peter Szemraj and Vincent Haines and {BEEspoke Data}},
title = {mega-encoder-small-16k-v1 (Revision 1476bcf)},
year = 2024,
url = {https://huggingface.co/BEE-spoke-data/mega-encoder-small-16k-v1},
doi = {10.57967/hf/1837},
publisher = {Hugging Face}
}