Downloads · 30 days
12
6% of all-time downloads
howey/HDT-E
HDT-E is a machine learning model from howey. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
To use the pre-trained model for masked language modeling, use the following snippet:
Downloads · 30 days
12
6% of all-time downloads
All-time downloads
209
Public
Parameters
108M
431 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors431 MB · 100%
From the Hugging Face model README
To use the pre-trained model for masked language modeling, use the following snippet:
from transformers import AutoModelForMaskedLM, AutoTokenizer
# See the `MDLM` collection page on the hub for list of available models.
tokenizer = transformers.AutoTokenizer.from_pretrained('howey/HDT-E')
model_name = 'howey/HDT-E'
model = AutoModelForMaskedLM.from_pretrained(model_name)
For more details, please see our github repository: HDT
The model, which has a context length of 8192 and is similar in size to BERT with approximately 110M parameters,
was trained on standard masked language modeling task with a Transformer-based architecture using our proposed hierarchical attention.
The training regimen comprised 24 hours on the ArXiv+Wikipedia+HUPD corpus, involving the processing of a total of 1.3 billion tokens.
For more details, please see our paper: HDT: Hierarchical Document Transformer.
Please cite our work using the bibtex below:
BibTeX:
@inproceedings{He2024COLM,
title={HDT: Hierarchical Document Transformer},
author={Haoyu He and Markus Flicke and Jan Buchmann and Iryna Gurevych and Andreas Geiger},
year={2024},
booktitle={Conference on Language Modeling}
}
Haoyu ([email protected])