Downloads · 30 days
0
gurvgupta/LayoutLM_rvl-cdip
LayoutLM_rvl-cdip is a machine learning model from gurvgupta. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Multimodal (text + layout/format + image) pre-training for document AI
Downloads · 30 days
0
Access
Public
Updated Jul 4, 2024
Repo size
451 MB
Likes
1
Public
Click a slice to open those files.
.pt451 MB · 100%
From the Hugging Face model README
Multimodal (text + layout/format + image) pre-training for document AI
Microsoft Document AI | GitHub
LayoutLM is a simple but effective pre-training method of text and layout for document image understanding and information extraction tasks, such as form understanding and receipt understanding. LayoutLM archives the SOTA results on multiple datasets. For more details, please refer to research paper:
LayoutLM: Pre-training of Text and Layout for Document Image Understanding Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, Ming Zhou, KDD 2020
I fine tuned the LayoutLM-Base, Uncased (11M documents, 2 epochs): 12-layer, 768-hidden, 12-heads, 113M parameters on RVL-CDIP Dataset(Total 3200 images : 1940 train, 640 test, 640 valid) for 50 epochs.
@misc{xu2019layoutlm,
title={LayoutLM: Pre-training of Text and Layout for Document Image Understanding},
author={Yiheng Xu and Minghao Li and Lei Cui and Shaohan Huang and Furu Wei and Ming Zhou},
year={2019},
eprint={1912.13318},
archivePrefix={arXiv},
primaryClass={cs.CL}
}