Downloads · 30 days
0
blakes/form-field-question-extraction
form-field-question-extraction is a machine learning model from blakes. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A complete pipeline that extends FFDNet (form field detection) with LayoutLMv3 (text understanding) to programmatically identify each form field, its associated question/prompt, and answer options.
Downloads · 30 days
0
Access
Public
Updated Apr 23, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.py18.1 KB · 65%
From the Hugging Face model README
A complete pipeline that extends FFDNet (form field detection) with LayoutLMv3 (text understanding) to programmatically identify each form field, its associated question/prompt, and answer options.
This project builds on the CommonForms paper (arXiv:2509.16506) by Joe Barrow, which introduced:
While FFDNet detects where form fields are (text inputs, checkboxes, signatures), it does not tell you what question each field answers. This pipeline:
{
"num_fields": 3,
"fields": [
{
"field_bbox": [285, 269, 442, 319],
"field_type": "Text Input",
"confidence": 0.92,
"question_text": "Full Name:",
"question_bbox": [100, 269, 275, 319],
"answer_options": []
},
{
"field_bbox": [500, 400, 530, 430],
"field_type": "Choice Button",
"confidence": 0.88,
"question_text": "Do you agree?",
"question_bbox": [300, 400, 490, 430],
"answer_options": [
{"text": "Yes", "bbox": [540, 400, 580, 430]},
{"text": "No", "bbox": [600, 400, 640, 430]}
]
}
]
}
Input PDF/Image
|
v
[EasyOCR] -----> words + bboxes
| |
v v
[FFDNet-L] [LayoutLMv3]
| |
v v
field bboxes token labels (Q/A/H/O)
| |
+-------> [Spatial Matching]
|
v
field -> question link
|
v
answer options
pip install transformers datasets evaluate easyocr Pillow torch numpy opencv-python-headless
pip install ultralytics # optional, for FFDNet field detection
from form_understanding_demo import FormUnderstandingPipeline
from PIL import Image
# Initialize pipeline (auto-detects CPU/GPU)
pipeline = FormUnderstandingPipeline(
layoutlmv3_model="nielsr/layoutlmv3-finetuned-funsd",
device="cuda", # or "cpu"
)
# Process an image
image = Image.open("form.png").convert("RGB")
# Option A: Provide pre-detected field boxes from FFDNet
field_boxes = [
{"bbox": [100, 200, 300, 220], "class_name": "Text Input", "confidence": 0.95},
{"bbox": [100, 300, 130, 330], "class_name": "Choice Button", "confidence": 0.92},
]
result = pipeline.process_image(image, field_boxes=field_boxes)
# Option B: Without field detection (text classification only)
result = pipeline.process_image(image)
# Print results
for field in result["fields"]:
print(f" [{field['field_type']}] Question: '{field['question_text']}'")
if field["answer_options"]:
print(f" Options: {[o['text'] for o in field['answer_options']]}")
# Process a single image
python form_understanding_demo.py form.png -o result.json
# Process a PDF (requires pdf2image)
python form_understanding_demo.py form.pdf -o result.json
# With pre-detected field boxes (from FFDNet or other detector)
python form_understanding_demo.py form.png --fields detected_boxes.json -o result.json
| Component | Model | HF Hub | Purpose |
|---|---|---|---|
| Field Detection | FFDNet-L | jbarrow/FFDNet-L | Detect form field boundaries (text input, checkbox, signature) |
| Field Detection | FFDNet-S | jbarrow/FFDNet-S | Faster variant (5ms/page vs 16ms) |
| Text Classification | LayoutLMv3 | nielsr/layoutlmv3-finetuned-funsd | Classify tokens as Question/Answer/Header/Other |
| Base Model | LayoutLMv3-base | microsoft/layoutlmv3-base | LayoutLMv3 base for fine-tuning |
| OCR | EasyOCR | N/A | Extract word-level text from images |
jbarrow/CommonForms): 480k form pages with field bounding boxes (no text annotations)nielsr/funsd-layoutlmv3): 199 annotated forms with question/answer/header labelsform_understanding_demo.py — Complete inference pipelineinference_form_understanding.py — Core pipeline classes (FFDNet + LayoutLMv3 + matching)test_pipeline_demo.py — Test script on FUNSD examplestrain_layoutlmv3_funsd_v2.py — Fine-tune LayoutLMv3 on FUNSDpreprocess_commonforms_for_layoutlmv3.py — Preprocess CommonForms with OCR for pseudo-labelingFFDNet (based on YOLO11) detects three classes of form fields:
Trained from scratch on CommonForms at 1216px resolution.
LayoutLMv3 token classification model fine-tuned on FUNSD classifies each word:
For each detected field bounding box, the pipeline:
To improve question detection on your own form dataset:
python train_layoutlmv3_funsd_v2.py
This fine-tunes microsoft/layoutlmv3-base on FUNSD token classification. Key hyperparameters:
The pre-trained LayoutLMv3 (nielsr/layoutlmv3-finetuned-funsd) achieves ~90.8% F1 on FUNSD token classification. Our pipeline on FUNSD test images achieves:
@article{barrow2025commonforms,
title={CommonForms: A Large, Diverse Dataset for Form Field Detection},
author={Barrow, Joe},
journal={arXiv preprint arXiv:2509.16506},
year={2025}
}
The CommonForms dataset and FFDNet models are licensed under Apache 2.0. This pipeline code is released under MIT License.