Downloads · 30 days
0
venki101/Venket
Venket is a machine learning model from venki101. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A modular, deterministic, configuration-driven pipeline that takes scanned bank document templates with blank signature fields, fills them with realistic handwritten signatures drawn from a public signature dataset, a…
Downloads · 30 days
0
Access
Public
Updated Aug 5, 2026
Repo size
560 MB
Likes
0
Public
Click a slice to open those files.
.png661 MB · 94%
From the Hugging Face model README
A modular, deterministic, configuration-driven pipeline that takes scanned bank document templates with blank signature fields, fills them with realistic handwritten signatures drawn from a public signature dataset, applies configurable physical/scanner degradation, and exports both the rendered document and a comprehensive per-sample metadata JSON — suitable for training and benchmarking document-AI / signature-verification models.
python main.py --config configs/generate.yaml --count 10000
Requires Python 3.12+ (developed and tested against 3.11.9 as well; no 3.12-only syntax is used, so 3.11 works too).
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
Core dependencies: OpenCV, NumPy, SciPy, Pillow, scikit-image, Shapely, PyYAML, Pydantic, Albumentations, OmegaConf, Matplotlib, tqdm, pytest.
cc/
├── image/ # 9 scanned bank document templates (provided)
├── dataset/ # Public handwritten signature dataset (provided)
├── annotations/ # Signature box annotations, one JSON per template
├── configs/
│ ├── default.yaml # Every config field with its default value
│ ├── generate.yaml # Example large-batch dataset config
│ └── examples/ # Focused example configs (see below)
├── output/ # Generated PNG + JSON land here
├── src/
│ ├── config/ # Pydantic schema + OmegaConf loader
│ ├── templates/ # TemplateLoader, annotation schema, CV box detector
│ ├── signatures/ # SignatureDataset discovery/sampling
│ ├── rendering/ # SignaturePreprocessor, Renderer (blending)
│ ├── placement/ # PlacementEngine
│ ├── augmentation/ # AugmentationEngine + effect implementations
│ ├── analysis/ # GeometryAnalyzer, DifficultyAnalyzer, ValidationEngine
│ ├── verification/ # Verification test-case registry + CASE_TABLE (case_type 1a-5b)
│ ├── metadata/ # MetadataExporter
│ ├── visualization/ # Debug figure rendering
│ ├── utils/ # Logging, seeding, geometry, I/O, names
│ └── generator.py # DatasetGenerator orchestrator (+ multiprocessing)
├── tests/ # pytest unit + integration tests
├── main.py # CLI entry point
└── requirements.txt
This repository ships with:
image/ — 9 real scanned bank document templates (guarantee letters,
loan agreements, account mandates), auto-discovered recursively regardless
of subfolder layout. Each gets a stable id bank_doc_01 … bank_doc_09,
assigned by a deterministic sort of their discovered paths.dataset/ — A public handwritten signature dataset (CEDAR-style:
genuine/original_<signer>_<sample>.png and
forged/forgeries_<signer>_<sample>.png), auto-discovered recursively.
Nested dataset/personNNN/*.png-style layouts are also supported — the
loader falls back to using each signature's parent folder name as its
signer group when the CEDAR filename pattern doesn't match, and to an
auto-generated id when no identity can be inferred at all. An optional
dataset/metadata.csv with path/filename + signer/source_type
columns can override both the identity association and genuine/forged
labeling for any file, plus an optional script column ("latin" by
default) recording the handwriting script a signature is written in.
dataset/hindi_bengali/ + the matching rows in dataset/metadata.csv
bundle 7 real, named, multi-script signers (Hindi/Devanagari-style,
Bengali-style, and Latin cursive) alongside the 55 anonymous CEDAR
identities, so generated samples aren't exclusively Latin-script English
names — SignerIdentityRegistry.display_name() (src/verification/registry.py)
returns a metadata.csv-sourced name as-is instead of generating one, and
every SignatureRecord/verification metadata block carries its script.
These signers have very few reference samples each (1-3, no forged
counterpart) — records_for_group(..., "forged") falls back to their
genuine samples with a logged warning rather than failing, so the 3b
skilled-forgery case degrades gracefully rather than erroring for them.Nothing about either directory's internal layout is hardcoded — add more template images or more signature images/folders and they will be picked up automatically on the next run.
# 1. Generate 10 samples with all defaults (configs/default.yaml)
python main.py
# 2. Generate a large, varied batch
python main.py --config configs/generate.yaml --count 10000
# 3. Generate one fully reproducible, fixed-identity sample with a debug figure
python main.py --config configs/examples/single_document_demo.yaml --save-visualization
# 4. Override arbitrary nested config values from the command line
python main.py --set placement.rotation_deg.max=6 --set augmentation.coffee_stain.enabled=true
# 5. Restrict to one template, use 8 worker processes, don't overwrite existing files
python main.py --template bank_doc_03 --count 5000 --workers 8
Every run writes output/<prefix>_<NNNNNN>.png + .json pairs, plus
output/generation_summary.json (batch statistics) and
output/resolved_config.yaml (the fully-resolved config actually used, for
provenance/reproduction).
| Flag | Effect |
|---|---|
--config PATH | User YAML merged on top of configs/default.yaml |
--count N | Number of samples |
--seed N | Master random seed |
--workers N | Worker process count (1 = sequential, in-process) |
--output DIR | Output directory |
--template ID | Force every sample to use one template (e.g. bank_doc_03) |
--start-index N | First sample number (for resuming/sharding a batch) |
--overwrite | Regenerate samples even if their files already exist |
--save-visualization | Also write a debug figure per sample |
--set key.path=value | Arbitrary dot-path config override (repeatable) |
--log-level LEVEL | DEBUG / INFO / WARNING / ERROR |
Configuration is layered: configs/default.yaml (every field, documented) →
an optional --config user YAML (merged on top) → --set key=value CLI
overrides (highest precedence). The merged result is validated against a
Pydantic schema (src/config/schema.py) — invalid types, out-of-range
values, or unknown keys fail fast with a clear error instead of silently
being ignored.
Any placement/augmentation numeric parameter accepts either a fixed scalar or a range, resolved independently per sample by a seeded RNG:
placement:
scale_factor: 0.84 # fixed
rotation_deg:
min: -5
max: 5 # sampled uniformly per sample
Every augmentation effect follows the same shape:
augmentation:
coffee_stain:
enabled: true
strength: 0.42 # or {min: ..., max: ...}
probability: 0.1 # per-sample chance of actually applying
seed: null # optional explicit seed; default derives from (master_seed, sample_index, name)
These effects are covered by 20 configurable families (a couple are
strict specializations of another and share one family with a type/tone
selector, so the config surface stays coherent):
| Family | Covers |
|---|---|
rotation | Extra probabilistic jitter on top of placement.rotation_deg |
stroke_dilation | Ink stroke thickness variation |
motion_blur, gaussian_blur | Signature-layer blur |
perspective | Whole-page perspective warp |
scanner_noise, salt_pepper | Sensor/impulse noise |
jpeg_compression | Lossy re-encode artifacts (applied last) |
contrast, gamma | Photometric shifts |
coffee_stain | Coffee-ring blotches |
stain (type: paper|water|random) | Paper stains + water stains |
shadow (type: edge|corner|fold|random) | Edge / corner / fold shadows |
crumple | Paper crumpling + wrinkle lines |
discoloration (tone: yellow|grey|sepia|random) | General discoloration + paper yellowing |
page_tilt | Crooked scanner feed |
dust, scanner_streaks | Sensor debris / dirty-glass streaks |
ink_fading, ink_bleed | Ink density effects |
See configs/default.yaml for every field and its default, and
configs/generate.yaml / configs/examples/*.yaml for worked examples
(a large varied batch, a fixed-identity single-document demo, an "easy"
preset, and a "hard/degraded" preset).
Each sample flows through the same 11 stages the spec describes, one class per stage:
| # | Stage | Class |
|---|---|---|
| 1 | Load template | TemplateLoader |
| 2 | Resolve signature box(es) | TemplateLoader (manual annotation, or CV auto-detect fallback) |
| 3 | Assign signer name/role | src.utils.names + config |
| 4 | Select a signature image | SignatureDataset |
| 5 | Clean/preprocess signature | SignaturePreprocessor |
| 6 | Fit & position in box | PlacementEngine |
| 7 | Signature-layer augmentation | AugmentationEngine.apply_to_signature |
| 8 | Blend onto document | Renderer |
| 9 | Document-level augmentation | AugmentationEngine.apply_to_document |
| 10 | Detect rendered signature + geometry | GeometryAnalyzer |
| 11 | Difficulty score | DifficultyAnalyzer |
| — | Export | MetadataExporter |
| — | Verify | ValidationEngine |
GeneratorContext (in src/generator.py) wires all of the above together
for one sample; DatasetGenerator drives it across a batch, either
sequentially (workers=1) or via a ProcessPoolExecutor where each worker
builds one GeneratorContext (via a pool initializer) and reuses it for
every sample it's assigned.
SignaturePreprocessor)Signature dataset images are flat grayscale/RGB scans with no alpha channel. The preprocessor estimates the local paper background level, derives a soft per-pixel alpha from how dark each pixel is relative to that background (gamma-softened to preserve antialiased stroke edges/natural texture rather than a hard binary mask), removes sub-pixel speckle noise via connected-component filtering, and tightly crops to the ink content. A flat ink color (randomly chosen per sample from a small realistic pool — black, blue-black, ballpoint blue) is applied to the RGB channels; the alpha channel carries all of the texture.
PlacementEngine)Rotates first (about the signature's own center, on an auto-expanding
transparent canvas so nothing clips), then scales to fit the box's
margin-adjusted interior while preserving aspect ratio (never stretches),
then positions via an anchor point (center, top_left, …) plus
pixel offsets. A configurable misplacement_probability optionally kicks
the signature away from its anchored position by up to
misplacement_strength * max(box_w, box_h) in a random direction — the
mechanism behind intentionally "hard" (badly-placed) samples.
A signature is never allowed to shrink below a legible minimum height
(16px) even inside a very short box; on the real bundled templates several
signature lines are only 12-24px tall, and a naive fit-to-box scale would
otherwise render an invisible, sub-pixel signature.
Renderer)Supports multiply, darken, and alpha blend modes, configurable ink
opacity, alpha-edge feathering (feather_px), an overall blur to match scan
resolution (blur_sigma), a light Gaussian sensor-noise pass baked into
every render, and a final mild anti-aliasing blur.
GeometryAnalyzer)Detection is scoped to the known signature field (the annotated box, and,
if nothing is found there, a widened window around the union of the box and
the pre-augmentation placement location) rather than searching the whole
page — exactly how a real pipeline would use its layout annotations, and
what makes the box-vs-detection coverage metrics meaningful in the first
place. Ink is separated from the local paper background via adaptive
thresholding; a heuristic filter discards thin, wide, solid components
(printed "sign here" rule lines) so a blank field is never mistaken for a
signature. From the resulting ink mask it computes: axis-aligned bounding
box, cv2.minAreaRect minimum rotated rectangle, a convex-hull polygon,
centroid, PCA-based orientation, ink pixel count, box_coverage_pct,
containment_pct, ink_fill_ratio, polygon_overlap_proxy (IoU),
distance from box center, and margin utilization.
One JSON per sample, always containing the full schema (fields for augmentations that weren't applied are present with a neutral value, not omitted):
{
"document_id": "sample_000001",
"template": "bank_doc_03",
"document_type": "deed_of_guarantee",
"page": 1, "page_size": [1024, 559], "dpi": 300,
"signer": {"name": "Jane Doe", "role": "Guarantor (Authorized Signatory)"},
"signature_source": {"signature_id": "sig_001437", "path": "dataset/genuine/original_14_7.png",
"signer_group": "signer_0014", "source_type": "genuine"},
"placement": {"x_offset": 8.0, "y_offset": -4.0, "rotation_deg": -3.0, "scale_factor": 0.84,
"anchor_point": "center", "misplaced": false, ...},
"rendering": {"blend_mode": "multiply", "ink_opacity": 0.91, "blur_sigma": 0.34, ...},
"difficulty": {"score": 0.57, "tier": "medium", "misplacement": 0.38, "whitespace": 0.14,
"faintness": 0.45, "smallness": 0.31, "low_texture": 0.12},
"quality": {"degraded": false, "scanner_quality": 0.85, "paper_quality": 0.85, "render_quality": 0.85},
"verification": {
"case_type": "3a", "verdict": "REJECT", "reject_code": "IDENTITY_MISMATCH",
"target_box_id": "sig1", "expected_signer_group": "signer_0014", "expected_signer_script": "latin",
"actual_signer_group": "signer_0027", "actual_source_type": "genuine", "actual_signer_script": "latin",
"companion": null
},
"augmentation": {
"rotate_deg": -2, "scale_factor": 0.84, "stroke": 0.12, "noise_amount": 0.06,
"stain_type": null, "stain_strength": 0.0, "shadow_type": "corner", "shadow_strength": 0.55,
"coffee_stain_strength": 0.42, "crumple_strength": 0.18, "tilt_deg": 1.3,
"discoloration_strength": 0.0, "blur_sigma": 0.35, "jpeg_quality": 82, "scanner_noise": 0.02,
"effects": { "...every single effect, always present, with enabled/applied/strength/probability/seed/params...": {} }
},
"geometry": {
"box_px": [1200, 830, 1530, 950], "target_box": [...], "actual_box": [...],
"signature_bbox_px": [...], "signature_polygon_px": [[x, y], ...], "convex_hull_px": [...],
"min_rotated_rect": {"center": [...], "size": [...], "angle_deg": ...},
"centroid_px": [...], "orientation_deg": ..., "ink_pixels": ..., "detected": true
},
"metrics": {
"box_coverage_pct": 84.2, "containment_pct": 96.3, "ink_fill_ratio": 0.28,
"polygon_overlap_proxy": 0.71, "distance_from_box_center_px": 12.4, "margin_utilization": 0.6,
"mean_dark_intensity": 142.1, "gray_stddev": 38.6, "signature_area_ratio": 0.011
},
"provenance": {"master_seed": 1234, "sample_index": 1}
}
The augmentation.effects.<name> block is the full audit trail for every
one of the 20 effect families: whether it was enabled in config, whether
its probability roll actually applied it, the resolved strength, and
any effect-specific params (e.g. {"resolved_type": "corner"} for
shadow, {"quality": 58} for JPEG compression) — this is present even when
applied: false, so the schema is identical across every sample regardless
of what happened to fire.
quality.degraded is true iff quality.case_type == "1d" or
quality.notes contains the substring "low-quality" (case-insensitive) —
both driven by quality.case_type / quality.notes in config (see
configs/examples/hard_degraded.yaml). verification is present (with
every field null) on every sample; see the next section for when it's
populated.
Setting quality.case_type to one of 12 recognized codes switches sample
generation into scenario-driven mode: instead of a random genuine
signature in a random box, the pipeline renders the specific scenario the
code describes and stamps the sample's verification metadata block with
the matching ground-truth verdict/reject_code — producing a labeled
benchmark for a downstream verifier. Any other value (or leaving it unset)
behaves exactly as before.
| Code | Scenario | Verdict | Reject code |
|---|---|---|---|
1a | Correct signer signs normally | PASS | - |
1b | Correct signer, natural pen variation (a different genuine sample of the same person) | PASS | - |
1c | Multi-signer form, both boxes sign correctly | PASS | - |
1d | Correct signer, scan degraded (pair with configs/examples/hard_degraded.yaml) | PASS | - |
2a | Box left completely blank | REJECT | NO_SIGNATURE |
2b | Printed/typed name only, no handwriting | REJECT | NO_SIGNATURE |
3a | Wrong real person signs the box (a different genuine signer) | REJECT | IDENTITY_MISMATCH |
3b | Skilled-forgery sample used (CEDAR forged/ sample of the authorized signer) | REJECT | IDENTITY_MISMATCH |
4a | Authorized signer, but wrong role's box (the other box's authorized signer signs here instead) | REJECT | ROLE_MISMATCH |
4b | Signer from a different form entirely (authorized elsewhere, not on this template) | REJECT | ROLE_MISMATCH |
5a | Required signer's box left empty (while a companion box is correctly signed) | REJECT | REQUIRED_SIGNER_ABSENT |
5b | Someone else signs in the required signer's place (while a companion box is correctly signed) | REJECT | REQUIRED_SIGNER_ABSENT |
python main.py --config configs/examples/verification_cases.yaml --set quality.case_type=3a --count 10
python scripts/generate_verification_suite.py --per-case 2 # one labeled batch covering all 12 codes
src/verification/registry.py's SignerIdentityRegistry deterministically
assigns an "authorized signer" identity (one of the discovered CEDAR signer
groups) to every (template_id, box_id) pair, purely as a function of the
master seed — there's no persisted state, consistent with the rest of the
pipeline's seeding model. Every case type is defined declaratively in
src/verification/case_types.py (CASE_TABLE) as what should actually be
rendered (signature_mode: genuine-correct, blank, printed-text,
wrong-signer, skilled-forgery, role-swap, foreign-template-signer, ...) plus
whether a companion box on the same page should also be filled with its own
correctly-authorized signature (1c/5a/5b — the cases whose label only
makes sense in a multi-signer-page context). The verification metadata
block records both the expected and actual signer group so the label is
independently auditable, not just asserted.
Companion-box rendering exists purely for visual/contextual realism (so a
"required signer absent" sample actually shows a properly multi-signed
page); it does not get its own geometry/difficulty block — the schema
stays single-target-box-centric, same as it's always been. signature_source
is null for the two blank scenarios (2a/5a); the expected identity in
that case lives only in verification.expected_signer_group.
Computed exactly as specified, from the post-render, post-augmentation detected geometry (a pure function — no randomness — so it is always reproducible from its own recorded components):
misplacement = clamp(1 - polygon_overlap_proxy, 0, 1)
whitespace = clamp(whitespace_ratio, 0, 1) # whitespace_ratio = 1 - ink_fill_ratio
faintness = clamp((mean_dark_intensity - 90) / 120, 0, 1)
smallness = clamp(1 - signature_area_ratio / 0.04, 0, 1)
low_texture = clamp(1 - gray_stddev / 60, 0, 1)
score = 0.35*misplacement + 0.20*whitespace + 0.20*faintness + 0.15*smallness + 0.10*low_texture
tier = "easy" if score < 0.33 else "medium" if score < 0.66 else "hard"
A note on the bundled templates: smallness uses a fixed reference of
4% of page area. On the 9 real bank forms shipped in image/, the actual
signature lines are (realistically) much smaller than that — typically
0.2%–1.5% of the page — so smallness, and with it the overall score, runs
structurally high (mostly medium/hard) even under the "clean" example
preset. This is an accurate reflection of how small real bank-form
signature fields are relative to a full page, not a bug in the scoring
formula, which is intentionally implemented exactly as specified (weights
and the 0.04 reference are not something the pipeline silently retunes).
If your use case wants a different-looking tier distribution, adjust
difficulty.easy_threshold / difficulty.hard_threshold in config, or
supply larger signature boxes in your own annotations.
Two modes, in order of preference:
Mode 2 — manual annotation (preferred, used for all 9 bundled
templates): annotations/<template_id>.json:
{
"id": "bank_doc_03",
"source_image": "image/guarantee/image.png",
"document_type": "deed_of_guarantee",
"pages": [{
"page": 1, "page_size": [1024, 559], "dpi": 300,
"signature_boxes": [
{"id": "sig1", "x1": 595, "y1": 399, "x2": 748, "y2": 431,
"role": "Guarantor (Authorized Signatory)", "name": "Guarantor"}
]
}]
}
The 9 bundled annotation files were hand-measured against each template image (using a coordinate-grid overlay for precision) — every signature field, its role, and page geometry are captured exactly.
Mode 1 — automatic CV detection (fallback): if no
annotations/<template_id>.json exists for a discovered template,
src/templates/box_detector.py looks for long horizontal "sign here"
rule lines (morphological opening + contour filtering) and places a
candidate box directly above each one. The result is cached back to
annotations/<template_id>.json so detection only ever runs once per
template — subsequent runs (including a fresh checkout with new
templates dropped into image/) reuse the cached annotation. If a
template has no such lines at all, a generic lower-right fallback box is
used so the pipeline never crashes on an unannotated template.
To add a 10th template: drop its image into any subfolder of image/ and
either let auto-detection run once, or hand-author its annotation file
(fastest way: temporarily overlay a coordinate grid on the image — see the
_grid_debug pattern used during development — and read off box corners).
Every random draw anywhere in the pipeline comes from an RNG whose seed is
sha256(master_seed | sample_index | stage_name) (src/utils/random_utils.py),
never from a shared/global RNG. Consequences:
(seed, config) always produces byte-identical PNGs and
semantically identical JSON, regardless of --workers or scheduling order
(verified in tests/test_end_to_end.py, and by hand: a workers=1 run
and a workers=4 run of the same seed produce cmp-identical PNGs).seed field in its metadata
record is independently derivable/overridable — set augmentation.<name>.seed
explicitly in config to pin one specific effect while leaving everything
else seed-derived.output/resolved_config.yaml captures the exact fully-resolved
configuration used for a run, so it can be handed to --config later to
reproduce that run exactly (given the same --seed).ValidationEngine runs (by default) after every sample and checks:
metadata schema completeness, the signature was actually detected, its
polygon is non-degenerate and geometrically valid, every bounding box is
well-formed, boxes lie on the page (a warning, or an error under
validation.strict: true), no NaN/Inf leaked into the JSON, and the
difficulty score/tier are exactly reproducible from their own recorded
components. Per-sample pass/fail rolls up into
output/generation_summary.json. Disable with generation.run_validation: false.
--save-visualization writes output/visualizations/<id>_viz.png per
sample: the full rendered document with the target box (green), detected
bounding box (blue), and detected polygon (red) overlaid; a crop of the
signature field; the reconstructed detected-ink mask; a coverage-metrics
readout; and a difficulty-component bar chart.
python -m pytest -q
80 tests covering geometry math (coverage/containment/IoU/margin
utilization), the exact difficulty formula (including weight-sum
validation and clamping), seed derivation/independence, placement
(determinism, aspect-ratio preservation, the minimum-height floor,
misplacement gating), signature/template discovery against the real
bundled assets (including the multi-script named signer pool), ink
detection (including the printed-rule-line false-positive guard), metadata
validation, full end-to-end reproducibility, and the case_type-driven
verification scenarios (tests/test_verification_cases.py: registry
determinism, ground truth for all 12 codes, multi-script signer handling,
and backward compatibility with the non-case-type path).
src/augmentation/document_effects.py or signature_effects.py, a field
to AugmentationsConfig in src/config/schema.py, and a branch in
AugmentationEngine.apply_to_document/apply_to_signature.image/; annotate manually or let
auto-detection + caching handle it.dataset/
(optionally with a metadata.csv); no code changes needed.difficulty.* weights/
thresholds in a config (must still sum to 1.0 — enforced by the schema).