Downloads · 30 days
0
youngPhilosopher/drywall-qa-clipseg
drywall-qa-clipseg is a image segmentation model from youngPhilosopher. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
<p align="center" <h1 align="center"Prompted Segmentation for Drywall QA</h1 <p align="center" Text-conditioned binary mask prediction for construction defect detection </p </p
Downloads · 30 days
0
Access
Public
Updated Apr 10, 2026
Repo size
615 MB
Likes
1
Public
Click a slice to open those files.
.pt603 MB · 98%
From the Hugging Face model README
Feed a construction photo and a text prompt. Get a binary segmentation mask back.
Two tasks — crack detection and drywall taping/joint detection — both driven by natural language at inference time. Change the prompt, change what gets segmented. No class heads, no retraining.
Input: image.jpg + "segment wall crack"
Output: image__segment_wall_crack.png (binary mask, {0, 255})
We fine-tune CLIPSeg (Luddecke & Ecker, CVPR 2022) — a text-conditioned segmentation model built on CLIP. The entire CLIP backbone (149.6M params) stays frozen. Only a lightweight 3-block transformer decoder with U-Net skip connections (1.13M params) is trained.

The model takes an RGB image and a text prompt. The CLIP vision encoder (ViT-B/16) and text encoder independently produce embeddings. The decoder fuses these via cross-attention and generates logits at 352x352, which are thresholded at 0.5 to produce binary masks.
<details> <summary><b>Why CLIPSeg over Grounded SAM, SEEM, X-Decoder?</b></summary> <br>| CLIPSeg | Grounded SAM | SEEM | X-Decoder | |
|---|---|---|---|---|
| Text-to-mask | Direct | Two-stage (text → bbox → mask) | Multi-modal | Yes |
| Small-data fine-tuning | Proven | Moderate | Difficult | Not ideal |
| Consumer GPU (Apple M4) | Yes | Decoder only | No | No |
| HuggingFace native | Yes | Yes | GitHub only | Limited |
CLIPSeg is the only architecture that gives direct text-to-mask conditioning without bounding box intermediates, fine-tunes reliably on small datasets, and runs on consumer hardware with mature HuggingFace support.
</details>| Parameter | Value |
|---|---|
| Base model | CIDAS/clipseg-rd64-refined |
| Trainable | 1,127,009 params (decoder only) |
| Frozen | 149,620,737 params (CLIP backbone) |
| Loss | BCEDiceLoss — 0.5 BCE + 0.5 Dice |
| Optimizer | AdamW (lr=1e-4, wd=1e-4) + CosineAnnealingLR |
| Early stopping | patience 7 on val mIoU |
| Device | Apple M4 (MPS backend) |
| Wall time | 97.2 min (18 epochs, best at epoch 11) |
Standard BCE alone fails on thin structures like cracks — the severe foreground/background imbalance means BCE happily predicts "all background" at low loss. Dice loss directly optimizes overlap, forcing the model to find crack pixels. The 50/50 blend gives gradient stability (BCE) and overlap-awareness (Dice).
</details>
Training converged at epoch 11 (val mIoU 0.1605). The remaining 7 epochs showed no improvement before early stopping triggered at epoch 18.
All hyperparameters: configs/train_config.yaml
Two datasets from Roboflow Universe, downloaded manually in COCO format:
| Dataset | Source | Images | Raw Annotation | Mask Strategy |
|---|---|---|---|---|
| Taping | drywall-join-detect | 1,186 | Bounding boxes only | Filled rectangles |
| Cracks | cracks-3ii36 | 5,369 | COCO polygons | Pixel-accurate binary masks via pycocotools |
Note: The cracks dataset had 0 generated Roboflow versions — the owner never created an exportable version, making API download impossible. The raw export was downloaded directly from the website.
pycocotools.mask. Some annotations had empty segmentation fields (edge case) — handled with try/except fallback to bounding box rendering.5 synonyms per class, randomly sampled each training iteration. This forces the decoder to learn semantic meaning from the text encoder rather than memorize exact strings:
| Class | Prompts |
|---|---|
| Cracks | "segment crack" · "segment wall crack" · "segment surface crack" · "segment drywall crack" · "segment fracture" |
| Taping | "segment taping area" · "segment joint tape" · "segment drywall seam" · "segment drywall joint" · "segment tape line" |

Stratified by class (taping vs cracks), seed 42:
| Train | Validation | Test |
|---|---|---|
| 4,588 (70%) | 982 (15%) | 985 (15%) |
Preprocessing code: src/data/preprocess.py · Dataset class: src/data/dataset.py
The model's strongest predictions reach IoU 0.78 on both cracks and taping:

| Class | mIoU | Dice | Samples |
|---|---|---|---|
| Taping | 0.1917 | 0.2780 | 179 |
| Cracks | 0.1639 | 0.2434 | 806 |
| Overall | 0.1689 | 0.2497 | 985 |
Taping outperforms cracks because filled-rectangle masks provide a stronger supervision signal (larger contiguous regions) compared to thin crack annotations where minor spatial offsets cause disproportionate IoU drops.
| Metric | Value |
|---|---|
| Avg inference time | 58.7 ms / image |
| Model size | 575.1 MB |
| Output format | PNG, single-channel {0, 255}, resized to original dimensions |
| Threshold | 0.5 (sigmoid → binary) |
The model's worst predictions (IoU near zero) reveal systematic failure patterns:

What's going wrong in these examples:
| # | Factor | Impact |
|---|---|---|
| 1 | Coarse taping annotations | Source dataset has bounding boxes, not pixel masks. Filled rectangles include background → model over-predicts. |
| 2 | Thin crack IoU sensitivity | A 1px crack shifted 2px = near-zero IoU despite visual similarity. Dominates aggregate. |
| 3 | 352x352 resolution ceiling | CLIPSeg's fixed input size discards fine detail from high-res construction photos. |
| 4 | Frozen backbone domain gap | CLIP was trained on internet images, not construction imagery. Cannot adapt feature extraction. |
| 5 | Small decoder (1.13M params) | Limited capacity to learn construction-specific visual patterns. |
| Limitation | Solution | Expected Impact |
|---|---|---|
| Coarse taping masks | Use SAM/SAM2 to generate pixel-accurate masks from bounding boxes before training | High — directly fixes the supervision signal |
| Frozen backbone | Unfreeze last 2–3 ViT blocks with 10x lower learning rate for domain adaptation | High — lets the model learn construction-specific features |
| 352x352 resolution | Switch to SAM2 with text-prompt conditioning or a higher-res architecture | High — preserves fine crack detail |
| Small decoder | Add decoder blocks or increase hidden dimension (monitor overfitting) | Medium — more capacity, but risk of overfitting on small data |
| Thin-crack metric sensitivity | Use boundary IoU or distance-tolerant evaluation instead of standard IoU | Low — doesn't improve the model, but gives fairer measurement |

| Path | Purpose |
|---|---|
configs/train_config.yaml | All hyperparameters in one file |
src/data/preprocess.py | Annotation inspection, mask rendering, stratified splits |
src/data/dataset.py | PyTorch Dataset + CLIPSegProcessor collation |
src/model/clipseg_wrapper.py | Model loading + backbone freezing |
src/model/losses.py | BCEDiceLoss implementation |
src/train.py | Training loop with early stopping + logging |
src/evaluate.py | Test metrics, mask generation, visual comparisons |
src/predict.py | Single-image CLI inference |
src/best_predictions.py | Per-sample IoU scoring, best/worst prediction figures |
reports/report.typ | Typst source → report.pdf |
Prerequisites: Python 3.11+, uv, Homebrew (macOS)
brew install graphviz plantuml typst d2
uv sync
Download both datasets from Roboflow Universe in COCO format → place under data/raw/:
data/raw/
├── taping/ # drywall-join-detect (COCO export)
│ ├── train/
│ └── valid/
└── cracks/ # cracks-3ii36 (COCO export)
└── train/
uv run python -m src.data.preprocess
uv run python -m src.train
uv run python -m src.evaluate
uv run python -m src.predict path/to/image.jpg "segment crack"
d2 reports/diagrams/pipeline.d2 reports/diagrams/pipeline.png
plantuml -tpng reports/diagrams/training.puml
uv run python reports/diagrams/architecture.py
typst compile reports/report.typ reports/report.pdf
configs/train_config.yaml.outputs/logs/.