Downloads · 30 days
50
8% of all-time downloads
phenobase/phenovisionL
phenovisionL is a image classification model from phenobase. Use it when you need a label for an image. It is set up for transformers. The card lists the license as mit.
PhenoVisionL is a Vision Transformer (ViT-Large) model fine-tuned to detect leaf phenological states in plant photographs: green leaves, colored (senescent) leaves, and breaking leaf buds. It was trained on 165,988 iN…
Downloads · 30 days
50
8% of all-time downloads
All-time downloads
611
Public
Parameters
303M
1.2 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors1.2 GB · 100%
From the Hugging Face model README
PhenoVisionL is a Vision Transformer (ViT-Large) model fine-tuned to detect leaf phenological states in plant photographs: green leaves, colored (senescent) leaves, and breaking leaf buds. It was trained on 165,988 iNaturalist records of deciduous woody plants using a two-stage semi-supervised approach, and has generated 5.6 million leaf phenology observations across 6,500+ species, filling major geographic gaps in global leaf phenology data.
| Green Leaves | Colored Leaves | Breaking Buds | |
|---|---|---|---|
| Expert validation accuracy | 98.6% | 99.4% | 87.0% |
| False positive rate | 1.2% | 0.6% | 9.4% |
PhenoVisionL is initialized from the trained PhenoVision reproductive structures model rather than from ImageNet or PlantCLEF directly. This leverages the reproductive model's learned representations of plant structure and morphology, providing a strong initialization for leaf phenology tasks. A new randomly initialized classification head replaces the original 2-class output with a 3-class output.
Primary use: Detecting leaf phenological states in field photographs of deciduous woody plants.
Suitable for:
Out of scope:
Requires transformers, torch, and torchvision (torchvision is needed by the
image processor in transformers v5+).
from transformers import AutoModelForImageClassification, AutoImageProcessor
from PIL import Image
import torch
# Load model and processor
processor = AutoImageProcessor.from_pretrained("phenobase/phenovisionL")
model = AutoModelForImageClassification.from_pretrained("phenobase/phenovisionL")
model.eval()
# Run inference
image = Image.open("plant_photo.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
probs = torch.sigmoid(outputs.logits)[0]
# Output order is [green, colored, breaking_buds].
green_prob = probs[0].item() # index 0 = green leaves
colored_prob = probs[1].item() # index 1 = colored leaves
breaking_buds_prob = probs[2].item() # index 2 = breaking leaf buds
print(f"Green leaves: {green_prob:.3f}")
print(f"Colored leaves: {colored_prob:.3f}")
print(f"Breaking buds: {breaking_buds_prob:.3f}")
Raw probabilities should be converted to detection calls using the optimized thresholds and uncertainty buffers provided as companion files. Predictions falling within the buffer zone are classified as "Equivocal" and should be excluded for research-quality outputs.
See the companion file epoch_1_threshold_buffers.csv for the specific threshold and buffer values for each class.
PhenoVisionL uses a two-stage semi-supervised training approach:
Independent expert review of high-confidence (unequivocal) model predictions:
| Phenophase | Accuracy | False Positive Rate |
|---|---|---|
| Green leaves | 98.6% | 1.2% |
| Colored leaves | 99.4% | 0.6% |
| Breaking leaf buds | 87.0% | 9.4% |
Breaking buds have lower accuracy due to inherent task difficulty — morphological similarity between breaking leaf buds and flower buds, and limited expert-only training data for this class.
The following files are uploaded alongside the model weights:
| File | Description |
|---|---|
epoch_1_threshold_buffers.csv | Decision thresholds and uncertainty buffer parameters per class. Used to convert probabilities to Detected/Not Detected/Equivocal calls. Note: Despite the .csv extension, this file is in RDS format and should be read with readRDS() in R. |
family_stats.csv | Per-family (57 families) accuracy statistics for each leaf class. |
family_stats.csv for family-level accuracy.If you use PhenoVisionL in your research, please cite:
@article{grady2025phenovisionL,
title={PhenoVision: A framework for automating and delivering research-ready plant phenology data from field images},
author={Grady, Erin L. and Denny, Ellen G. and Seltzer, Carrie E. and Deck, John and Li, Daijiang and Dinnage, Russell and Guralnick, Robert P.},
journal={bioRxiv},
year={2025},
doi={10.1101/2025.09.26.678778}
}
Also cite the original PhenoVision framework paper:
@article{dinnage2025phenovision,
title={PhenoVision: A framework for automating and delivering research-ready plant phenology data from field images},
author={Dinnage, Russell and Grady, Erin and Neal, Nevyn and Deck, Jonn and Denny, Ellen and Walls, Ramona and Seltzer, Carrie and Guralnick, Robert and Li, Daijiang},
journal={Methods in Ecology and Evolution},
volume={16},
pages={1763--1780},
year={2025},
doi={10.1111/2041-210X.14346}
}