Downloads · 30 days
16
35% of all-time downloads
itspublu/EgoSieve-S
EgoSieve-S is a video classification model from itspublu. Use it for the video classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
EgoSieve-S ranks manipulation-ready spans in first-person video. It produces three readiness logits (KEEP, REVIEW, REJECT), seven observable issue scores, diagnostic start/end boundary proposals, and a normalized retr…
Downloads · 30 days
16
35% of all-time downloads
All-time downloads
46
Public
Parameters
23.9M
95.7 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors95.7 MB · 99%
From the Hugging Face model README
EgoSieve-S ranks manipulation-ready spans in first-person video. It produces
three readiness logits (KEEP, REVIEW, REJECT), seven observable issue
scores, diagnostic start/end boundary proposals, and a normalized retrieval
embedding. It is a dataset-curation model, not a robot policy.
from transformers import AutoModelForVideoClassification, AutoProcessor
processor = AutoProcessor.from_pretrained(
"itspublu/EgoSieve-S", revision="v0.1.0", trust_remote_code=True
)
model = AutoModelForVideoClassification.from_pretrained(
"itspublu/EgoSieve-S", revision="v0.1.0", trust_remote_code=True
).eval()
outputs = model(**processor(frames, return_tensors="pt"))
The timestamp-aware scanner and JSONL compiler are provided by the egosieve
package. The checkpoint expects 12 center-sampled
RGB frames per window; use its bundled processor.
Data represented in the held-out evaluation: HoloAssist (CDLA-Permissive-2.0), HoloAssist controlled corruptions (CDLA-Permissive-2.0). Splits are grouped by original capture unit. Readiness, calibration, and boundary results use 142 human-grounded readiness rows: 0 direct human and 142 human-derived. Boundary results use 97 human-grounded boundary rows. Issue results use 145 labeled rows: 0 human, 109 human-derived, and 36 programmatic controlled corruptions. Unlabeled task targets are masked. The release bundle includes raw held-out predictions with task-level provenance, split assignments, exact run configuration, and metric provenance.
Human-derived rows are not direct, independent EgoSieve rubric judgments. Treat them as proxy evidence and consult the source dataset cards and their dataset-specific proxy details before comparing or interpreting these metrics.
For v0.1, readiness and boundaries come from a fixed-grid occupancy rule over
reviewed HoloAssist fine-action intervals. low_hand_activity is an occupancy
proxy, and acting_hand_not_visible follows HoloAssist's acting-hand modifier;
it does not assert that every hand is absent. The other five issue metrics
measure injected-corruption versus unmodified-reference discrimination. Those
references were not independently audited as natural issue negatives.
Use the model to rank raw egocentric windows, route uncertain spans for review, and create embeddings for near-duplicate search. Readiness remains dependent on the published rubric and capture domain. RGB cannot establish force, physical success, consent, safety, metric depth, or legal publishability. Boundary scores are proposals and are diagnostic-only in the v0.1 compiler.
First-person recordings can contain faces, screens, homes, and bystanders. Apply a separate privacy and consent review before sharing any media.