Downloads · 30 days
0
RichardScottOZ/comics-analysis-closure-lite-simple
comics-analysis-closure-lite-simple is a machine learning model from RichardScottOZ. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
The model code and documentation repository is at https://github.com/RichardScottOZ/comic-analysis
Downloads · 30 days
0
Access
Public
Updated Mar 1, 2026
Repo size
2.6 GB
Likes
1
Public
Click a slice to open those files.
.pth2.6 GB · 100%
From the Hugging Face model README
The model code and documentation repository is at https://github.com/RichardScottOZ/comic-analysis
Using transformers multimodal fusion of image and text to make embeddings to query comics for similarity or text.
More more detail the repo above.
language: en tags:
ClosureLiteSimple is the Version 1 precursor to the Stage 3 panel encoder within the Comic Analysis Framework.
It is a multimodal neural network designed to fuse image crops, textual dialogue, and compositional metadata into a unified 384-dimensional embedding per comic panel, and can also aggregate these panels into a single Page-level embedding using an attention mechanism.
(Note: This model is considered deprecated in favor of the newer comic-panel-encoder-v1 which utilizes SigLIP, ResNet, and an improved Adaptive Fusion Gate).
The ClosureLiteSimple model consists of the PanelAtomizerLite and a SimpleAttention mechanism:
google/vit-base-patch16-224):
roberta-base):
GatedFusion):
SimpleAttention):
The codebase for this model resides in the src/version1/ directory of the repository.
import torch
from PIL import Image
import torchvision.transforms as T
from transformers import AutoTokenizer
# Requires cloning the GitHub repo
from closure_lite_simple_framework import ClosureLiteSimple
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
# 1. Initialize Model
model = ClosureLiteSimple(d=384, num_heads=4, temperature=0.1).to(device)
# Load weights from Hugging Face
state_dict = torch.hub.load_state_dict_from_url(
"https://huggingface.co/RichardScottOZ/closure-lite-simple/resolve/main/best_model.pt",
map_location=device
)
if 'model_state_dict' in state_dict:
state_dict = state_dict['model_state_dict']
model.load_state_dict(state_dict)
model.eval()
# 2. Prepare Inputs (Example: A page with 2 panels)
transform = T.Compose([
T.Resize((224, 224)),
T.ToTensor(),
T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
])
# Dummy Image Crops (B=1 page, N=2 panels, C=3, H=224, W=224)
images = torch.stack([
transform(Image.new('RGB', (224, 224))),
transform(Image.new('RGB', (224, 224)))
]).unsqueeze(0).to(device)
# Dummy Text
tokenizer = AutoTokenizer.from_pretrained("roberta-base")
text_enc = tokenizer(["Panel 1 text", "Panel 2 text"], return_tensors='pt', padding=True)
input_ids = text_enc['input_ids'].unsqueeze(0).to(device)
attention_mask = text_enc['attention_mask'].unsqueeze(0).to(device)
# Dummy Composition (B=1, N=2, F=7)
comp_feats = torch.zeros((1, 2, 7)).to(device)
# Valid Panel Mask (B=1, N=2)
panel_mask = torch.tensor([[True, True]]).to(device)
# 3. Generate Embeddings
with torch.no_grad():
panel_embeddings, page_embedding = model(
images, input_ids, attention_mask, comp_feats, panel_mask
)
print(f"Panel Embeddings Shape: {panel_embeddings.shape}") # (1, 2, 384)
print(f"Page Embedding Shape: {page_embedding.shape}") # (1, 384)
GatedFusion mechanism struggled to fall back gracefully to the visual features, often resulting in collapsed or non-discriminative embeddings for single-modality queries.comic-panel-encoder-v1), which utilizes independent modality projection and a masked Adaptive Fusion gate to solve the dominance issues.Please reference the Comic Analysis GitHub Repository when utilizing this architecture.