Downloads ยท 30 days
0
chijiokejackson35/jack-lora
jack-lora is a machine learning model from chijiokejackson35. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Autonomous 4K Character Photography & Cinema Generation Suite Subject: jackthelastcreator Author & Architect: Antigravity AI & Jack Last Updated: September 2026 Repository: /Users/gameboy/Documents/Dev Apps/Comfy + Moโฆ
Downloads ยท 30 days
0
Access
Public
Updated Sep 28, 2026
Repo size
18.8 GB
Likes
0
Public
Click a slice to open those files.
.safetensors647 MB ยท 100%
From the Hugging Face model README
Autonomous 4K Character Photography & Cinema Generation Suite
Subject:jackthelastcreator
Author & Architect: Antigravity AI & Jack
Last Updated: September 2026
Repository:/Users/gameboy/Documents/Dev Apps/Comfy + Modal
Hugging Face:chijiokejackson35/jack-lora
This project is an end-to-end, serverless generative AI studio designed to produce 100% photorealistic 4K imagery and 97-frame cinematic videos of Jack (jackthelastcreator). The system accepts a Pinterest link or text prompt, reverse-engineers the camera optics, lighting, and composition using Google Gemini 3.8 Flash, generates high-fidelity latent representations using Tongyi Z-Image Turbo conditioned on a custom LoRA, upscales the result with SeedVR2 VideoUpscaler, and animates still frames into 4K video using Lightricks LTX-Video 2.3 22B.
flowchart TD
A["Pinterest URL / Prompt"] --> B["Pinterest Scraper (web_ui/pinterest_gemini_engine.py)"]
B --> C["Google Gemini 3.8 Flash Vision"]
C -->|Optical Prompt + Negative + Aspect Ratio| D["Headless ComfyUI on Modal A100"]
subgraph ComfyUI_Engine ["ComfyUI 4K Pipeline"]
E["Z-Image Turbo BF16 (8 Steps, FlowMatch)"] --> G["Custom LoRA (comfy_jack_zimage_v1_recipe)"]
F["Qwen 2.5 VL Text Encoder (FP8)"] --> G
G --> H["Base KSampler (Euler Simple, Denoise 1.0)"]
H --> I["VAE Decode (z_image_ae.safetensors)"]
I --> J["SeedVR2 VideoUpscaler (DiT 3B FP8, input_noise_scale 0.18)"]
end
D --> ComfyUI_Engine
J --> K["4K Master Still (PNG)"]
K --> L["Motion Gemini Engine (web_ui/motion_gemini_engine.py)"]
L --> M["LTX-Video 2.3 22B Distilled (Modal A100-80GB)"]
M --> N["97-Frame 4K MP4 Video (24fps)"]
The initial model (comfy_jack_lora_v1.safetensors) produced incredible facial likeness. When we audited its original configuration (ai-toolkit/config/jack_zimage_turbo.json), we uncovered the exact parameters responsible:
linear: 32, linear_alpha: 32).conv: null, no convolution layers).lokr_full_rank: false).[512, 768, 1024].Subsequent attempts to expand the dataset to multi-angle shots degraded likeness:
jack_zimage_turbo_3000 / jack_zimage_master_lora):
conv: 16, enabled lokr_full_rank: true, locked resolution to square [1024], trained to 3,000 steps.conv: 16 and lokr_full_rank: true enabled.| Parameter | V1 Recipe (Success) | 3,000-Step Run (Failure) | Technical / Mathematical Reason |
|---|---|---|---|
linear vs conv | linear: 32, conv: 0 | linear: 32, conv: 16 | Convolutional LoRA layers inject rank updates into spatial feature maps. In diffusion architectures, this distorts the spatial cross-attention geometry of the face, stretching the skull and narrowing the nose bridge. Pure linear LoRA only tunes the attention projection matrices (Q, K, V). |
| Lokr Full Rank | false | true | Full rank matrix reconstruction over-parameterized the adapter, memorizing camera lens focal lengths and burning rigid spatial proportions into the weights. |
| Resolution Bucketing | [512, 768, 1024] | [1024] (Locked Square) | Locking training to square 1024 force-stretched vertical phone portraits and 16:9 landscape b-roll, teaching the model unnatural face proportions. Bucketing preserves true aspect ratios. |
| Dataset Dilution | 35 portraits (100% face) | 37 wide body/desk shots + 5 faces | When 88% of images had small or obscured faces, the trigger token jackthelastcreator learned computer monitors, office chairs, and desks rather than facial morphology. |
| Step Count & Epochs | 1,500 steps (32 epochs) | 3,000 steps (64 epochs) | Z-Image Turbo is an 8-step distilled flowmatch model. 64 epochs at lr: 1e-4 causes latent saturation (waxy plastic specular shine and blown highlights). |
jack_zimage_v1_recipe)We combined the pristine facial fidelity of V1 with multi-angle capability:
jack_zimage_v1_recipe.yamlmodal_train_v1_recipe.pylinear: 32, linear_alpha: 32, conv: 0.adamw8bit, lr: 0.0001, weight_decay: 0.0001.flowmatch, loss_type: mse.comfy_jack_zimage_v1_recipe.safetensors (162.2 MB).Located at lora_training/master_dataset/, containing exactly 47 curated images:
DSC_2667 to DSC_2723). Pristine lighting, sharp eye focus, varying facial angles.DSC_2691_crop.png, king_of_kings_crop.png.broll_front45_01.png to 04.png.broll_overhead_01.png to 04.png.Every image is paired with a .txt file containing the trigger word jackthelastcreator. Captions explicitly describe the camera angle and environment so the model decouples facial identity from the scene:
"A sharp eye-level portrait of jackthelastcreator, young Black man with clean fade haircut and short trimmed chin beard, neutral grey background.""A high-angle bird's-eye overhead shot of jackthelastcreator seated at a dark desk typing on a laptop, hands on keyboard, modern workspace."Z-Image Turbo is distilled down to 8 sampling steps. While fast, 8-step flowmatch schedulers lack the Brownian motion variance of traditional 30-step Euler schedulers, tending to produce overly smoothed, plastic skin.
To eliminate the waxy look, we engineered a two-stage pipeline:
1024x1024 or 768x1152).jack_zimage_seedvr2_api.jsonseedvr2_ema_3b_fp8_e4m3fn.safetensors + ema_vae_fp16.safetensors.input_noise_scale: 0.18: Setting this parameter to 0.18 (up from 0.06) forces the 3B diffusion transformer to inject high-frequency photographic noise, restoring natural skin pores, beard stubble, cloth weave, and 35mm film grain.1.0 for frontal/side portraits, and tuned to 0.85 for extreme overhead bird's-eye shots to prevent facial weights from distorting downward angles."mutated hands, extra limbs, floating arms, third arm, extra fingers, malformed limbs, disembodied hands, extra keyboard".Located at web_ui/pinterest_gemini_engine.py:
pin.it/...) and Pinterest pin pages, bypassing redirects and extracting the uncompressed source image (i.pinimg.com/736x/...).jackthelastcreator (young Black man, clean fade, short trimmed goatee).Kodak Portra 400, 35mm DSLR capture, visible epidermal pores, natural skin sheen) while banning generic AI buzzwords (photorealistic, 8k, hyperrealistic).Located at web_ui/motion_gemini_engine.py and modal_generate_jack_video.py:
All models and outputs are persisted across dedicated Modal Volumes on workspace the-john-secret:
jack-zimage-seedvr2-volume: Foundation diffusion models, VAEs, text encoders, ComfyUI custom nodes, and converted LoRAs.ostris-outputs: AI-Toolkit training checkpoints, database, and intermediate validation samples.flux-lora-models: LTX-Video base checkpoints and text encoders.[!IMPORTANT] Zero Idle Cost ($0.00): All Modal worker functions spin down to 0 replicas immediately after execution. The deployed web studio (
jack-studio-web) is serverless ASGI; it consumes zero GPU credits while waiting for user requests.
Our execution scripts (modal_headless_zimage_seedvr2.py and modal_cloud_studio.py) run ComfyUI headlessly:
127.0.0.1:8188./cache directly into ComfyUI's internal folder hierarchy.POST /prompt.GET /history for task completion.os.killpg).https://the-john-secret--jack-studio-web-web-endpoint.modal.runjack-studio-next/. Built with React 19, Tailwind CSS, TypeScript, and Lucide icons for mobile browser testing.If you ever need to retrain or redeploy from scratch, execute these exact steps:
# Launch detached training on cloud GPU
.venv311/bin/modal run --detach modal_train_v1_recipe.py
ostris-outputs every 250 steps./cache/models/loras/comfy_jack_zimage_v1_recipe.safetensors.# Test Frontal Portrait
.venv311/bin/modal run modal_headless_zimage_seedvr2.py \
--prompt "A sharp cinematic portrait of jackthelastcreator in a modern studio, warm directional key lighting, visible skin pores" \
--output "rendered_clips/test_portrait.png"
# Test True 90-Degree Bird's-Eye Desk View
.venv311/bin/modal run modal_headless_zimage_seedvr2.py \
--prompt "A true 90-degree direct overhead bird's-eye view looking straight down from the ceiling at a dark wooden desk at night. Crown of head of jackthelastcreator with clean buzzcut fade at bottom edge, arms typing on laptop, warm desk lamp, notebook, raw 35mm film photograph" \
--negative-prompt "mutated hands, extra limbs, floating arms, third arm, extra fingers, 3d render, plastic skin" \
--width 768 --height 1152 \
--output "rendered_clips/test_overhead.png"
# Deploy permanent web app to Modal
.venv311/bin/modal deploy modal_cloud_studio.py
https://the-john-secret--jack-studio-web-web-endpoint.modal.run| File / Resource | Location / URI | Description |
|---|---|---|
| Production LoRA | chijiokejackson35/jack-lora / /cache/models/loras/ | comfy_jack_zimage_v1_recipe.safetensors (162.2 MB) |
| Original V1 LoRA | chijiokejackson35/jack-lora / /cache/models/loras/ | comfy_jack_lora_v1.safetensors (162.2 MB) |
| Master Training Config | jack_zimage_v1_recipe.yaml | Full training recipe specification |
| Modal Training Runner | modal_train_v1_recipe.py | Cloud training daemon with Telegram alerts |
| ComfyUI 4K Workflow | jack_zimage_seedvr2_api.json | Z-Image Turbo + SeedVR2 API execution graph |
| Headless 4K Runner | modal_headless_zimage_seedvr2.py | Python CLI for running headless renders |
| Pinterest Gemini Engine | web_ui/pinterest_gemini_engine.py | Pin scraper & Gemini 3.8 Flash optical prompt extractor |
| Motion Gemini Engine | web_ui/motion_gemini_engine.py | Motion dynamics analysis for video animation |
| Cloud Studio Web App | modal_cloud_studio.py | Serverless FastAPI web server on Modal |
| Next.js Studio App | jack-studio-next/ | Mobile-first Next.js React frontend |
| Master Training Dataset | lora_training/master_dataset/ | 47 curated images + 47 matching .txt captions |
| Telegram Alerts Bot | Bot: 8934320817:..., Chat: 6291627175 | Automatic completion & error dispatch |