Downloads · 30 days
0
Ma7ee7/Bad-Apple_Unified-Model-Test
Bad-Apple_Unified-Model-Test is a machine learning model from Ma7ee7. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch.
This is a compact coordinate-based neural representation of the complete Bad Apple!! shadow video and its stereo audio.
Downloads · 30 days
0
Access
Public
Updated Jul 26, 2026
Repo size
36.3 MB
Likes
0
Public
Click a slice to open those files.
.mp427.4 MB · 75%
From the Hugging Face model README
This is a compact coordinate-based neural representation of the complete Bad Apple!! shadow video and its stereo audio.
The checkpoint does not store ordinary video frames or compressed audio. Given a normalized time and pixel coordinate, the network predicts the image brightness. Given a normalized time coordinate with audio conditioning, the same network predicts the stereo waveform.
Download or open the generated video
Install the Python requirements and make sure ffmpeg is available on your
system. Then render the checkpoint:
python bad_apple_nn.py render model.pt --audio-source generated
The generated MP4 is written to outputs/.
This is one unified multimodal model, not two independently trained models.
Video and audio share:
The shared trunk ends in two small task-specific output heads:
Separate output heads are necessary because pixels and audio samples have different output shapes, but the representation and main network are shared and trained together in one checkpoint.
| Property | Value |
|---|---|
| Architecture | Unified coordinate neural field (unified-v3) |
| Trainable parameters | 1,111,571 |
| Training video resolution | 192 x 144 |
| Rendered resolution at default 3x scale | 576 x 432 |
| Frames | 6,572 |
| Frame rate | 30 FPS |
| Duration | About 3 minutes 39 seconds |
| Generated audio | 16 kHz stereo |
| Best checkpoint step | 59,000 / 80,000 |
| Video pixel accuracy | 99.6% |
| Silhouette IoU | 0.991 |
The inference weights occupy approximately 4.45 MB in FP32. The uploaded training checkpoint may be larger because it contains both regular and exponential-moving-average weights. Rendering uses the EMA weights by default.
| File | Purpose |
|---|---|
model.pt | Unified video-and-audio checkpoint |
bad_apple_nn.py | Model definition and renderer |
requirements.txt | Python dependencies |
demo.mp4 | Example generated result |
ffmpeg is required to encode and combine the rendered streams.