Downloads · 30 days
57
20% of all-time downloads
rzgar/Bernini-R-S2V
Bernini-R-S2V is a image-to-image model from rzgar. Use it when you need one image transformed into another. The card lists the license as apache-2.0.
Bernini S2V Conditioning v2 - Bernini in-context video/image conditioning with masked lip-sync for one or two speakers. <video controls width="100%" height="480" <source src="https://huggingface.co/rzgar/Bernini-R-S2V…
Downloads · 30 days
57
20% of all-time downloads
All-time downloads
288
Public
Repo size
143 GB
Likes
46
Trending 1
Click a slice to open those files.
.safetensors143 GB · 100%
From the Hugging Face model README
Bernini S2V Conditioning v2 - Bernini in-context video/image conditioning with masked lip-sync for one or two speakers. <video controls width="100%" height="480"> <source src="https://huggingface.co/rzgar/Bernini-R-S2V/resolve/main/video/ComfyUI__00003-audio.mp4" type="video/mp4"> Your browser does not support the video tag. </video>
Unzip ComfyUI-WanBerniniS2V_v2.zip into ComfyUI/custom_nodes/, then restart ComfyUI.
or save the Python files in: ComfyUI/custom_nodes/ComfyUI-WanBerniniS2V_v2/
Disable or remove the older ComfyUI-WanBerniniS2V folder if you only want v2.
audio_1 + mask_1audio_2 / mask_2 unwiredaudio_1 + mask_1 - first speakeraudio_2 + mask_2 - second speakerspeaker_2_start_frame = -1 - second audio starts when the first clip endsPaint on the output frame where each speaker's face appears. Masks control lip-sync placement only, they are not tied to reference_image_N slots.
<img src="https://huggingface.co/rzgar/Bernini-R-S2V/resolve/main/ComfyUI-WanBerniniS2V_v2/demo_assets/Screenshot_ComfyUI-WanBerniniS2V_v2.png" width="1280" height="720" />
Speech-driven video on Bernini-R , T2V, I2V, and V2V with lip-sync.
This model adds single-speaker audio support to Bernini-R, so you can drive video with speech in text-to-video, image-to-video, and video-to-video setups. It is not state-of-the-art audio-to-video, but it removes the need for post-processing or extra models just to add speech to Wan videos. For basic talking-head work, or longer videos built from short sequences, it is a handy all-in-one option on top of Bernini's motion and editing strengths.
ComfyUI detects these as WAN22_S2V / 'WanModel_S2V' (audio keys trigger S2V model type).
ComfyUI/models/diffusion_models/
ComfyUI/models/audio_encoders/
(Create a folder in 'ComfyUI/custom_nodes/' named 'ComfyUI-WanBerniniS2V', then save the Python files in that folder.)
| Setting | Recommendation |
|---|---|
| Channels | Mono, wav2vec2 downmixes stereo internally |
| Sample rate | 44.1 khz or 48 khz (resampled to 16 kHz) |
| Content | Clear speech, less background music = better sync |
| Length | in my testing, max 9 to 15 seconds |