Downloads · 30 days
6
2% of all-time downloads
youzarsif/wav2vec2bert_2_diffusion
wav2vec2bert_2_diffusion is a feature extraction model from youzarsif. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as apache-2.0.
This is a Wav2Vec2-BERT model that can be used as an audio conditioning mechanism for Stable Diffusion instead of the CLIP text encoder.
Downloads · 30 days
6
2% of all-time downloads
All-time downloads
297
Public
Parameters
631M
2.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.5 GB · 100%
From the Hugging Face model README
This is a Wav2Vec2-BERT model that can be used as an audio conditioning mechanism for Stable Diffusion instead of the CLIP text encoder.
<video controls src="https://cdn-uploads.huggingface.co/production/uploads/6674757dc1ccf20bff7089cf/6JL3EDXBcP-h26PX4T0zs.webm"></video> <img src="https://cdn-uploads.huggingface.co/production/uploads/6674757dc1ccf20bff7089cf/NmlyRozQileSdXUgpB_kW.png" width="400" height="300"> <video controls src="https://cdn-uploads.huggingface.co/production/uploads/6674757dc1ccf20bff7089cf/9ubrZlD1XD7xZrLgLAMh6.webm"></video> <img src="https://cdn-uploads.huggingface.co/production/uploads/6674757dc1ccf20bff7089cf/0QX906Q09xZYg5mJKFFwR.png" width="400" height="300">
<video controls src="https://cdn-uploads.huggingface.co/production/uploads/6674757dc1ccf20bff7089cf/QWP4ofLxT5zPf6p67Yh8H.webm"></video> <img src="https://cdn-uploads.huggingface.co/production/uploads/6674757dc1ccf20bff7089cf/8DJb-Cp650mIeb0VwJrmO.png" width="400" height="300">
The audio2img project aims to enhance generative modeling by integrating audio embeddings into the conditioning process of models like Stable Diffusion. This integration allows for the exploration of new creative possibilities by leveraging the rich semantic information contained in audio data.
Potential Users:
Researchers and Developers
Artists and Creatives
Content Creators
https://huggingface.co/datasets/nateraw/fsd50k
The core idea behind our training process is to achieve cross-modal alignment between audio and text embeddings using a two-stream architecture. This involves leveraging the powerful CLIPTextModel to generate text embeddings that serve as true labels for the audio embeddings produced by our Wav2Vec2Bert model.