Downloads Ā· 30 days
33
0% of all-time downloads
camenduru/potat1
potat1 is a text-to-video model from camenduru. Use it when you need video from a text prompt. It is set up for diffusers.
š£ Please follow me for new updates https://twitter.com/camenduru <br / š„ Please join our discord server https://discord.gg/k5BwmmvJJU
Downloads Ā· 30 days
33
0% of all-time downloads
All-time downloads
6.7K
Public
Repo size
12.2 GB
Likes
160
Public
Click a slice to open those files.
.bin4.4 GB Ā· 100%
From the Hugging Face model README
š£ Please follow me for new updates https://twitter.com/camenduru <br /> š„ Please join our discord server https://discord.gg/k5BwmmvJJU
First Open-Source 1024x576 Text To Video Model š„³
https://huggingface.co/vdo/potat1-5000/tree/main <br /> https://huggingface.co/vdo/potat1-10000/tree/main <br /> https://huggingface.co/vdo/potat1-10000-base-text-encoder/tree/main <br /> https://huggingface.co/vdo/potat1-15000/tree/main <br /> https://huggingface.co/vdo/potat1-20000/tree/main <br /> https://huggingface.co/vdo/potat1-25000/tree/main <br /> https://huggingface.co/vdo/potat1-30000/tree/main <br /> https://huggingface.co/vdo/potat1-35000/tree/main <br /> https://huggingface.co/vdo/potat1-40000/tree/main <br /> https://huggingface.co/vdo/potat1-45000/tree/main <br /> https://huggingface.co/vdo/potat1-50000/tree/main <br /> https://huggingface.co/vdo/potat1-50000-base-text-encoder/tree/main = https://huggingface.co/camenduru/potat1 (you are here) <br />
Prototype Model <br /> Trained with https://lambdalabs.com ⤠1xA100 (40GB) <br /> 2197 clips, 68388 tagged frames ( salesforce/blip2-opt-6.7b-coco ) <br /> train_steps: 10000 <br />
https://huggingface.co/camenduru/potat1_dataset/tree/main
https://github.com/Breakthrough/PySceneDetect <br /> https://github.com/ExponentialML/Video-BLIP2-Preprocessor <br /> https://github.com/ExponentialML/Text-To-Video-Finetuning <br /> https://github.com/camenduru/Text-To-Video-Finetuning-colab <br />
https://huggingface.co/damo-vilab/modelscope-damo-text-to-video-synthesis <br /> https://www.modelscope.cn/models/damo/text-to-video-synthesis <br />
Thanks to damo-vilab ⤠ExponentialML ⤠kabachuha ⤠@DiffusersLib ⤠@LambdaAPI ⤠@cerspense ⤠@CiaraRowles1 ⤠@p1atdev_art ⤠<br />
Thanks to Orellius ⤠(important bug report) <br />
Please try it š£ <br /> https://github.com/camenduru/text-to-video-synthesis-colab <br />
<video src="https://github-production-user-asset-6210df.s3.amazonaws.com/54370274/244223223-c5201c8a-2815-4533-9474-1e312c564f4e.mp4" data-canonical-src="https://github-production-user-asset-6210df.s3.amazonaws.com/54370274/244223223-c5201c8a-2815-4533-9474-1e312c564f4e.mp4" controls="controls" muted="muted" class="d-block rounded-bottom-2 border-top width-fit" style="max-height:640px; min-height: 200px"></video>
Potat 2ļøā£ is in the oven ⨠<br />