Downloads · 30 days
0
optimizerai/vocos
vocos is a machine learning model from optimizerai. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This modelcard aims to be a base template for new models. It has been generated using this raw template.
Downloads · 30 days
0
Access
Public
Updated Nov 21, 2024
Repo size
1.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors1.7 GB · 100%
From the Hugging Face model README
This modelcard aims to be a base template for new models. It has been generated using this raw template.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
pip install "huggingface-hub[cli]"
huggingface-cli login # Paste access token w/ read access to this repository.
# Tokens look like this: hf_*****
export TEMP_DIR=$(mktemp -d)
huggingface-cli download optimizerai/vocos --exclude "*.safetensors" --local-dir $TEMP_DIR
pip install "file://$TEMP_DIR"
Or for an automated approach:
pip install "huggingface-hub[cli]"
export HF_TOKEN=hf_******
export TEMP_DIR=$(mktemp -d)
huggingface-cli download optimizerai/vocos --exclude "*.safetensors" --local-dir $TEMP_DIR
pip install "file://$TEMP_DIR"
If you want to hardcode your token for some reason:
pip install "huggingface-hub[cli]"
export TEMP_DIR=$(mktemp -d)
huggingface-cli download optimizerai/vocos --exclude "*.safetensors" --local-dir $TEMP_DIR --token hf_*****
pip install "file://$TEMP_DIR"
import torch
from vocos import get_voco
mel_voco = get_voco("mel")
encodec_voco = get_voco("encodec")
dac_voco = get_voco("dac")
dac_vae_voco = get_voco("dacvae")
oobleck_voco = get_voco("oobleck")
audio = torch.randn(1, 44100, 2) # [batch, audio_length, audio_channels]
latents = oobleck_voco.encode(audio) # [batch, encoded_length, latent_dim]
recon = oobleck_voco.decode(latents) # [batch, recon_length, audio_channels]
Sampling rate: oobleck_voco.sampling_rate
Audio channels: oobleck_voco.channel
Length conversion:
import torch
from vocos import get_voco
oobleck_voco = get_voco("oobleck")
audio_length = 44100
encode_length = oobleck_voco.encode_length(audio_length)
recon_length = oobleck_voco.decode_length(encode_length)
audio = torch.randn(1, audio_length, oobleck_voco.channel)
latent = oobleck_voco.encode(audio)
recon = oobleck_voco.decode(latent)
assert encode_length == latent.shape[1]
assert recon_length == recon.shape[1]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]