Downloads · 30 days
15
9% of all-time downloads
videoloc/seamless-translation
seamless-translation is a machine learning model from videoloc. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as cc-by-nc-4.0.
This is a SeamlessTranslation model that processes audio and text inputs with translation awareness to predict Time To Edit (TTE) for subtitle segments. Given an audio segment and its corresponding subtitle text, the…
Downloads · 30 days
15
9% of all-time downloads
All-time downloads
165
Public
Repo size
9.7 GB
Likes
0
Public
Click a slice to open those files.
.bin4.9 GB · 100%
From the Hugging Face model README
This is a SeamlessTranslation model that processes audio and text inputs with translation awareness to predict Time To Edit (TTE) for subtitle segments. Given an audio segment and its corresponding subtitle text, the model predicts how much time (in seconds) would be required to edit/refine that subtitle segment, while taking into account whether the subtitle is a translation or original content.
The model extends the basic SeamlessM4T architecture with a translation feature that helps distinguish between original and translated subtitle content, improving TTE prediction accuracy across 5 languages: English, French, Spanish, Italian, and German with various translation pairs between them.
The model extends the basic SeamlessM4T architecture with translation awareness:
Audio Processing:
Text Processing:
Translation Feature Processing:
Feature Fusion:
Regression Head:
pip install transformers torch torchaudio huggingface_hub
from transformers import AutoModel, AutoConfig
from huggingface_hub import hf_hub_download
import torch
import numpy as np
import importlib.util
# Load model - custom architecture requires importing the model class
model_files = hf_hub_download(repo_id="videoloc/seamless-translation", filename="modeling_seamless_translation.py")
spec = importlib.util.spec_from_file_location("modeling_seamless_translation", model_files)
modeling_module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(modeling_module)
# Now load the model using the custom class
config = modeling_module.SeamlessTranslationConfig.from_pretrained("videoloc/seamless-translation")
model = modeling_module.HFSeamlessTranslation.from_pretrained("videoloc/seamless-translation")
# Load the data collator (included in this repo)
collator_file = hf_hub_download(repo_id="videoloc/seamless-translation", filename="data_collator.py")
spec = importlib.util.spec_from_file_location("data_collator", collator_file)
collator_module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(collator_module)
# Initialize data collator
data_collator = collator_module.DataCollatorSimpleSeamless(
processor="facebook/hf-seamless-m4t-medium",
max_audio_length_sec=8.0,
max_text_length=256
)
# Prepare your data with translation information
your_data = [
{
'raw_audio': np.random.randn(16000 * 5), # 5 seconds at 16kHz
'raw_text': "Your subtitle text here",
'is_translation': 1, # 1 for translated content, 0 for original
}
]
# Process and run inference
batch = data_collator(your_data)
model.eval()
with torch.no_grad():
outputs = model(**batch)
tte_prediction = outputs.logits.item()
print(f"Predicted Time To Edit (TTE): {tte_prediction:.2f} seconds")
Your input data should be a list of dictionaries with:
raw_audio: NumPy array of audio samples (16kHz sampling rate)raw_text: String of subtitle textis_translation: Binary flag (1 for translated, 0 for original content)labels: Target TTE values in seconds (optional, for training)Example:
data = [
{
'raw_audio': audio_samples, # shape: (num_samples,) at 16kHz
'raw_text': "Subtitle text content",
'is_translation': 1, # 1 = translated, 0 = original
'labels': 2.5 # optional TTE target value in seconds
}
]
The model was trained with the following specifications:
seamless-basicseamless-langpairs