Downloads · 30 days
8
8% of all-time downloads
malper/abjadsr-he
abjadsr-he is a machine learning model from malper. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Fine-tuned from outputs/pretrain on ILSpeech (~2h Hebrew, studio quality). Given audio, outputs Hebrew text paired with IPA transcription.
Downloads · 30 days
8
8% of all-time downloads
All-time downloads
100
Public
Parameters
809M
9.7 GB on disk
Likes
1
Public
Click a slice to open those files.
.pt6.5 GB · 67%
From the Hugging Face model README
Fine-tuned from outputs/pretrain on ILSpeech (~2h Hebrew, studio quality). Given audio, outputs Hebrew text paired with IPA transcription.
Stage 2 of 2 — use this model for inference.
Hebrew text paired with ASCII IPA transcription:
החליט יוז = hexl'it j'uz
IPA special characters are mapped to ASCII: ʃ→S, ʒ→Z, dʒ→dZ, tʃ→tS, ʔ→q, ˈ→', ʁ→r, χ→x, ɡ→g.
import torch
import soundfile as sf
from transformers import WhisperForConditionalGeneration, WhisperProcessor
model_id = "malper/abjadsr-he"
processor = WhisperProcessor.from_pretrained(model_id)
model = WhisperForConditionalGeneration.from_pretrained(model_id)
model.eval()
# Load audio (must be 16 kHz mono float32)
audio, sr = sf.read("audio.wav", dtype="float32", always_2d=False)
# resample if needed: torchaudio.functional.resample(torch.from_numpy(audio), sr, 16000).numpy()
inputs = processor(audio, sampling_rate=16000, return_tensors="pt")
forced_ids = processor.get_decoder_prompt_ids(language="he", task="transcribe")
with torch.no_grad():
generated = model.generate(
inputs.input_features,
forced_decoder_ids=forced_ids,
max_new_tokens=444,
)
output = processor.batch_decode(generated, skip_special_tokens=True)[0].strip()
print(output)
# e.g. "hex'lit j'uzem lena'tsel"