Downloads · 30 days
0
fjmgAI/whisper-large-v3-ATC
whisper-large-v3-ATC is a automatic speech recognition model from fjmgAI. Use it when you need speech turned into text. It is set up for unsloth. The card lists the license as apache-2.0.
<img src="https://cdn-avatars.huggingface.co/v1/production/uploads/67b2f4e49edebc815a3a4739/R1g957j1aBbx8lhZbWmxw.jpeg" width="200"/
Downloads · 30 days
0
Access
Public
Updated Jun 13, 2025
Repo size
126 MB
Likes
0
Public
Click a slice to open those files.
.safetensors126 MB · 99%
From the Hugging Face model README
fjmgAI/whisper-large-v3-ATC
unsloth/whisper-large-v3
Fine-tuning was performed using unsloth, an efficient fine-tuning framework optimized for low-resource environments.
This dataset contains 14,830 examples transcriptions and corresponding audio files from two main sources: ATCO2 and the UWB-ATCC corpus, specifically selected for aviation-related communications.
First install the dependencies:
Colab Version
%%capture
!pip install --no-deps bitsandbytes accelerate xformers==0.0.29.post3 peft trl==0.15.2 triton cut_cross_entropy unsloth_zoo
!pip install sentencepiece protobuf "datasets>=3.4.1" huggingface_hub hf_transfer
!pip install transformers==4.51.3
!pip install --no-deps unsloth
!pip install librosa soundfile evaluate jiwer
No Colab Version
pip install unsloth
pip install librosa soundfile evaluate jiwer
Then you can load this model and run inference.
import torch
from unsloth import FastModel
from transformers import pipeline
from transformers import WhisperForConditionalGeneration
model, tokenizer = FastModel.from_pretrained(
model_name = "fjmgAI/whisper-large-v3-ATC",
dtype = None,
load_in_4bit = False,
auto_model = WhisperForConditionalGeneration,
whisper_language = "English",
whisper_task = "transcribe",
)
model.generation_config.language = "<|en|>"
model.generation_config.task = "transcribe"
model.config.suppress_tokens = []
model.generation_config.forced_decoder_ids = None
whisper = pipeline(
"automatic-speech-recognition",
model=model,
tokenizer=tokenizer.tokenizer,
feature_extractor=tokenizer.feature_extractor,
processor=tokenizer,
return_language=True,
torch_dtype=torch.float16
)
audio_file = "audio_example.flac"
transcribed_text = whisper(audio_file)
print(transcribed_text["text"])
This fine-tuned model is designed for Speech-to-Text (STT) applications in Air Traffic Control (ATC) environments, leveraging a specialized ATC dataset to enhance robustness and precision in transcribing ATC recordings. The model aims to deliver accurate and reliable transcription while maintaining efficient performance.