Downloads · 30 days
15
14% of all-time downloads
bigdefence/Bigvox-Midm-Audio
Bigvox-Midm-Audio is a audio-text-to-text model from bigdefence. Use it for the audio-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
- Bigvox은 한국어 음성 인식에 특화된 고성능, 저지연 음성 언어 멀티모달 모델입니다. K-intelligence/Midm-2.0-Mini-Instruct 기반으로 구축되었습니다. 🚀 - End-to-End 음성 멀티모달 구조를 채택하여 음성 입력부터 텍스트 출력까지 하나의 파이프라인에서 처리하며, 추가적인 중간 모델 없이 자연스럽게 멀티모달 처리를 지원합니다.
Downloads · 30 days
15
14% of all-time downloads
All-time downloads
105
Public
Parameters
3B
6.9 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors5.9 GB · 100%
From the Hugging Face model README

| 항목 | 세부사항 |
|---|---|
| 기반 모델 | K-intelligence/Midm-2.0-Mini-Instruct |
| 언어 | 한국어 (Korean) |
| 모델 크기 | ~2B 파라미터 |
| 작업 유형 | Speech-to-Text 음성 멀티모달 |
| 라이선스 | Apache 2.0 |
Bigvox을 시작하려면 다음과 같이 레포지토리를 클론하고 환경을 설정하세요. 🛠️
레포지토리 클론:
git clone https://github.com/bigdefence/bigvox-midm
cd bigvox-midm
의존성 설치:
bash setting.sh
Huggingface CLI 사용:
pip install -U huggingface_hub
huggingface-cli download bigdefence/Bigvox-Midm-Audio --local-dir ./checkpoints
Snapshot Download 사용:
pip install -U huggingface_hub
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="bigdefence/Bigvox-Midm-Audio",
local_dir="./checkpoints",
resume_download=True
)
Git 사용:
git lfs install
git clone https://huggingface.co/bigdefence/Bigvox-Midm-Audio
Bigvox으로 추론을 수행하려면 다음 단계를 따라 모델을 설정하고 로컬에서 실행하세요. 📡
모델 준비:
./models/speech_encoder/ 디렉토리에 배치 🎤추론 실행:
python3 omni_speech/infer/bigvox.py --query_audio test_audio.wav
python3 omni_speech/infer/bigvox_streaming.py --query_audio test_audio.wav
이 모델은 Apache 2.0 라이선스 하에 배포됩니다. 상업적 사용이 가능하며, 자세한 내용은 LICENSE 파일을 참조하세요.
BigDefence와 함께 한국어 AI 음성 인식의 미래를 만들어가세요! 🚀🇰🇷
"Every voice matters, every word counts - 모든 목소리가 중요하고, 모든 말이 가치 있습니다"