Downloads · 30 days
0
aTrapDeer/ace-step15-endpoint
ace-step15-endpoint is a machine learning model from aTrapDeer. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Downloads · 30 days
0
Access
Public
Updated Feb 15, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.txt4.8 MB · 82%
From the Hugging Face model README
Train ACE-Step 1.5 LoRA adapters, deploy your own Hugging Face Space, and run production-style inference through a Dedicated Endpoint.
app.py, lora_ui.pylora_train.pyqwen_caption_app.py, qwen_audio_captioning.py, scripts/annotations/af3_chatgpt_pipeline.py, scripts/pipeline/, services/pipeline_api.pyreact-ui/handler.py, acestep/scripts/hf_clone.pyscripts/endpoint/, scripts/jobs/python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python app.py
Open http://localhost:7860.
Use this sequence when setting up from scratch.
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
.env from .env.example and fill secretsHF_TOKEN=hf_xxx
HF_AF3_ENDPOINT_URL=https://YOUR_AF3_ENDPOINT.endpoints.huggingface.cloud
OPENAI_API_KEY=sk-...
OPENAI_MODEL=gpt-5-mini
AF3_MODEL_ID=nvidia/audio-flamingo-3-hf
python scripts/hf_clone.py space --repo-id YOUR_USERNAME/YOUR_SPACE_NAME
python scripts/hf_clone.py af3-nvidia-endpoint --repo-id YOUR_USERNAME/YOUR_AF3_NVIDIA_ENDPOINT_REPO
custom.handler.py exists in the endpoint repo.HF_TOKEN, AF3_NV_DEFAULT_MODE=think).python scripts/pipeline/run_af3_chatgpt_pipeline.py \
--dataset-dir ./train-dataset \
--backend hf_endpoint \
--endpoint-url "$HF_AF3_ENDPOINT_URL" \
--openai-api-key "$OPENAI_API_KEY"
python scripts/pipeline/refine_dataset_json_with_openai.py \
--dataset-dir ./train-dataset \
--enable-web-search
This script keeps core fields needed by ACE-Step LoRA training and preserves rich analysis context in source.rich_details.
python app.py
Then in UI:
scripts/endpoint/..zip or adapter files),
then load it without retraining in this Space..env (never commit this file):HF_TOKEN=hf_xxx
HF_AF3_ENDPOINT_URL=https://bc3r76slij67lskb.us-east-1.aws.endpoints.huggingface.cloud
OPENAI_API_KEY=sk-...
OPENAI_MODEL=gpt-5-mini
AF3_MODEL_ID=nvidia/audio-flamingo-3-hf
python af3_gui_app.py
PowerShell alternative:
.\scripts\dev\run_af3_gui.ps1
This command builds the React UI and serves it from the FastAPI backend.
Open http://127.0.0.1:8008.
Use the two buttons near the top of this README to create target repos in your HF account, then run:
Set token once:
# Linux/macOS
export HF_TOKEN=hf_xxx
# Windows PowerShell
$env:HF_TOKEN="hf_xxx"
Clone your own Space:
python scripts/hf_clone.py space --repo-id YOUR_USERNAME/YOUR_SPACE_NAME
Clone your own Endpoint repo:
python scripts/hf_clone.py endpoint --repo-id YOUR_USERNAME/YOUR_ENDPOINT_REPO
Clone a Qwen2-Audio caption endpoint repo:
python scripts/hf_clone.py qwen-endpoint --repo-id YOUR_USERNAME/YOUR_QWEN_ENDPOINT_REPO
Clone an Audio Flamingo 3 caption endpoint repo:
python scripts/hf_clone.py af3-endpoint --repo-id YOUR_USERNAME/YOUR_AF3_ENDPOINT_REPO
When creating that endpoint, set task to custom so it loads the custom handler.py.
Clone an AF3 NVIDIA-stack endpoint repo (matches NVIDIA Space stack better):
python scripts/hf_clone.py af3-nvidia-endpoint --repo-id YOUR_USERNAME/YOUR_AF3_NVIDIA_ENDPOINT_REPO
Use this path when you want think/long quality behavior closer to NVIDIA's public demo.
Clone both in one run:
python scripts/hf_clone.py all \
--space-repo-id YOUR_USERNAME/YOUR_SPACE_NAME \
--endpoint-repo-id YOUR_USERNAME/YOUR_ENDPOINT_REPO
.
|- app.py
|- lora_ui.py
|- lora_train.py
|- qwen_caption_app.py
|- qwen_audio_captioning.py
|- af3_chatgpt_pipeline.py
|- af3_gui_app.py
|- handler.py
|- acestep/
|- scripts/
| |- hf_clone.py
| |- dev/
| | |- run_af3_gui.py
| | `- run_af3_gui.ps1
| |- annotations/
| | `- qwen_caption_dataset.py
| |- pipeline/
| | `- run_af3_chatgpt_pipeline.py
| |- endpoint/
| | |- generate_interactive.py
| | |- test.ps1
| | |- test.bat
| | |- test_rnb.bat
| | `- test_rnb_2min.bat
| `- jobs/
| `- submit_hf_lora_job.ps1
| `- submit_hf_qwen_caption_job.ps1
|- services/
| `- pipeline_api.py
|- react-ui/
|- utils/
| `- env_config.py
|- docs/
| |- deploy/
| `- guides/
|- summaries/
| `- findings.md
`- templates/hf-endpoint/
Supported audio:
.wav, .flac, .mp3, .ogg, .opus, .m4a, .aacOptional sidecar metadata per track:
song_001.wavsong_001.json{
"caption": "melodic emotional rnb pop with warm pads",
"lyrics": "[Verse]\\n...",
"bpm": 92,
"keyscale": "Am",
"timesignature": "4/4",
"vocal_language": "en",
"duration": 120
}
Run the dedicated annotation UI:
python qwen_caption_app.py
Batch caption from CLI:
python scripts/annotations/qwen_caption_dataset.py \
--dataset-dir ./dataset_inbox \
--backend local \
--model-id Qwen/Qwen2-Audio-7B-Instruct \
--output-dir ./qwen_annotations \
--copy-audio
This also writes .json sidecars next to source audio by default for direct ACE-Step LoRA training.
Then train LoRA on the exported dataset:
python lora_train.py --dataset-dir ./qwen_annotations/dataset --model-config acestep-v15-base
This stack runs:
CLI single track:
python scripts/pipeline/run_af3_chatgpt_pipeline.py \
--audio "./train-dataset/Andrew Spacey - Wonder (Prod Beat It AT).mp3" \
--backend hf_endpoint \
--endpoint-url "$HF_AF3_ENDPOINT_URL" \
--hf-token "$HF_TOKEN" \
--openai-api-key "$OPENAI_API_KEY" \
--artist-name "Andrew Spacey" \
--track-name "Wonder"
CLI dataset batch:
python scripts/pipeline/run_af3_chatgpt_pipeline.py \
--dataset-dir ./train-dataset \
--backend hf_endpoint \
--endpoint-url "$HF_AF3_ENDPOINT_URL" \
--openai-api-key "$OPENAI_API_KEY"
Refine already-generated JSON files in place:
python scripts/pipeline/refine_dataset_json_with_openai.py \
--dataset-dir ./train-dataset \
--enable-web-search
Write refined files to a separate folder:
python scripts/pipeline/refine_dataset_json_with_openai.py \
--dataset-dir ./train-dataset \
--recursive \
--enable-web-search \
--output-dir ./train-dataset-refined
Single-command GUI (recommended):
python af3_gui_app.py
Manual API + React UI:
uvicorn services.pipeline_api:app --host 0.0.0.0 --port 8008 --reload
cd react-ui
npm install
npm run dev
Open http://localhost:5173 (manual) or http://127.0.0.1:8008 (single-command).
python scripts/endpoint/generate_interactive.py
Or run scripted tests:
scripts/endpoint/test.ps1scripts/endpoint/test.batCurrent baseline analysis and improvement ideas are tracked in:
summaries/findings.mddocs/deploy/SPACE.mddocs/deploy/QWEN_SPACE.mddocs/deploy/ENDPOINT.mddocs/deploy/AF3_ENDPOINT.mddocs/deploy/AF3_NVIDIA_ENDPOINT.mddocs/guides/qwen2-audio-train.md, docs/guides/af3-chatgpt-pipeline.mdHF_TOKEN, HF_AF3_ENDPOINT_URL, OPENAI_API_KEY, .env)..gitignore..env is git-ignored; keep real credentials only in local .env.git status
git add .
git commit -m "Consolidate AF3/Qwen pipelines, endpoint templates, and docs"
git push github main