Downloads · 30 days
836
41% of all-time downloads
cstr/beat-this-GGUF
beat-this-GGUF is a audio classification model from cstr. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for crispasr. The card lists the license as mit.
GGUF conversions of Beat This! (Foscarin, Schlüter & Widmer, CPJKU — ISMIR 2024), a transformer beat and downbeat tracker, for use with CrispASR / ggml.
Downloads · 30 days
836
41% of all-time downloads
All-time downloads
2K
Public
Repo size
122 MB
Likes
0
Public
Click a slice to open those files.
.gguf122 MB · 100%
From the Hugging Face model README
GGUF conversions of Beat This! (Foscarin, Schlüter & Widmer, CPJKU — ISMIR 2024), a transformer beat and downbeat tracker, for use with CrispASR / ggml.
Checkpoint: final0 (trained on all datasets except GTZAN, seed 0).
Nearly every published beat tracker post-processes its framewise output with madmom's Dynamic Bayesian Network, which is Böck-patented and licensed for non-commercial use only. Beat This! reaches state-of-the-art without one — its postprocessing is plain peak-picking, and its dependency list (numpy / torch / torchaudio / einops / rotary-embedding-torch / soxr) contains no part of madmom.
Both the upstream code and the published weights are MIT. That combination — SOTA accuracy, no DBN, MIT weights — is why this is usable in a commercial product where most beat trackers are not.
| file | size | notes |
|---|---|---|
beat-this-f16.gguf | 41 MB | default. Recommended for all normal use. |
beat-this-f32.gguf | 81 MB | Reference build for exact-parity debugging. |
20.25 M parameters, 147 tensors. Both files bake in the [513, 128] mel
filterbank the model was exported with, so the front end never re-derives it
(slaney-vs-htk and the freq/mel layout are classic silent-drift sources).
Ported stage by stage against a PyTorch reference driven by the original checkpoint, comparing every sub-block rather than only the final output.
At f32, every stage is numerically exact:
| stage | cos | max rel err |
|---|---|---|
stem | 1.00000000 | 3.0e-7 |
blk0_partial | 1.00000000 | 3.0e-7 |
blk2 | 1.00000000 | 9.5e-7 |
linear | 1.00000000 | 1.3e-6 |
transformer | 1.00000000 | 9.3e-7 |
out_beat / out_downbeat | 1.00000000 | 2.7e-6 / 1.4e-6 |
At f16 the same stages score cos ≥ 0.99999973 with max rel err ~5e-4, flat across all 12 attention/FF sub-blocks rather than compounding — i.e. the residual is weight quantisation, not drift. The log-mel front end matches torchaudio at cos = 1.00000000.
Windowing (1500-frame chunks, 6-frame border, keep_first overlap) and the
peak-picking postprocessor reproduce upstream exactly: on a 45 s two-chunk
fixture, running the reference's own logits through the ported peak-picker gives
identical beat and downbeat times to 1e-6 s.
# one line per beat: time_sec <TAB> beat|downbeat
crispasr --beats -m beat-this-f16.gguf -f song.wav
# JSON, including a median-interval tempo estimate
crispasr --beats -m beat-this-f16.gguf --beats-format json -f song.wav
Input is decoded to the model's native 22.05 kHz mono automatically, and long files are chunked internally.
Every downbeat is also reported as a beat. The postprocessor snaps each downbeat onto its nearest detected beat, so downbeats are a strict subset and you never have to merge two lists to reconstruct the grid.
Code and weights are MIT, per upstream. Note that upstream's own README records that some of the training audio is copyrighted or under restrictive Creative Commons terms; the licence on the released weights is unambiguous, but the provenance of the training corpus is a separate question that upstream, not this conversion, is the authority on.
@inproceedings{foscarin2024beatthis,
title = {Beat this! Accurate beat tracking without DBN postprocessing},
author = {Foscarin, Francesco and Schl{\"u}ter, Jan and Widmer, Gerhard},
booktitle = {Proceedings of the 25th International Society for
Music Information Retrieval Conference (ISMIR)},
year = {2024}
}