Downloads · 30 days
59
54% of all-time downloads
stillerman/fdt-disfluency-medium-41m
fdt-disfluency-medium-41m is a token classification model from stillerman. Use it when you need labels on individual words, such as names. It is set up for transformers.js. The card lists the license as apache-2.0.
Disfluency deletion tagger for live speech transcripts: tags every whitespace word KEEP / DELETE / KEEPSTRIPCOMMA / KEEPCAPITALIZE, then a ~15-line reconstruction turns tags into cleaned text. Deletion-only by constru…
Downloads · 30 days
59
54% of all-time downloads
All-time downloads
110
Public
Parameters
41.1M
206 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors164 MB · 79%
From the Hugging Face model README
Disfluency deletion tagger for live speech transcripts: tags every
whitespace word KEEP / DELETE / KEEP_STRIP_COMMA / KEEP_CAPITALIZE,
then a ~15-line reconstruction turns tags into cleaned text. Deletion-only by
construction — it cannot rephrase, hallucinate, or alter names and numbers.
onnx/model_quantized.onnx (int8) is ready for transformers.js
(device: "webgpu", dtype: "q8"); runs at ~10–50 ms per utterance
in-browser.Trained on a DGX Spark as part of the FluencyAI digital-twin project.