Downloads · 30 days
22
19% of all-time downloads
stillerman/fdt-disfluency-tiny-4m
fdt-disfluency-tiny-4m is a token classification model from stillerman. Use it when you need labels on individual words, such as names. It is set up for transformers.js. The card lists the license as apache-2.0.
Disfluency deletion tagger for live speech transcripts: tags every whitespace word KEEP / DELETE / KEEPSTRIPCOMMA / KEEPCAPITALIZE, then a ~15-line reconstruction turns tags into cleaned text. Deletion-only by constru…
Downloads · 30 days
22
19% of all-time downloads
All-time downloads
113
Public
Parameters
4.4M
21.9 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors17.5 MB · 76%
From the Hugging Face model README
Disfluency deletion tagger for live speech transcripts: tags every
whitespace word KEEP / DELETE / KEEP_STRIP_COMMA / KEEP_CAPITALIZE,
then a ~15-line reconstruction turns tags into cleaned text. Deletion-only by
construction — it cannot rephrase, hallucinate, or alter names and numbers.
onnx/model_quantized.onnx (int8) is ready for transformers.js
(device: "webgpu", dtype: "q8"); runs at ~10–50 ms per utterance
in-browser.Trained on a DGX Spark as part of the FluencyAI digital-twin project.