Downloads · 30 days
25
6% of all-time downloads
vasanth009/LC-lfm2.5-350m
LC-lfm2.5-350m is a text generation model from vasanth009. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as other.
A small, fast on-device voice-dictation cleanup model: it turns messy spoken transcripts into clean written text, and — unlike most cleanup models — it honors spoken self-corrections ("book the 7pm flight no wait the…
Downloads · 30 days
25
6% of all-time downloads
All-time downloads
423
Public
Parameters
354M
250 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors250 MB · 98%
How the weights are stored.
U32354M · 100%
From the Hugging Face model README
A small, fast on-device voice-dictation cleanup model: it turns messy spoken transcripts into clean written text, and — unlike most cleanup models — it honors spoken self-corrections ("book the 7pm flight no wait the 9pm one" → "Book the 9pm flight.").
Built for MacWispr. This repo ships
the fused model (LoRA baked in) so you can pull and run it directly; the
standalone LoRA adapter is under lora-adapter/.
juanquivilla/sotto-cleanup-lfm25-350m-mlx-5bit (LFM2.5-350M, MLX 5-bit)mlx-swift-lm (LFM2)Trained on a raw completion format (not a chat template):
### Input:
{raw dictation}
### Output:
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tok = load("vasanth009/LC-lfm2.5-350m")
raw = "set the oven to three fifty no wait three seventy five for the lasagna"
prompt = f"### Input:\n{raw}\n\n### Output:\n"
out = generate(model, tok, prompt=prompt, max_tokens=64, sampler=make_sampler(temp=0.0))
print(out.split("###")[0].strip())
# -> Set the oven to 375 for the lasagna.
To apply the LoRA to the base yourself instead of using the fused weights:
mlx_lm.generate --model juanquivilla/sotto-cleanup-lfm25-350m-mlx-5bit \
--adapter-path lora-adapter --prompt "### Input:\n...\n\n### Output:\n"
Graded by an LLM judge on a 94-item held-out set generated on topics disjoint from training (0/94 overlap with the training data — verified). This is a real generalization test, not memorized phrases.
| Model | Course-correction | Light cleanup | Preserve (anti over-edit) |
|---|---|---|---|
| Base (Sotto LFM2.5-350M) | 10/16 | 12/12 | 6/8 |
| This model (+LoRA) | 13/16 | 12/12 | 7/8 |
Course-correction is the headline improvement (10→13/16). Light cleanup was already strong (tie). Latency ~50–100 ms/utterance on Apple Silicon.
Training data (course-correction / light-cleanup / preserve pairs) was generated
and QC-filtered with an LLM on fresh topics, with a hard leakage gate against the
held-out eval. See the MacWispr repo's bench/polish_finetune/ for the full,
reproducible pipeline.