Downloads · 30 days
176
14% of all-time downloads
anudit/lfm25-strudel
lfm25-strudel is a text generation model from anudit. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as other.
LoRA fine-tune of LiquidAI/LFM2.5-350M for natural-language - Strudel.cc live-coding music generation, fused with mlx-lm.
Downloads · 30 days
176
14% of all-time downloads
All-time downloads
1.2K
Public
Parameters
354M
8.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.onnx2.1 GB · 39%
From the Hugging Face model README
LoRA fine-tune of LiquidAI/LFM2.5-350M for natural-language -> Strudel.cc live-coding music generation, fused with mlx-lm.
Root: fused MLX weights (model.safetensors + config/tokenizer), ready for mlx-lm inference.
adapters/: LoRA adapter checkpoints saved during training (mlx-lm LoRA format, rank 32 / alpha 64).
onnx/: ONNX exports of the fused model for cross-platform / non-MLX inference:
model_fp32.onnx — full precisionmodel_bf16.onnx — bfloat16 weightsmodel_fp8.onnx — float8 (e4m3fn) weightsThe ONNX graphs take input_ids and attention_mask and return logits (no KV cache; each call is a full forward pass). They were exported from the fused weights after correcting mlx-lm's depthwise-conv weight layout ((dim, kernel, 1)) to the transformers Conv1d layout ((dim, 1, kernel)) expected by Lfm2ForCausalLM. The bf16/fp8 variants are weight-only casts of the fp32 graph (storage-size quants); verify operator/EP support for these dtypes before relying on them for compute.
pip install mlx-lm
mlx_lm.generate --model <this-repo> --prompt "// a fun pop indian lofi beat"
import onnxruntime as ort
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("<this-repo>")
sess = ort.InferenceSession("onnx/model_fp32.onnx", providers=["CPUExecutionProvider"])
inputs = tok("// a fun pop indian lofi beat\n", return_tensors="np")
logits = sess.run(None, {"input_ids": inputs["input_ids"], "attention_mask": inputs["attention_mask"]})[0]