Downloads · 30 days
50
16% of all-time downloads
GameGC/questions-lfm2-4bit
questions-lfm2-4bit is a text generation model from GameGC. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
LFM2.5-230M fine-tuned with MLX LoRA to extract the most recent question from a multi-turn dialogue. Input is a transcript tagged with [S] and [M] line prefixes. The model prioritizes questions appearing under the [S]…
Downloads · 30 days
50
16% of all-time downloads
All-time downloads
306
Public
Parameters
230M
129 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors129 MB · 96%
How the weights are stored.
U32230M · 100%
From the Hugging Face model README
LFM2.5-230M fine-tuned with MLX LoRA to extract the most recent question from a multi-turn dialogue. Input is a transcript tagged with [S] and [M] line prefixes. The model prioritizes questions appearing under the [S] tag; if no question is present there, it falls back to the latest [M] block.
End-to-end latency on M-series Mac: 100–150 ms per call (greedy decoding, max 80 tokens, 4-bit quantized).
Transcript with [S] and [M] line prefixes:
[S] what would you like to discuss today
[M] i was thinking about the architecture of the new service
[S] ok
[M] could you walk me through the current approach
A single extracted question (no question mark, no quotes). Priority order:
[S][M][M] contentfrom mlx_lm import load, generate
model, tokenizer = load("GameGC/questions-lfm2-4bit")
messages = [
{"role": "system", "content": "Extract the most recent question from the dialogue."},
{"role": "user", "content": "[S] what would you like to discuss\n[M] i was thinking about the architecture\n[S] ok\n[M] could you walk me through the current approach"},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
out = generate(model, tokenizer, prompt=prompt, max_tokens=80)
print(out)
import MLXLMCommon
let config = ModelConfiguration(id: "GameGC/questions-lfm2-4bit")
| Metric | Value |
|---|---|
| Latency (M-series Mac, 4-bit) | 100–150 ms per call |
| Model size on disk | ~134 MB |
| Max context | 1024 tokens |
| Quantization | 4-bit, group size 64 |
[M] containing question-sounding words ("how", rhetorical "right") may trigger false-positive extraction. Will be addressed in v2.| Hyperparameter | Value |
|---|---|
| Base | LiquidAI/LFM2.5-230M |
| Method | MLX LoRA |
| Rank | 32 |
| Alpha | 64 |
| LoRA keys | self_attn.{q,k,v,out}_proj, feed_forward.{w1,w2,w3} |
| Trainable params | 6.19M (2.696%) |
| Iters | 200 |
| Learning rate | 2e-4 (cosine decay, warmup 10) |
| Max seq length | 1024 |
| Batch size | 4 × 4 grad accum |
| Val loss | 3.264 → 0.360 |
| Train loss | 3.144 → 0.368 |
| Duration | 110s on M-series Mac |