Downloads · 30 days
0
mlboydaisuke/Mordant-3B-Think-LiteRT
Mordant-3B-Think-LiteRT is a text generation model from mlboydaisuke. Use it when you need the model to write or continue text. It is set up for litert. The card lists the license as apache-2.0.
On-device conversion of Kezmark/Mordant-3B-Think — a full fine-tune of ibm-granite/granite-4.1-3b for AI image-generation prompt composition with chain-of-thought reasoning — to a .litertlm bundle for the LiteRT-LM ru…
Downloads · 30 days
0
Access
Public
Updated Oct 7, 2026
Repo size
7.5 GB
Likes
1
Public
Click a slice to open those files.
.litertlm3.8 GB · 100%
From the Hugging Face model README
On-device conversion of Kezmark/Mordant-3B-Think —
a full fine-tune of ibm-granite/granite-4.1-3b
for AI image-generation prompt composition with chain-of-thought reasoning — to a .litertlm
bundle for the LiteRT-LM runtime. All credit for
the model itself goes to its author; this repo only packages it for phones and desktops.
Requires litert-lm ≥ 0.16 to run.
| file | quant | size |
|---|---|---|
Mordant-3B-Think_int8.litertlm | dynamic int8 (linears + embedding) | 3.76 GB |
2026-09-21: chat template updated to accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical.
Converted with one command by hf-to-litertlm
(python scripts/convert.py Kezmark/Mordant-3B-Think, 2026-08-25):
<think> — is embedded verbatim (byte-equal to the checkpoint's
chat_template.jinja, 1474/1474).bos == eos == <|end_of_text|> and its template never renders a leading BOS, so an engine-prepended
start token reads as "this document already ended" — measured on this checkpoint, it flips
HF bf16 greedy output into a code-fence loop.-p 256 -d 256 --runs 3 --cache no)| backend | prefill tok/s | decode tok/s | TTFT |
|---|---|---|---|
| CPU | 97.3 | 20.2 | 2.68 s |
| GPU | 1129 | 71.9 | 0.24 s |
pip install litert-lm
litert-lm run Mordant-3B-Think_int8.litertlm \
--prompt "A cat sitting on a windowsill at sunset" --max-num-tokens 4096
The model answers with a <think>…</think> block followed by the composed image prompt —
budget generation length accordingly. On Android, load the bundle in an app embedding the
LiteRT-LM engine (e.g. Google AI Edge Gallery-style hosts).
apache-2.0, inherited from the source model and its granite-4.1 base.