Downloads · 30 days
22
22% of all-time downloads
liminalstoat/osim-4b-mlx-4bit
osim-4b-mlx-4bit is a text generation model from liminalstoat. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as mit.
A 4-bit MLX build of cmu-lti/osim-4b — CMU's OSim / OdysSim human-behavior-simulation model — for running on Apple Silicon (Mac, iPhone, iPad).
Downloads · 30 days
22
22% of all-time downloads
All-time downloads
100
Public
Parameters
4B
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
How the weights are stored.
U324B · 100%
From the Hugging Face model README
A 4-bit MLX build of cmu-lti/osim-4b — CMU's OSim / OdysSim human-behavior-simulation model — for running on Apple Silicon (Mac, iPhone, iPad).
This is the instruct-derived OSim 4B (built on Qwen/Qwen3-4B), so it carries the correct Qwen3 chat template and <|im_end|> stop token — chat works out of the box.
0.31.3 on 2026-06-28OSim (OdysSim) is a family of foundation models for human-behavior simulation from CMU LTI. It's trained to simulate how a person behaves in a conversation — to play the user, not the helpful assistant. Prompt it like a chatbot and it will do un-assistant-like things: ask its own questions, act like someone seeking help, hold a persona. That's the model working as intended. Use it where you want a synthetic human counterpart — dialogue-system testing, user simulation, behavioral data generation.
cmu-lti/osim-4b (MIT)Qwen/Qwen3-4B (Apache-2.0)This is a 4-bit quant of a 4B model, so there's some loss versus full precision — expect occasional arithmetic/reasoning slips and the odd repetition. For more headroom, convert a higher-bit MLX build (5/6/8-bit) from the same source, or run cmu-lti/osim-4b directly on a larger machine. None of this is a prompting problem; it's the 4-bit size trade.
pip install mlx-lm
mlx_lm.generate --model liminalstoat/osim-4b-mlx-4bit \
--prompt "Hi, what can you help me with?" --max-tokens 256
from mlx_lm import load, generate
model, tokenizer = load("liminalstoat/osim-4b-mlx-4bit")
print(generate(model, tokenizer, prompt="Hi, what can you help me with?", max_tokens=256))
The chat template ships with the model, so mlx_lm applies it automatically.
MLX runs on-device through mlx-swift. The most direct path is Apple's mlx-swift-examples app — point it at this repo or a local copy — or your own mlx-swift harness. Some MLX-based iOS chat apps can also load custom Hugging Face MLX repos; if yours supports adding a model by ID, use liminalstoat/osim-4b-mlx-4bit.
source: cmu-lti/osim-4b (instruct-derived; Qwen3-4B foundation)
tool: mlx_lm.convert --quantize --q-bits 4 --q-group-size 64
mlx-lm: 0.31.3
A straight 4-bit MLX conversion of CMU's published weights — no fine-tuning or merging, built from full-precision source (not from a pre-quantized model).
A research / tinkering artifact for on-device human-behavior simulation. It inherits the intended uses and limitations of the base cmu-lti/osim-4b, plus quantization loss. Not validated for production or factual QA. Because it simulates human behavior, outputs can be inconsistent, opinionated, or persona-driven by design.
cmu-lti/osim-4b — MIT (CMU LTI).Qwen/Qwen3-4B — Apache-2.0 (Qwen Team).