Downloads · 30 days
227
100% of all-time downloads
minjaechoi/Midm-2.0-Mini-Instruct-W4A16
Midm-2.0-Mini-Instruct-W4A16 is a text generation model from minjaechoi. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
4-bit weight-only (W4A16) quantization of K-intelligence/Midm-2.0-Mini-Instruct, produced with LLM Compressor's GPTQModifier. All Linear layers are quantized except lmhead, which is kept at full precision.
Downloads · 30 days
227
100% of all-time downloads
All-time downloads
227
Public
Parameters
2.3B
1.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.5 GB · 99%
How the weights are stored.
I322.1B · 90%
From the Hugging Face model README
4-bit weight-only (W4A16) quantization of K-intelligence/Midm-2.0-Mini-Instruct, produced with
LLM Compressor's GPTQModifier. All Linear layers are quantized
except lm_head, which is kept at full precision.
| Method | GPTQ (GPTQModifier, LLM Compressor) |
| Scheme | W4A16 (4-bit weights, 16-bit activations) |
| Weight dtype | INT4, symmetric |
| Group size | 128 |
| Quantized targets | Linear (all layers except lm_head) |
| Activation ordering | static |
| Dampening frac | 0.01 |
| Output format | compressed-tensors (pack-quantized) |
Recipe used (recipe.yaml, included in this repo):
default_stage:
default_modifiers:
GPTQModifier:
targets: [Linear]
ignore: [lm_head]
scheme: W4A16
block_size: 128
dampening_frac: 0.01
actorder: static
requires_calibration_data: true
256 packed sequences × 2048 tokens = 524,288 calibration tokens, sampled (seed=42) from a mixed Korean/English instruction + function-calling corpus, targeting the following source composition:
| Source | Target % | Actual % (this model) | Tokens (this model) |
|---|---|---|---|
| KRX-Data/Won-Instruct | 35% | 34.19% | 179,256 |
| heegyu/glaive-function-calling-v2-ko | 30% | 29.80% | 156,212 |
| NousResearch/hermes-function-calling-v1 (json-mode-agentic.json) | 10% | 10.35% | 54,259 |
| heegyu/open-korean-instructions | 10% | 10.16% | 53,265 |
| kuotient/gsm8k-ko | 5% | 5.28% | 27,698 |
| in-house synthetic data (SafeCommit project, programmatically generated) | 10% | 10.22% | 53,598 |
safecommit_synth is unreleased in-house synthetic data from the SafeCommit project, not a public HF dataset.
compressed-tensors W4A16 kernels used here)vllm serve minjaechoi/Midm-2.0-Mini-Instruct-W4A16 --served-model-name midm2-mini
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "midm2-mini",
"messages": [{"role": "user", "content": "Hello!"}]
}'
compressed-tensors package for int4 dequant kernels)from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "minjaechoi/Midm-2.0-Mini-Instruct-W4A16"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
msgs = [{"role": "user", "content": "Hello!"}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=256)
print(tok.decode(out[0], skip_special_tokens=True))
Evaluated on this W4A16 checkpoint (not compared against an FP16 baseline run in this project yet):
--local-model-path was not
substituted into the API request, so every call 404'd against the vLLM server) — the resulting scores are invalid
and are intentionally not published here. Will be updated after a corrected re-run.model.safetensors — quantized weights (compressed-tensors pack-quantized format)config.json — includes the quantization_config (compressed-tensors) needed by vLLM/transformers to load this checkpointrecipe.yaml — the exact LLM Compressor recipe used to produce this checkpointtokenizer.json, tokenizer_config.json, chat_template.jinja — tokenizer/chat template, copied unmodified from the base modelLICENSE — base model license, included per its termsThis checkpoint is a derivative of K-intelligence/Midm-2.0-Mini-Instruct and is distributed
under the same license (mit, see LICENSE in this repo). No additional restrictions are added beyond
the base model's license.