Downloads · 30 days
58
8% of all-time downloads
diverWayne/mikky-64m
mikky-64m is a text generation model from diverWayne. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
mikky-64m is a 63,912,192-parameter small language model named mikky. It was trained by HUANG JUNZHE 黄俊哲 with the minimind-scratch codebase, based on the MiniMind project/data format.
Downloads · 30 days
58
8% of all-time downloads
All-time downloads
718
Public
Parameters
68.8M
413 MB on disk
Likes
1
Public
Click a slice to open those files.
.gguf138 MB · 33%
From the Hugging Face model README
mikky-64m is a 63,912,192-parameter small language model named mikky.
It was trained by HUANG JUNZHE 黄俊哲 with the minimind-scratch codebase, based on the MiniMind project/data format.
This release is intended as a compact learning and experimentation checkpoint for local inference, model-format conversion, and small-model alignment workflows.
The released checkpoint uses the completed alignment path:
pretrain -> SFT -> mikky LoRA identity SFT -> DPO
GRPO was only run as a probe and is not used as the final release checkpoint. PPO was skipped because the local reward signal was not strong enough to justify another RL stage.
The model identity/persona is:
mikky-64m.pth: native minimind_scratch state dict, BF16 tensors.model.safetensors: Qwen3-compatible Hugging Face tensor names, BF16 tensors.mikky-64m-bf16.gguf: llama.cpp GGUF export, BF16, not quantized.tokenizer.json, tokenizer_config.json: MiniMind tokenizer files.config.json, generation_config.json: Qwen3-compatible metadata used for conversion and loading.The final source checkpoint was checkpoints/dpo_768_resume.pth.
The training code uses MiniMind chat markers:
<|im_start|>user
你的问题<|im_end|>
<|im_start|>assistant
Use the project code for native scratch inference:
python -m minimind_scratch.cli chat \
--weight out/hf/mikky-64m/mikky-64m.pth \
--prompt "请用一句话介绍你自己"
The GGUF file is BF16 and intentionally not quantized:
llama-cli -m mikky-64m-bf16.gguf \
-p "<|im_start|>user\n请用一句话介绍你自己<|im_end|>\n<|im_start|>assistant\n" \
-n 128
The GGUF export maps the scratch model to a Qwen3-compatible tensor layout because the model uses RMSNorm, SwiGLU MLP, grouped-query attention, RoPE, and q/k normalization. The GGUF structure and metadata were verified locally. Always verify generation quality in your target runtime before treating the GGUF file as production-ready.
This model was trained with the MiniMind small-data recipe from
jingyaogong/minimind_dataset.
For this release, the dataset reference follows the MiniMind small dataset license: Apache-2.0.
Main data files used by this run:
pretrain_t2t_mini.jsonl: pretraining data.sft_t2t_mini.jsonl: supervised fine-tuning data.dpo.jsonl: preference data for DPO.lora_identity_mikky.jsonl: project-authored identity/persona data for mikky.The model card, exported native checkpoint, Safetensors checkpoint, and GGUF artifact are released under Apache-2.0.