Downloads · 30 days
0
kevinzhow/kwja-mlx
kwja-mlx is a machine learning model from kevinzhow. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for mlx.
This repository contains MLX-converted weights for upstream KWJA checkpoints.
Downloads · 30 days
0
Access
Public
Updated Apr 10, 2026
Repo size
8.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors8.3 GB · 100%
From the Hugging Face model README
This repository contains MLX-converted weights for upstream KWJA checkpoints.
The files here are converted from the official KWJA checkpoint releases under kwja==2.5.1 / model version v2.4.
The original checkpoint files are downloaded from:
https://lotus.kuee.kyoto-u.ac.jp/kwja/v2.4/Converted modules are provided for three model sizes:
tinybaselargeEach size includes three modules:
charseq2seqword| Size | Module | Upstream checkpoint | Base encoder / model |
|---|---|---|---|
tiny | char | char_deberta-v2-tiny-wwm.ckpt | ku-nlp/deberta-v2-tiny-japanese-char-wwm |
tiny | seq2seq | seq2seq_t5-small.ckpt | retrieva-jp/t5-small-short |
tiny | word | word_deberta-v2-tiny.ckpt | ku-nlp/deberta-v2-tiny-japanese |
base | char | char_deberta-v2-base-wwm.ckpt | ku-nlp/deberta-v2-base-japanese-char-wwm |
base | seq2seq | seq2seq_t5-base.ckpt | retrieva-jp/t5-base-long |
base | word | word_deberta-v2-base.ckpt | ku-nlp/deberta-v2-base-japanese |
large | char | char_deberta-v2-large-wwm.ckpt | ku-nlp/deberta-v2-large-japanese-char-wwm |
large | seq2seq | seq2seq_t5-large.ckpt | retrieva-jp/t5-large-long |
large | word | word_deberta-v2-large.ckpt | ku-nlp/deberta-v2-large-japanese |
The conversion is a format conversion from PyTorch Lightning checkpoints to MLX runtime bundles.
High-level steps:
.ckpt file on CPU.hyper_parameters.state_dict..mlx.safetensors..mlx.json file.Module-specific notes:
char: exports the DeBERTa encoder and tagging heads; the convolution weight is transposed to match the MLX layout.seq2seq: exports the encoder-decoder weights after stripping the encoder_decoder. prefix and sanitizing them for the MLX T5 implementation.word: exports the DeBERTa encoder, prediction heads, and CRF transition parameters used at inference time.The conversion command used in the companion tooling is:
python -m hira kwja-export-mlx --model-size base --module all
The same conversion flow is used for tiny and large.
Each converted module is stored as a pair of files:
<size>/<module>.mlx.safetensors: MLX weights<size>/<module>.mlx.json: metadata needed to reconstruct the tokenizer/config/runtime bundleExamples:
base/char_deberta-v2-base-wwm.mlx.safetensorsbase/char_deberta-v2-base-wwm.mlx.jsonbase/seq2seq_t5-base.mlx.safetensorsbase/seq2seq_t5-base.mlx.jsonbase/word_deberta-v2-base.mlx.safetensorsbase/word_deberta-v2-base.mlx.jsonThis repository is intended to host converted MLX artifacts. It does not add new training, fine-tuning, or evaluation results beyond the upstream KWJA checkpoints.