Downloads · 30 days
0
leoncynn/paper8-cross-model-compression
paper8-cross-model-compression is a machine learning model from leoncynn. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Artifacts for the paper by Yeonseong Cynn (River Lab, May 2026).
Downloads · 30 days
0
Access
Public
Updated May 7, 2026
Repo size
438 MB
Likes
0
Public
Click a slice to open those files.
.pt438 MB · 100%
From the Hugging Face model README
Artifacts for the paper by Yeonseong Cynn (River Lab, May 2026).
Decomposes transformer FFN layers into structural (format-preserving) and classification-relevant components across BERT and GPT-2.
Key findings:
bert_sst2_int4_ste.pt — BERT SST-2 with L1-L3 INT4 quantization + STE retraining. Standard BERT state_dict, loadable directly. Accuracy: 90.1% (original FP32: 92.4%).results/bert/)bert_structural_prune.json — Per-layer structural pruning results (head/FFN reduction, accuracy)bert_sst2_all_prune.json — All-layer simultaneous FFN pruning resultsbert_l8_prune_results.json — L8 FFN correction + pruning (multi-seed)bert_quantize_results.json — INT4/INT8 post-training quantization resultsbert_quantize_retrain.json — INT4 STE retraining resultsresults/gpt2/)gpt2_structural_prune.json — Per-layer structural pruning (head + FFN)gpt2_each_layer_prune.json — Individual layer compression resultsgpt2_prune_validate.json — Pruning validation (PPL, accuracy)figures/fig1_ratio.png — FFN dual role ratio: BERT vs GPT-2 (log scale)figures/fig2_compression.png — Per-layer compression rates comparisonfigures/fig3_pruning.png — BERT SST-2 FFN neuron pruning curvefigures/fig4_quantization.png — INT4 quantization results (PTQ vs STE)MIT