Downloads · 30 days
0
Dat1710/gemma-dgmc-piqa
gemma-dgmc-piqa is a machine learning model from Dat1710. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Adapter DGMC (Dual-Gated Memory Consolidation) huấn luyện trên bộ dữ liệu PIQA, gắn thêm vào mô hình nền tảng đóng băng hoàn toàn google/gemma-4-E2B-it-qat-q40-unquantized. Chỉ các tham số của module DGMC được huấn lu…
Downloads · 30 days
0
Access
Public
Updated Jul 4, 2026
Repo size
56.8 MB
Likes
0
Public
Click a slice to open those files.
.pt56.7 MB · 100%
From the Hugging Face model README
Adapter DGMC (Dual-Gated Memory Consolidation) huấn luyện trên bộ dữ liệu
PIQA, gắn thêm vào mô hình nền tảng
đóng băng hoàn toàn google/gemma-4-E2B-it-qat-q4_0-unquantized. Chỉ các tham số của module DGMC được
huấn luyện (dgmc_params/1e6:.1fM tham số).
| Model | Method | Trainable Params | Accuracy | F1 Macro |
|---|---|---|---|---|
| Gemma (Zero-Shot) | Next-token logit scoring | 0 | 75.73% | 0.7561 |
| Gemma + DGMC | DGMC fine-tuned (LM) | 28.3M | 77.20% | 0.7719 |
Cải thiện: +1.47 điểm phần trăm so với zero-shot.
block_size: 64memory_dim: 1536decay_alpha: 0.1Tải file dgmc_piqa_weights.pt rồi load lại vào module DGMC tương ứng với
kiến trúc trong notebook huấn luyện gốc:
import torch
ckpt = torch.load("dgmc_piqa_weights.pt", map_location="cpu")
dgmc.load_state_dict(ckpt["dgmc_state_dict"])
Model nền tảng google/gemma-4-E2B-it-qat-q4_0-unquantized cần được tải riêng và giữ đóng băng (frozen);
adapter DGMC này chỉ bổ sung một module memory consolidation nhẹ, không thay
thế attention gốc của model.