Downloads · 30 days
46
44% of all-time downloads
Rubin-Wei/MemoryDecoder-Pythia-6.9B-general
MemoryDecoder-Pythia-6.9B-general is a text generation model from Rubin-Wei. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
- Project Page: Memory Decoder at Scale - GitHub Repository: LUMIA-Group/MemoryDecoder-at-Scale - Paper: Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory - Hugging Face Collection: MemoryDecoder-at-S…
Downloads · 30 days
46
44% of all-time downloads
All-time downloads
104
Public
Parameters
6.9B
13.7 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors13.7 GB · 100%
From the Hugging Face model README
This repository contains the 6.9B general Memory Decoder released with Memory Decoder at Scale. It is a pretrained parametric long-term memory designed to be paired with a frozen Pythia-family language model. The general-memory suite was pretrained on 300B tokens and scales memory capacity independently from the backbone.
This checkpoint is a memory component, not a chat- or instruction-tuned model. Use it through the Memory Decoder integration in the released codebase rather than treating it as a standalone assistant.
| Field | Value |
|---|---|
| Memory size | approximately 6.9B parameters |
| Architecture/tokenizer family | GPT-NeoX / Pythia |
| Memory scope | General |
| Training scale | 300B-token general-memory pretraining suite |
| Intended backbone | Frozen, tokenizer-compatible Pythia model |
| Release contents | Inference weights, configuration, and tokenizer files |
Install the environments from
LUMIA-Group/MemoryDecoder-at-Scale, then use the
hf-memdec adapter. For example:
cd eval/lm-evaluation-harness
BACKBONE=EleutherAI/pythia-410m-deduped
MEMORY=Rubin-Wei/MemoryDecoder-Pythia-6.9B-general
lm-eval \
--model hf-memdec \
--model_args pretrained=$BACKBONE,memdec_path=$MEMORY \
--tasks arc_easy,piqa,mmlu \
--batch_size 1
The adapter provides the task-specific interpolation settings used for the
released Pythia memories; an explicit lmbda can be supplied as an override.
This checkpoint is intended for research on parametric memory, memory scaling, and evaluation with frozen language-model backbones. Its behavior depends on the selected backbone and interpolation settings. It may reproduce biases or errors present in its training data and should not be treated as an authoritative knowledge source.
If you use this checkpoint, please cite:
@misc{wei2026memorydecoderscalepretrained,
title={Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory},
author={Rubin Wei and Jiaqi Cao and Jiarui Wang and Junming Zhang and Qipeng Guo and Bowen Zhou and Zhouhan Lin},
year={2026},
eprint={2607.27919},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2607.27919},
}