Downloads · 30 days
25
17% of all-time downloads
linear-moe-hub/GSA-340M
GSA-340M is a machine learning model from linear-moe-hub. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Model of the paper MoM: Linear Sequence Modeling with Mixture-of-Memories and Gated Slot Attention for Efficient Linear-Time Sequence Modeling.
Downloads · 30 days
25
17% of all-time downloads
All-time downloads
151
Public
Parameters
355M
711 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors710 MB · 99%
From the Hugging Face model README
Model of the paper MoM: Linear Sequence Modeling with Mixture-of-Memories and Gated Slot Attention for Efficient Linear-Time Sequence Modeling.
The model was trained on a sample of SlimPajama with 15B tokens.
Due to changes in the MLP layer structure in the latest version of fla, the weights cannot be loaded. You can use the version at fla instead.