Downloads · 30 days
21
25% of all-time downloads
H0ARK/wave-density-130m
wave-density-130m is a text generation model from H0ARK. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This repository contains a 130M parameter causal language model built with Wave-Density Attention (WDA), a novel alternative to standard dot-product self-attention.
Downloads · 30 days
21
25% of all-time downloads
All-time downloads
84
Public
Parameters
134M
535 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors535 MB · 100%
How the weights are stored.
F32133M · 99%
From the Hugging Face model README
This repository contains a 130M parameter causal language model built with Wave-Density Attention (WDA), a novel alternative to standard dot-product self-attention.
WDA reframes attention as a wave-interference and density-rendering process, replacing the traditional $QK^\top$ similarity computation with learned frequency-based interactions. This allows attention patterns to emerge from constructive and destructive interference rather than explicit pairwise dot products.
⸻
This combination provides both general language coverage and instruction-following coherence, while allowing the WDA mechanism to learn stable long-range structure.
⸻
Despite using a fundamentally different attention formulation, WDA achieves competitive perplexity and strong qualitative coherence at this scale.
⸻
To use this model, install or clone the reference implementation from the official repository:
Example loading snippet:
from wave_dencity import WaveCharLM
import torch
import json
# Load model configuration
with open("config.json", "r") as f:
config = json.load(f)
model = WaveCharLM(**config)
# Load weights from model.safetensors
# model.load_state_dict(...)
model.eval()
Note: This model is intended for research and experimentation with alternative attention mechanisms. The codebase exposes WDA internals for inspection and modification.
⸻
Traditional attention relies on sharp token-to-token similarity. WDA instead:
This approach avoids explicit dot-product similarity while still supporting coherent, causal language modeling.
⸻
If you use this model or the Wave-Density Attention mechanism in your work, please cite the official repository and paper.