Downloads · 30 days
18
20% of all-time downloads
dxtech-asia/deepx-embedding-v09
deepx-embedding-v09 is a feature extraction model from dxtech-asia. Use it when you need embeddings to search or compare text. It is set up for pytorch. The card lists the license as apache-2.0.
⚠️ Preview Release — This is a preview version for testing and evaluation. Not recommended for production use. Final v1.0 release will include improved quality and optimizations.
Downloads · 30 days
18
20% of all-time downloads
All-time downloads
91
Public
Repo size
1.8 GB
Likes
12
Public
Click a slice to open those files.
.pt1.8 GB · 98%
From the Hugging Face model README
⚠️ Preview Release — This is a preview version for testing and evaluation. Not recommended for production use. Final v1.0 release will include improved quality and optimizations.
DeepX Embedding 0.9 is a Vietnamese-focused embedding model built on a novel Gated DeltaNet-2 (GDN-2) linear attention architecture with O(n) complexity. It uses Hyperloop weight sharing (9 unique layers, 35 compute passes) with architecture design inspired by Google Gemma 4 E2B.
Key features:
| Component | Detail |
|---|---|
| Base | Gemma 4 E2B (262K vocab, 1536 hidden) |
| Attention | GDN-2 (Gated DeltaNet-2), pure linear O(n) |
| Structure | Hyperloop: Begin(4) + Phase1×2(5) + Phase2×4(5) + End(1) = 35 passes |
| Unique layers | 9 (shared via loop + per-iteration LoRA) |
| Total params | 889M (486M trainable backbone + 403M frozen token embedding) |
| Embedding dim | 1536 (Matryoshka: 256/512/768/1024/1536) |
| Max sequence | 2048 tokens (training), 131K (theoretical via YaRN RoPE) |
| Pooling | Attention-weighted pooling |
| Depth signal | RoDE (Rotary Depth Encoding) |
Unlike standard softmax attention (O(n²)), GDN-2 uses a gated delta rule recurrence computed via chunked parallel scan. This gives:
The model reuses 9 unique layer parameter sets across 35 compute passes:
Per-iteration LoRA (rank 16) differentiates each pass within a loop.
| Metric | DeepX v0.9 | mE5-large | vietlegal-harrier (SOTA) |
|---|---|---|---|
| nDCG@10 | 0.7449 | 0.6660 | 0.7813 |
| MRR@10 | 0.6921 | — | 0.7303 |
| Recall@10 | 0.9086 | — | 0.9321 |
| Sequence Length | DeepX (GDN-2) | Softmax Equivalent |
|---|---|---|
| 512 | 0.31 it/s | 0.31 it/s |
| 2048 | 0.18 it/s | 0.01 it/s |
| 8192 | ~0.18 it/s | OOM |
Linear attention maintains constant speed regardless of sequence length.
import torch
from transformers import AutoTokenizer
# Load
tokenizer = AutoTokenizer.from_pretrained("google/gemma-4-E2B-it")
# Load model (see repo for full pipeline code)
model = load_deepx_model("deepx_v09.pt")
# Encode
text = "Mức phạt khi vượt đèn đỏ là bao nhiêu?"
inputs = tokenizer(text, return_tensors="pt", max_length=2048, truncation=True)
with torch.no_grad():
embedding = model(inputs["input_ids"], attention_mask=inputs["attention_mask"], normalize=True)
# embedding.shape = (1, 1536)
# Matryoshka: use first N dims
embedding_256d = embedding[:, :256] # 90% quality, 6x less storage
@misc{deepx2026,
title={DeepX: Vietnamese Embedding Model with Gated DeltaNet-2 Linear Attention},
author={DXTech Asia},
year={2026},
url={https://huggingface.co/dxtech-asia/deepx-embedding-v09}
}
Apache 2.0 (code) / Model weights follow Gemma license terms.