Downloads · 30 days
0
undertheseanlp/bamboo-1
bamboo-1 is a token classification model from undertheseanlp. Use it when you need labels on individual words, such as names. It is set up for underthesea. The card lists the license as mit.
A Vietnamese dependency parser trained on the UDD-1 dataset using the Biaffine architecture.
Downloads · 30 days
0
Access
Public
Updated Jun 1, 2026
Repo size
2.8 GB
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 95%
From the Hugging Face model README
A Vietnamese dependency parser trained on the UDD-1 dataset using the Biaffine architecture.
Bamboo-1 is a neural dependency parser for Vietnamese that uses:
cd ~/projects/workspace_underthesea/bamboo-1
uv sync
from src.inference import Parser
parser = Parser("undertheseanlp/bamboo-1") # downloads the released safetensors model
sent = parser.parse("Tôi yêu Việt Nam")
print(sent.to_conllu())
# Train with default parameters
uv run scripts/train.py
# Train with custom parameters
uv run scripts/train.py --output models/bamboo-1 --max-epochs 200 --feat char
# Train with BERT embeddings
uv run scripts/train.py --feat bert --bert vinai/phobert-base
# Train with Weights & Biases logging
uv run scripts/train.py --wandb
# Evaluate trained model
uv run scripts/evaluate.py --model models/bamboo-1
# Interactive prediction
uv run scripts/predict.py --model models/bamboo-1
# Predict from file
uv run scripts/predict.py --model models/bamboo-1 --input input.txt --output output.conllu
The UDD-1 dataset is automatically downloaded from HuggingFace:
undertheseanlp/UDD-1Input: Vietnamese sentence
↓
Word Embeddings + Character LSTM Embeddings
↓
BiLSTM Encoder (3 layers, 400 hidden units)
↓
Biaffine Attention (Arc + Relation)
↓
Output: Dependency tree (head indices + relation labels)
The released checkpoint is the XLM-RoBERTa + Biaffine (Trankit-style) variant
models/bamboo-1.0.0-20260601-xlmr-udd1, trained on UDD-1 (whitespace-tokenized input).
| Split | UAS | LAS |
|---|---|---|
| UDD-1 dev | 88.70% | 82.37% |
| UDD-1 test | 89.25% | 82.87% |
Trained 100 epochs (batch 32, encoder LR 1e-5, head LR 1e-4, AdamW, FP16) on a single RTX 3090.
bamboo-1/
├── README.md
├── requirements.txt
├── scripts/
│ ├── train.py # Training script
│ ├── evaluate.py # Evaluation script
│ └── predict.py # Prediction script
├── bamboo1/
│ └── corpus.py # UDD-1 corpus loader
├── models/ # Trained models (generated)
└── data/ # Downloaded dataset (generated)
MIT License