Downloads · 30 days
10
12% of all-time downloads
ngovanphuoc2006/Legal-embedding
Legal-embedding is a feature extraction model from ngovanphuoc2006. Use it when you need embeddings to search or compare text. It is set up for peft.
This model is a parameter-efficient fine-tuned (PEFT) version of Qwen/Qwen3-Embedding-8B specifically adapted for the Vietnamese Legal Domain. It uses LoRA (Low-Rank Adaptation) to capture the nuances of legal termino…
Downloads · 30 days
10
12% of all-time downloads
All-time downloads
84
Public
Repo size
26.8 MB
Likes
1
Public
Click a slice to open those files.
.safetensors15.4 MB · 49%
From the Hugging Face model README
This model is a parameter-efficient fine-tuned (PEFT) version of Qwen/Qwen3-Embedding-8B specifically adapted for the Vietnamese Legal Domain. It uses LoRA (Low-Rank Adaptation) to capture the nuances of legal terminology and semantics in Vietnamese statutory documents.
Bạn có thể sử dụng model này với thư viện transformers và peft theo cấu trúc sau:
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel, PeftConfig
import torch
# Đường dẫn repo
model_id = "ngovanphuoc2006/Legal-embedding"
# Load cấu hình và model
config = PeftConfig.from_pretrained(model_id)
base_model = AutoModel.from_pretrained(config.base_model_name_or_path, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(config.base_model_name_or_path)
# Merge Adapter
model = PeftModel.from_pretrained(base_model, model_id)
# Ví dụ sử dụng
sentences = ["Quy định về tội giết người", "Các hình phạt đối với hành vi cố ý gây thương tích"]
inputs = tokenizer(sentences, padding=True, truncation=True, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
# Lấy embedding từ Last Hidden State (thường là CLS token hoặc mean pooling)
embeddings = outputs.last_hidden_state[:, 0, :]