Downloads · 30 days
3
4% of all-time downloads
lordChipotle/SimaQian
SimaQian is a machine learning model from lordChipotle. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Ancient Chinese Translator + Phonology Model (SimaQian)
Downloads · 30 days
3
4% of all-time downloads
All-time downloads
83
Public
Parameters
9.2B
18.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors18.5 GB · 100%
From the Hugging Face model README
Ancient Chinese Translator + Phonology Model (SimaQian)
Name Origin:
The origin of the model name comes from famous ancient chinese historian Qian Sima (司馬遷), known for his Records of the Grand Historian, a general history of China covering more than two thousand years.
This model combines two key functionalities for Ancient Chinese texts:
1. Translation: Converts Ancient Chinese passages into modern Chinese.
2. Phonological Reconstruction: Provides historical pronunciations for characters or entire sentences across multiple eras (e.g., Middle Tang, Song, Yuan, Ming/Qing).
Model Description
• Architecture: Fine-tuned on top of Google’s Gemma 2 model using LoRA.
• Input Format: Special tokens <start_of_turn> / <end_of_turn> define user vs. model turns.
• Output: Era identification (optional), phonetic renderings, and modern Chinese translations.
Training Data • Translation: Erya dataset from RUCAIBox/Erya-dataset. • Phonology: Ancient-Chinese-Phonology (ACP) for multi-era reconstructions. • Fine-Tuning: LoRA-based parameter-efficient approach on Gemma 2 Instruct.
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("lordChipotle/SimaQian")
model = AutoModelForCausalLM.from_pretrained("lordChipotle/SimaQian")
prompt = """ <start_of_turn>user Given the ancient text: 「子曰:學而時習之,不亦說乎?」
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=256)
print(tokenizer.decode(outputs[0]))
Limitations and Biases
• Era Estimation: Model may not always correctly guess the historical era.
• Pronunciations: Reconstructions are approximate and can vary by scholarly consensus.
• Contextual Accuracy: For highly contextual Ancient Chinese passages, translations may need further review by domain experts.