Downloads · 30 days
33
5% of all-time downloads
songhieng/khmer-mt5-summarization
khmer-mt5-summarization is a summarization model from songhieng. Use it when you need a shorter version of a longer text. It is set up for transformers. The card lists the license as mit.
This repository contains a fine-tuned mT5 model for Khmer text summarization. The model is based on Google's mT5-small and fine-tuned on a dataset of Khmer text and corresponding summaries.
Downloads · 30 days
33
5% of all-time downloads
All-time downloads
728
Public
Parameters
300M
1.9 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors1.2 GB · 64%
From the Hugging Face model README
This repository contains a fine-tuned mT5 model for Khmer text summarization. The model is based on Google's mT5-small and fine-tuned on a dataset of Khmer text and corresponding summaries.
Fine-tuning was performed using the Hugging Face Trainer API, optimizing the model to generate concise and meaningful summaries of Khmer text.
google/mt5-smallkimleang123/khmer-text-datasettransformersEnsure you have transformers, torch, and datasets installed:
pip install transformers torch datasets
To load and use the fine-tuned model:
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
model_name = "songhieng/khmer-mt5-summarization"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)
def summarize_khmer(text, max_length=150):
input_text = f"summarize: {text}"
inputs = tokenizer(input_text, return_tensors="pt", truncation=True, max_length=512)
summary_ids = model.generate(**inputs, max_length=max_length, num_beams=5, length_penalty=2.0, early_stopping=True)
summary = tokenizer.decode(summary_ids[0], skip_special_tokens=True)
return summary
khmer_text = "កម្ពុជាមានប្រជាជនប្រមាណ ១៦ លាននាក់ ហើយវាគឺជាប្រទេសនៅតំបន់អាស៊ីអាគ្នេយ៍។"
summary = summarize_khmer(khmer_text)
print("🔹 Khmer Summary:", summary)
For a simpler approach:
from transformers import pipeline
summarizer = pipeline("summarization", model="songhieng/khmer-mt5-summarization")
khmer_text = "កម្ពុជាមានប្រជាជនប្រមាណ ១៦ លាននាក់ ហើយវាគឺជាប្រទេសនៅតំបន់អាស៊ីអាគ្នេយ៍។"
summary = summarizer(khmer_text, max_length=150, min_length=30, do_sample=False)
print("🔹 Khmer Summary:", summary[0]['summary_text'])
You can create a simple API for summarization:
from fastapi import FastAPI
app = FastAPI()
@app.post("/summarize/")
def summarize(text: str):
inputs = tokenizer(f"summarize: {text}", return_tensors="pt", truncation=True, max_length=512)
summary_ids = model.generate(**inputs, max_length=150, num_beams=5, length_penalty=2.0, early_stopping=True)
summary = tokenizer.decode(summary_ids[0], skip_special_tokens=True)
return {"summary": summary}
# Run with: uvicorn filename:app --reload
The model was evaluated using ROUGE scores, which measure how similar the generated summaries are to the ground truth summaries.
from datasets import load_metric
rouge = load_metric("rouge")
def compute_metrics(pred):
labels_ids = pred.label_ids
pred_ids = pred.predictions
decoded_preds = tokenizer.batch_decode(pred_ids, skip_special_tokens=True)
decoded_labels = tokenizer.batch_decode(labels_ids, skip_special_tokens=True)
return rouge.compute(predictions=decoded_preds, references=decoded_labels)
trainer.evaluate()
After fine-tuning, the model was uploaded to Hugging Face Hub:
model.push_to_hub("songhieng/khmer-mt5-summarization")
tokenizer.push_to_hub("songhieng/khmer-mt5-summarization")
To download it later:
model = AutoModelForSeq2SeqLM.from_pretrained("songhieng/khmer-mt5-summarization")
tokenizer = AutoTokenizer.from_pretrained("songhieng/khmer-mt5-summarization")
| Feature | Details |
|---|---|
| Base Model | google/mt5-small |
| Task | Summarization |
| Language | Khmer (ខ្មែរ) |
| Dataset | kimleang123/khmer-text-dataset |
| Framework | Hugging Face Transformers |
| Evaluation Metric | ROUGE Score |
| Deployment | Hugging Face Model Hub, API (FastAPI), Python Code |
Contributions are welcome! Feel free to open issues or submit pull requests if you find any improvements.
If you have any questions, feel free to reach out via Hugging Face Discussions or create an issue in the repository.
📌 Built for Khmer NLP Community 🇰🇭 🚀