Downloads · 30 days
3
9% of all-time downloads
channudam/stable-diffusion-khm-53
stable-diffusion-khm-53 is a text-to-image model from channudam. Use it when you need an image from a text prompt. It is set up for diffusers.
This repository hosts a fine-tuned Stable Diffusion model customized for Khmer text image generation. The model aims to generate high-quality synthetic data, particularly for applications such as Khmer OCR, document l…
Downloads · 30 days
3
9% of all-time downloads
All-time downloads
34
Public
Parameters
147M
468 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors468 MB · 100%
From the Hugging Face model README
This repository hosts a fine-tuned Stable Diffusion model customized for Khmer text image generation. The model aims to generate high-quality synthetic data, particularly for applications such as Khmer OCR, document layout analysis, and AI-based Khmer text systems.
Despite rapid advances in AI and generative models, Khmer remains a low-resource language, lacking high-quality datasets and models for tasks like text-to-image generation, OCR, and scene text analysis. Compared to languages like Thai or Vietnamese, Khmer lacks sufficient publicly available data, especially in image form, making it difficult to develop robust AI systems.
The primary objective of this project was to:
Goal: To build a complete, scalable, and publicly accessible pipeline that can transform Khmer text into realistic images for downstream use in OCR and machine learning.
Scope:
This model was developed as part of a 4-month internship at Factory.io under the Cambodia Academy of Digital Technology (CADT), with the main objective of generating synthetic images of Khmer script from text prompts. The final output is an end-to-end text-to-image generation pipeline, fine-tuned on Khmer word images using the Stable Diffusion architecture.
Hugging Face Collection:
🔗 https://huggingface.co/collections/channudam/textimagegeneration-khm-35-67d916c2505635db1ba8fc3c
| Model Type | Output Quality | Params | Image Size |
|---|---|---|---|
| DCGAN | Low | 239K | 64x64, 1-chan |
| UNet2D | Good | 106M | 64x64, 1-chan |
| UNet2DConditional | Very Good | 147M | 64x32, 1-chan |
| Stable Diffusion | Excellent | 881M (Total) | 128x64, RGB |
✔️ Stable Diffusion outperformed all other methods, producing sharper, more accurate Khmer text images.
✔️ Works well under 12GB GPU constraints using compressed latent representation.
This is a base model and is intended to be fine-tuned for specific tasks or datasets. The model was trained on images with a resolution of 128×64 in RGB color channel, but this can be adjusted during fine-tuning to match your desired output size.
For best results, it is recommended to fine-tune the following three main components rather than just the core UNet model:
RobertaModel]AutoencoderKL]UNet2DConditionModel]import matplotlib.pyplot as plt
import torch
from diffusers import StableDiffusionPipeline
pipe = StableDiffusionPipeline.from_pretrained(
"channudam/stable-diffusion-khm-53",
torch_dtype=torch.float16,
).to("cuda")
images = pipe("បាត់ដំបង", guidance_scale=2).images[0]
plt.imshow(images)
plt.show()

Made with ❤️ by Channudam Ray | Factory.io & CADT, 2025