Downloads · 30 days
0
0% of all-time downloads
dtp-fine-tuning/multi-turn_chatbot_diploy
multi-turn_chatbot_diploy is a text generation model from dtp-fine-tuning. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
This model is a fine-tuned version of [aitfindonesia/Bakti-8B-Base] designed specifically for multi-turn conversational capabilities in the Indonesian language. It was trained using the Unsloth library for faster and…
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
3
Public
Parameters
8.2B
16.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.4 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of [aitfindonesia/Bakti-8B-Base] designed specifically for multi-turn conversational capabilities in the Indonesian language. It was trained using the Unsloth library for faster and memory-efficient training, utilizing LoRA (Low-Rank Adaptation).
The model is optimized to handle context retention across multiple turns of conversation, making it suitable for interview simulations, customer support, and general-purpose Indonesian assistants.
The model is designed for:
Dataset: dtp-fine-tuning/dtp-multiturn-interview-valid-15k
The model was fine-tuned using Unsloth on a single NVIDIA A100 (80GB) GPU. It utilizes 4-bit quantization (NF4) to reduce memory usage while maintaining performance via QLoRA.
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_projThe model demonstrates strong convergence on the multi-turn dataset.
Note: The model outperforms the standard Qwen3-8B baseline on this specific Indonesian dataset, achieving lower loss values faster.