Downloads ยท 30 days
23
23% of all-time downloads
Lyon28/caca-2M-untrained
caca-2M-untrained is a text generation model from Lyon28. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads ยท 30 days
23
23% of all-time downloads
All-time downloads
101
Public
Parameters
2M
8 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 MB ยท 98%
From the Hugging Face model README
2,001,216 parameters โข 2.00M โข 7 layers โข 512 tokens
๐ Documentation โข ๐ป Usage โข โ๏ธ Configuration โข ๐ฌ Architecture
</div>Status Model:
Widget di atas hanya menunjukkan format input yang diharapkan. Setelah model dilatih dengan dataset yang tepat, format yang sama akan menghasilkan output berkualitas tinggi.
| โ Bisa | โ Belum Bisa |
|---|---|
| Load model architecture | Generate teks bermakna |
| Test forward pass | Menjawab pertanyaan |
| Measure memory & speed | Reasoning & understanding |
| Start training | Production deployment |
| Fine-tuning experiments | Real-world applications |
Caca adalah arsitektur Large Language Model (LLM) generasi terbaru yang menggabungkan berbagai teknik state-of-the-art dalam deep learning. Model ini dirancang dengan fokus pada efisiensi komputasi, skalabilitas, dan performa tinggi.
<blockquote style="border-left: 4px solid #4A90E2; padding-left: 16px; margin: 16px 0; background: #f8f9fa; padding: 12px;"> <p><strong>๐ Tentang Project Caca</strong></p> <p><em>Caca</em> adalah eksperimen open-source Indonesian LLM yang dibuat dari nol secara individual dan bertahap. Bukan kompetitor siapa-siapa, cuma pengen eksplorasi apa yang bisa dilakukan dengan budget terbatas, passion unlimited, dan mindset collaborative.</p> <p>Kalau berguna buat orang lain, alhamdulillah. Kalau enggak, ya tetap fun kok. Ini proyek eksplorasi, jadi kalau gagal ya bagian dari proses belajar. Kalau berhasil, itu bonus.</p> <p>โ <strong>Lyon</strong>, Creator</p> </blockquote>| Fitur | Caca caca-2M | LLaMA-2 2.00M | GPT-3 2.00M |
|---|---|---|---|
| Attention Type | GQA | GQA | MHA |
| Position Encoding | RoPE + ALiBI | RoPE | Learned |
| Activation | SwiGLU | SwiGLU | GELU |
| Flash Attention | โ v2 | โ v1/v2 | โ |
| Long Context | Sliding Window + Sink | โ | Limited |
| MoE Support | โ Optional | โ | โ |
| Multimodal | โ Optional | โ | โ |
| Quantization | 4/8-bit | 4/8-bit | Limited |
๐ฌ Research & Development
๐ Academic & Education
๐ Base Model for Fine-tuning
๐ก Prototyping
โ Grouped Query Attention (GQA) - Efisiensi memori dan komputasi superior
โ Rotary Position Embeddings (RoPE) - Generalisasi konteks panjang lebih baik
โ RMSNorm - Normalisasi lebih stabil dan ~50% lebih cepat dari LayerNorm
โ SwiGLU Activation - Performa 10-15% lebih baik dari ReLU/GELU
โ Flash Attention 2 - Akselerasi hingga 3x dengan memory efficiency
๐ก Note: KV cache bertambah secara linear dengan panjang sequence. Untuk context 8K, kalikan nilai KV cache dengan 4.
CacaForCausalLM (2.00M)
โ
โโ Embedding: 4,000 ร 128
โ
โโ Transformer Layers (7x)
โ โโ RMSNorm
โ โโ Attention (GQA)
โ โ โโ Q: 4 heads ร 32 dim
โ โ โโ KV: 1 heads ร 32 dim
โ โ โโ RoPE (ฮธ=10,000)
โ โ โโ Flash Attention v2
โ โโ Residual
โ โโ RMSNorm
โ โโ FFN (SwiGLU)
โ โ โโ Gate: 128 โ 256
โ โ โโ Up: 128 โ 256
โ โ โโ Down: 256 โ 128
โ โโ Residual
โ
โโ Final RMSNorm
โโ LM Head: 128 โ 4,000
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
๐ PARAMETER BREAKDOWN:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Embeddings: 512,000 ( 25.6%)
Transformer Layers: 974,848 ( 48.7%)
โโ Attention: 286,720
โโ FFN: 688,128
Final Norm: 128 ( 0.0%)
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
TOTAL: 2,001,216 (100.0%)
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Key Design Decisions:
# Core dependencies (REQUIRED)
pip install torch>=2.0.0 transformers>=4.35.0 accelerate safetensors
# Optional: Untuk performa maksimal
pip install flash-attn --no-build-isolation # Flash Attention 2 (3x speedup)
pip install xformers # Memory efficient attention
pip install bitsandbytes # 4/8-bit quantization
# Optional: Untuk monitoring & profiling
pip install tensorboard wandb # Training monitoring
pip install gputil psutil # Resource monitoring
Compatibility Matrix:
| Component | Version | Note |
|---|---|---|
| Python | 3.8 - 3.11 | 3.11 recommended |
| PyTorch | โฅ 2.0.0 | 2.1+ untuk SDPA optimal |
| CUDA | 11.8 / 12.1 | Untuk Flash Attention |
| Transformers | โฅ 4.35.0 | Untuk AutoModel support |
from transformers import AutoConfig, AutoModelForCausalLM, AutoTokenizer
import torch
# Load configuration
config = AutoConfig.from_pretrained(
"Lyon28/caca-2M-untrained",
trust_remote_code=True
)
# Load model (FP16 untuk efisiensi)
model = AutoModelForCausalLM.from_pretrained(
"Lyon28/caca-2M-untrained",
config=config,
trust_remote_code=True,
torch_dtype=torch.float16,
device_map="auto" # Automatic device placement
)
# Model ini UNTRAINED - butuh training dulu!
print(f"Model loaded: {model.num_parameters():,} parameters")
print("โ ๏ธ Model ini belum dilatih dan belum bisa digunakan untuk inference")
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch
# 4-bit quantization config
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True
)
# Load model dengan quantization
model = AutoModelForCausalLM.from_pretrained(
"Lyon28/caca-2M-untrained",
trust_remote_code=True,
quantization_config=bnb_config,
device_map="auto"
)
print(f"Memory footprint: ~0.00GB (4-bit)")
from transformers import TrainingArguments, Trainer
# Training configuration
training_args = TrainingArguments(
output_dir="./output",
per_device_train_batch_size=1,
gradient_accumulation_steps=16,
learning_rate=2e-4,
max_steps=10000,
lr_scheduler_type="cosine",
warmup_steps=500,
logging_steps=10,
save_steps=500,
fp16=True, # Mixed precision
gradient_checkpointing=True, # Memory efficient
)
# Initialize trainer
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
)
# Start training
trainer.train()
model.gradient_checkpointing_enable()
print("โ
Gradient checkpointing enabled - saves ~40% memory")
from torch.optim import AdamW
from torch.cuda.amp import autocast, GradScaler
optimizer = AdamW(model.parameters(), lr=2e-4)
scaler = GradScaler()
for batch in dataloader:
# Mixed precision forward
with autocast(dtype=torch.bfloat16):
outputs = model(**batch)
loss = outputs.loss
# Backward with gradient scaling
scaler.scale(loss).backward()
scaler.step(optimizer)
scaler.update()
optimizer.zero_grad()
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel
# Initialize process group
dist.init_process_group(backend="nccl")
# Wrap model
model = DistributedDataParallel(
model,
device_ids=[local_rank],
find_unused_parameters=False
)
{
"architectures": ["CacaForCausalLM"],
"model_type": "caca",
"vocab_size": 4000,
"hidden_size": 128,
"intermediate_size": 256,
"num_hidden_layers": 7,
"num_attention_heads": 4,
"num_key_value_heads": 1,
"head_dim": 32,
"max_position_embeddings": 512,
"rope_theta": 10000,
"rms_norm_eps": 1e-06,
"use_cache": true,
"use_qk_norm": true,
"use_flash_attn": true,
"attention_dropout": 0.0,
"hidden_dropout": 0.1,
"torch_dtype": "float16"
}
from transformers import AutoConfig
# Load dan modifikasi config
config = AutoConfig.from_pretrained("Lyon28/caca-2M-untrained")
# Custom modifications
config.max_position_embeddings = 16384 # Extend context
config.rope_scaling = {"type": "linear", "factor": 2.0}
config.use_flash_attn = True
config.hidden_dropout = 0.05
# Save custom config
config.save_pretrained("./custom_config")
Input Tokens
โ
Embedding Layer (4,000 โ 128)
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Decoder Block ร 7 โ
โ โ
โ โโ RMSNorm โ
โ โโ Multi-Head Attention (GQA) โ
โ โ - Flash Attention v2 โ
โ โ - 4 Query heads, 1 KV heads โ
โ โ - RoPE position encoding โ
โ โโ Residual Connection โ
โ โ โ
โ โโ RMSNorm โ
โ โโ Feed-Forward Network (SwiGLU) โ
โ โ - Gate: 128 โ 256 โ
โ โ - Up: 128 โ 256 โ
โ โ - Down: 256 โ 128 โ
โ โโ Residual Connection โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
RMSNorm (Final)
โ
LM Head (128 โ 4,000)
โ
Output Logits
Query: [4 heads ร 32 dim] = 128
Key: [1 heads ร 32 dim] = 32
Value: [1 heads ร 32 dim] = 32
Grouped Query Attention:
- Setiap 4 query heads berbagi 1 KV head
- Memory KV cache: 75% lebih kecil dari Multi-Head Attention
- Kualitas mendekati MHA, speed mendekati MQA
FFN(x) = (SiLU(xW_gate) โ xW_up) W_down
Where:
- W_gate: 128 ร 256
- W_up: 128 ร 256
- W_down: 256 ร 128
- SiLU(x) = x ยท sigmoid(x)
- โ = element-wise multiplication
Model mendukung format chat standar untuk conversational AI:
# Format chat template bawaan
chat_template = """
{% for message in messages %}
{% if message['role'] == 'system' %}
System: {{ message['content'] }}
{% elif message['role'] == 'user' %}
User: {{ message['content'] }}
{% elif message['role'] == 'assistant' %}
Assistant: {{ message['content'] }}
{% endif %}
{% endfor %}
{% if add_generation_prompt %}Assistant:{% endif %}
"""
# Contoh penggunaan
messages = [
{"role": "system", "content": "Kamu adalah asisten AI yang membantu dan ramah."},
{"role": "user", "content": "Jelaskan tentang fotosintesis"},
{"role": "assistant", "content": "Fotosintesis adalah proses di mana tumbuhan mengubah cahaya matahari menjadi energi kimia..."},
{"role": "user", "content": "Apa manfaatnya bagi manusia?"},
]
# Apply template
formatted = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
print(formatted)
# Output:
# System: Kamu adalah asisten AI yang membantu dan ramah.
#
# User: Jelaskan tentang fotosintesis
# Assistant: Fotosintesis adalah proses di mana tumbuhan...
# User: Apa manfaatnya bagi manusia?
# Assistant:
Model ini dirancang untuk berbagai aplikasi NLP setelah melalui proses training:
โ ๏ธ Model belum melalui evaluasi karena status untrained
Setelah training, model akan dievaluasi pada:
# Rule of thumb untuk 2.00M model
# GPU Memory โ Batch size per device
if gpu_memory >= 80: # A100 80GB
batch_size = 7995
gradient_accumulation = 1
elif gpu_memory >= 40: # A100 40GB
batch_size = 3997
gradient_accumulation = 1
elif gpu_memory >= 24: # RTX 3090/4090
batch_size = 1
gradient_accumulation = 1
# Effective batch size = batch_size ร gradient_accumulation ร num_gpus
# Recommended untuk 2.00M model
learning_rate = 0.0005 # Base LR
warmup_ratio = 0.05 # 5% of total steps
lr_scheduler = "cosine" # atau "linear"
# Learning rate scaling rule:
# LR โ sqrt(batch_size)
# Untuk batch size 256: LR = 0.0005
# Untuk batch size 512: LR = 7.07e-04
# Prevent gradient explosion
max_grad_norm = 1.0 # Clip at 1.0
# Monitor gradients
from torch.nn.utils import clip_grad_norm_
grad_norm = clip_grad_norm_(model.parameters(), max_grad_norm)
if grad_norm > 10.0:
print(f"โ ๏ธ High gradient norm: {grad_norm:.2f}")
# Tips untuk stable training:
1. **Warmup**: Mulai dengan LR rendah
2. **Gradient Checkpointing**: Kurangi memory footprint
3. **Mixed Precision**: Gunakan BF16 jika tersedia (lebih stable dari FP16)
4. **Batch Size**: Start small, increase gradually
5. **Monitor**: Track loss, perplexity, gradient norms
# Solusi OOM saat training:
โ
1. Enable gradient checkpointing
model.gradient_checkpointing_enable()
โ
2. Reduce batch size
per_device_train_batch_size = 1
โ
3. Increase gradient accumulation
gradient_accumulation_steps = 32
โ
4. Use quantization
load_in_8bit = True # atau load_in_4bit
โ
5. Reduce sequence length
max_length = 512 # Start dengan ini
โ
6. CPU offloading (jika perlu)
device_map = "auto"
offload_folder = "offload"
# Optimasi kecepatan training:
โ
1. Flash Attention
config.use_flash_attn = True # 2-3x speedup
โ
2. Compile model (PyTorch 2.0+)
model = torch.compile(model, mode="reduce-overhead")
โ
3. DataLoader optimization
dataloader = DataLoader(
dataset,
batch_size=batch_size,
num_workers=4, # Parallel data loading
pin_memory=True, # Faster GPU transfer
prefetch_factor=2
)
โ
4. Mixed precision
use_fp16 = True # atau bf16
โ
5. Optimize communication (multi-GPU)
find_unused_parameters = False
gradient_as_bucket_view = True
# Jika loss menjadi NaN:
โ
1. Reduce learning rate
learning_rate = learning_rate * 0.1
โ
2. Check gradient norms
clip_grad_norm_(model.parameters(), 1.0)
โ
3. Use BF16 instead of FP16
torch_dtype = torch.bfloat16 # Lebih stable
โ
4. Add epsilon to RMSNorm
rms_norm_eps = 1e-5 # Increase jika perlu
โ
5. Check data
# Pastikan tidak ada inf/nan di dataset
assert not torch.isnan(input_ids).any()
assert not torch.isinf(attention_mask).any()
Model ini TIDAK BOLEH digunakan untuk:
Violation consequences: Model access revocation + legal action jika applicable
</div>Model ini dirilis di bawah Apache License 2.0
โ Anda BEBAS untuk:
โ ๏ธ Dengan syarat:
โ Tanpa jaminan apapun (use at your own risk)
</div>Full license text: Apache-2.0
Jika Anda menggunakan model ini dalam penelitian, mohon sitasi:
@misc{cacacaca2m,
author = {Lyon},
title = {Caca-caca-2M: Modern Transformer Architecture with Grouped Query Attention},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face Model Hub},
howpublished = {\url{https://huggingface.co/Lyon28/caca-2M-untrained}},
note = {Untrained model with 2,001,216 parameters}
}
APA Style:
Lyon. (2026). Caca-caca-2M: Modern Transformer Architecture with Grouped
Query Attention [Untrained model]. Hugging Face.
https://huggingface.co/Lyon28/caca-2M-untrained
MLA Style:
Lyon. "Caca-caca-2M: Modern Transformer Architecture with Grouped Query Attention."
Hugging Face, 2026, huggingface.co/Lyon28/caca-2M-untrained.
Model ini berdiri di pundak para raksasa! Terima kasih kepada:
<details> <summary><b>๐๏ธ Klik untuk daftar lengkap acknowledgments</b></summary>Special thanks to Indonesian NLP researchers & practitioners yang telah membangun foundation untuk Indonesian language AI.
</details>Model ini dirilis di bawah Apache License 2.0.
Lihat LICENSE untuk detail lengkap.
Kami sangat terbuka untuk kontribusi! Berikut cara Anda bisa berkontribusi:
Terima kasih kepada komunitas open-source yang telah berkontribusi pada:
Jika model ini berguna, jangan lupa โญ repository kami!
<div align="center"> <table> <tr> <td align="center">โญ<br/><b>Star Repo</b><br/><sub>Show your support</sub></td> <td align="center">๐<br/><b>Share</b><br/><sub>Tell your friends</sub></td> <td align="center">๐ฌ<br/><b>Join Discussion</b><br/><sub>Ask questions</sub></td> <td align="center">๐ค<br/><b>Contribute</b><br/><sub>Make it better</sub></td> </tr> </table>Model ini menunggu untuk dilatih dan menjadi foundation untuk aplikasi AI Anda.
๐ฅ Download Model โข ๐ Read Docs โข ๐ฌ Join Community
</div>| Metric | Value |
|---|---|
| ๐ Total Parameters | 2,001,216 |
| ๐๏ธ Layers | 7 |
| ๐ฏ Attention Heads | 4 |
| ๐ Max Context | 512 tokens |
| ๐พ Size (FP16) | 0.00 GB |
| ๐พ Size (INT4) | 0.00 GB |
<br/><br/>
๐ "Dari nol, untuk semua" ๐
<sub>Last updated: january 2026</sub>
</div>