Downloads · 30 days
23
23% of all-time downloads
Lyon28/caca-2M-untrained
caca-2M-untrained is a text generation model from Lyon28. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
23
23% of all-time downloads
All-time downloads
101
Public
Parameters
2M
8 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 MB · 98%
From the Hugging Face model README
2,001,216 parameters • 2.00M • 7 layers • 512 tokens
📚 Documentation • 💻 Usage • ⚙️ Configuration • 🔬 Architecture
</div>Status Model:
Widget di atas hanya menunjukkan format input yang diharapkan. Setelah model dilatih dengan dataset yang tepat, format yang sama akan menghasilkan output berkualitas tinggi.
| ✅ Bisa | ❌ Belum Bisa |
|---|---|
| Load model architecture | Generate teks bermakna |
| Test forward pass | Menjawab pertanyaan |
| Measure memory & speed | Reasoning & understanding |
| Start training | Production deployment |
| Fine-tuning experiments | Real-world applications |
Caca adalah arsitektur Large Language Model (LLM) generasi terbaru yang menggabungkan berbagai teknik state-of-the-art dalam deep learning. Model ini dirancang dengan fokus pada efisiensi komputasi, skalabilitas, dan performa tinggi.
<blockquote style="border-left: 4px solid #4A90E2; padding-left: 16px; margin: 16px 0; background: #f8f9fa; padding: 12px;"> <p><strong>📖 Tentang Project Caca</strong></p> <p><em>Caca</em> adalah eksperimen open-source Indonesian LLM yang dibuat dari nol secara individual dan bertahap. Bukan kompetitor siapa-siapa, cuma pengen eksplorasi apa yang bisa dilakukan dengan budget terbatas, passion unlimited, dan mindset collaborative.</p> <p>Kalau berguna buat orang lain, alhamdulillah. Kalau enggak, ya tetap fun kok. Ini proyek eksplorasi, jadi kalau gagal ya bagian dari proses belajar. Kalau berhasil, itu bonus.</p> <p>— <strong>Lyon</strong>, Creator</p> </blockquote>| Fitur | Caca caca-2M | LLaMA-2 2.00M | GPT-3 2.00M |
|---|---|---|---|
| Attention Type | GQA | GQA | MHA |
| Position Encoding | RoPE + ALiBI | RoPE | Learned |
| Activation | SwiGLU | SwiGLU | GELU |
| Flash Attention | ✅ v2 | ✅ v1/v2 | ❌ |
| Long Context | Sliding Window + Sink | ✅ | Limited |
| MoE Support | ✅ Optional | ❌ | ❌ |
| Multimodal | ✅ Optional | ❌ | ❌ |
| Quantization | 4/8-bit | 4/8-bit | Limited |
🔬 Research & Development
📚 Academic & Education
🚀 Base Model for Fine-tuning
💡 Prototyping
✅ Grouped Query Attention (GQA) - Efisiensi memori dan komputasi superior
✅ Rotary Position Embeddings (RoPE) - Generalisasi konteks panjang lebih baik
✅ RMSNorm - Normalisasi lebih stabil dan ~50% lebih cepat dari LayerNorm
✅ SwiGLU Activation - Performa 10-15% lebih baik dari ReLU/GELU
✅ Flash Attention 2 - Akselerasi hingga 3x dengan memory efficiency
💡 Note: KV cache bertambah secara linear dengan panjang sequence. Untuk context 8K, kalikan nilai KV cache dengan 4.
CacaForCausalLM (2.00M)
│
├─ Embedding: 4,000 × 128
│
├─ Transformer Layers (7x)
│ ├─ RMSNorm
│ ├─ Attention (GQA)
│ │ ├─ Q: 4 heads × 32 dim
│ │ ├─ KV: 1 heads × 32 dim
│ │ ├─ RoPE (θ=10,000)
│ │ └─ Flash Attention v2
│ ├─ Residual
│ ├─ RMSNorm
│ ├─ FFN (SwiGLU)
│ │ ├─ Gate: 128 → 256
│ │ ├─ Up: 128 → 256
│ │ └─ Down: 256 → 128
│ └─ Residual
│
├─ Final RMSNorm
└─ LM Head: 128 → 4,000
═══════════════════════════════════════════════════════════
📊 PARAMETER BREAKDOWN:
═══════════════════════════════════════════════════════════
Embeddings: 512,000 ( 25.6%)
Transformer Layers: 974,848 ( 48.7%)
├─ Attention: 286,720
└─ FFN: 688,128
Final Norm: 128 ( 0.0%)
───────────────────────────────────────────────────────────
TOTAL: 2,001,216 (100.0%)
═══════════════════════════════════════════════════════════
Key Design Decisions:
# Core dependencies (REQUIRED)
pip install torch>=2.0.0 transformers>=4.35.0 accelerate safetensors
# Optional: Untuk performa maksimal
pip install flash-attn --no-build-isolation # Flash Attention 2 (3x speedup)
pip install xformers # Memory efficient attention
pip install bitsandbytes # 4/8-bit quantization
# Optional: Untuk monitoring & profiling
pip install tensorboard wandb # Training monitoring
pip install gputil psutil # Resource monitoring
Compatibility Matrix:
| Component | Version | Note |
|---|---|---|
| Python | 3.8 - 3.11 | 3.11 recommended |
| PyTorch | ≥ 2.0.0 | 2.1+ untuk SDPA optimal |
| CUDA | 11.8 / 12.1 | Untuk Flash Attention |
| Transformers | ≥ 4.35.0 | Untuk AutoModel support |
from transformers import AutoConfig, AutoModelForCausalLM, AutoTokenizer
import torch
# Load configuration
config = AutoConfig.from_pretrained(
"Lyon28/caca-2M-untrained",
trust_remote_code=True
)
# Load model (FP16 untuk efisiensi)
model = AutoModelForCausalLM.from_pretrained(
"Lyon28/caca-2M-untrained",
config=config,
trust_remote_code=True,
torch_dtype=torch.float16,
device_map="auto" # Automatic device placement
)
# Model ini UNTRAINED - butuh training dulu!
print(f"Model loaded: {model.num_parameters():,} parameters")
print("⚠️ Model ini belum dilatih dan belum bisa digunakan untuk inference")
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch
# 4-bit quantization config
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True
)
# Load model dengan quantization
model = AutoModelForCausalLM.from_pretrained(
"Lyon28/caca-2M-untrained",
trust_remote_code=True,
quantization_config=bnb_config,
device_map="auto"
)
print(f"Memory footprint: ~0.00GB (4-bit)")
from transformers import TrainingArguments, Trainer
# Training configuration
training_args = TrainingArguments(
output_dir="./output",
per_device_train_batch_size=1,
gradient_accumulation_steps=16,
learning_rate=2e-4,
max_steps=10000,
lr_scheduler_type="cosine",
warmup_steps=500,
logging_steps=10,
save_steps=500,
fp16=True, # Mixed precision
gradient_checkpointing=True, # Memory efficient
)
# Initialize trainer
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
)
# Start training
trainer.train()
model.gradient_checkpointing_enable()
print("✅ Gradient checkpointing enabled - saves ~40% memory")
from torch.optim import AdamW
from torch.cuda.amp import autocast, GradScaler
optimizer = AdamW(model.parameters(), lr=2e-4)
scaler = GradScaler()
for batch in dataloader:
# Mixed precision forward
with autocast(dtype=torch.bfloat16):
outputs = model(**batch)
loss = outputs.loss
# Backward with gradient scaling
scaler.scale(loss).backward()
scaler.step(optimizer)
scaler.update()
optimizer.zero_grad()
import torch.distributed as dist
from torch.nn.parallel import DistributedDataParallel
# Initialize process group
dist.init_process_group(backend="nccl")
# Wrap model
model = DistributedDataParallel(
model,
device_ids=[local_rank],
find_unused_parameters=False
)
{
"architectures": ["CacaForCausalLM"],
"model_type": "caca",
"vocab_size": 4000,
"hidden_size": 128,
"intermediate_size": 256,
"num_hidden_layers": 7,
"num_attention_heads": 4,
"num_key_value_heads": 1,
"head_dim": 32,
"max_position_embeddings": 512,
"rope_theta": 10000,
"rms_norm_eps": 1e-06,
"use_cache": true,
"use_qk_norm": true,
"use_flash_attn": true,
"attention_dropout": 0.0,
"hidden_dropout": 0.1,
"torch_dtype": "float16"
}
from transformers import AutoConfig
# Load dan modifikasi config
config = AutoConfig.from_pretrained("Lyon28/caca-2M-untrained")
# Custom modifications
config.max_position_embeddings = 16384 # Extend context
config.rope_scaling = {"type": "linear", "factor": 2.0}
config.use_flash_attn = True
config.hidden_dropout = 0.05
# Save custom config
config.save_pretrained("./custom_config")
Input Tokens
↓
Embedding Layer (4,000 → 128)
↓
┌─────────────────────────────────────┐
│ Decoder Block × 7 │
│ │
│ ┌─ RMSNorm │
│ ├─ Multi-Head Attention (GQA) │
│ │ - Flash Attention v2 │
│ │ - 4 Query heads, 1 KV heads │
│ │ - RoPE position encoding │
│ ├─ Residual Connection │
│ │ │
│ ├─ RMSNorm │
│ ├─ Feed-Forward Network (SwiGLU) │
│ │ - Gate: 128 → 256 │
│ │ - Up: 128 → 256 │
│ │ - Down: 256 → 128 │
│ └─ Residual Connection │
│ │
└─────────────────────────────────────┘
↓
RMSNorm (Final)
↓
LM Head (128 → 4,000)
↓
Output Logits
Query: [4 heads × 32 dim] = 128
Key: [1 heads × 32 dim] = 32
Value: [1 heads × 32 dim] = 32
Grouped Query Attention:
- Setiap 4 query heads berbagi 1 KV head
- Memory KV cache: 75% lebih kecil dari Multi-Head Attention
- Kualitas mendekati MHA, speed mendekati MQA
FFN(x) = (SiLU(xW_gate) ⊙ xW_up) W_down
Where:
- W_gate: 128 × 256
- W_up: 128 × 256
- W_down: 256 × 128
- SiLU(x) = x · sigmoid(x)
- ⊙ = element-wise multiplication
Model mendukung format chat standar untuk conversational AI:
# Format chat template bawaan
chat_template = """
{% for message in messages %}
{% if message['role'] == 'system' %}
System: {{ message['content'] }}
{% elif message['role'] == 'user' %}
User: {{ message['content'] }}
{% elif message['role'] == 'assistant' %}
Assistant: {{ message['content'] }}
{% endif %}
{% endfor %}
{% if add_generation_prompt %}Assistant:{% endif %}
"""
# Contoh penggunaan
messages = [
{"role": "system", "content": "Kamu adalah asisten AI yang membantu dan ramah."},
{"role": "user", "content": "Jelaskan tentang fotosintesis"},
{"role": "assistant", "content": "Fotosintesis adalah proses di mana tumbuhan mengubah cahaya matahari menjadi energi kimia..."},
{"role": "user", "content": "Apa manfaatnya bagi manusia?"},
]
# Apply template
formatted = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
print(formatted)
# Output:
# System: Kamu adalah asisten AI yang membantu dan ramah.
#
# User: Jelaskan tentang fotosintesis
# Assistant: Fotosintesis adalah proses di mana tumbuhan...
# User: Apa manfaatnya bagi manusia?
# Assistant:
Model ini dirancang untuk berbagai aplikasi NLP setelah melalui proses training:
⚠️ Model belum melalui evaluasi karena status untrained
Setelah training, model akan dievaluasi pada:
# Rule of thumb untuk 2.00M model
# GPU Memory → Batch size per device
if gpu_memory >= 80: # A100 80GB
batch_size = 7995
gradient_accumulation = 1
elif gpu_memory >= 40: # A100 40GB
batch_size = 3997
gradient_accumulation = 1
elif gpu_memory >= 24: # RTX 3090/4090
batch_size = 1
gradient_accumulation = 1
# Effective batch size = batch_size × gradient_accumulation × num_gpus
# Recommended untuk 2.00M model
learning_rate = 0.0005 # Base LR
warmup_ratio = 0.05 # 5% of total steps
lr_scheduler = "cosine" # atau "linear"
# Learning rate scaling rule:
# LR ∝ sqrt(batch_size)
# Untuk batch size 256: LR = 0.0005
# Untuk batch size 512: LR = 7.07e-04
# Prevent gradient explosion
max_grad_norm = 1.0 # Clip at 1.0
# Monitor gradients
from torch.nn.utils import clip_grad_norm_
grad_norm = clip_grad_norm_(model.parameters(), max_grad_norm)
if grad_norm > 10.0:
print(f"⚠️ High gradient norm: {grad_norm:.2f}")
# Tips untuk stable training:
1. **Warmup**: Mulai dengan LR rendah
2. **Gradient Checkpointing**: Kurangi memory footprint
3. **Mixed Precision**: Gunakan BF16 jika tersedia (lebih stable dari FP16)
4. **Batch Size**: Start small, increase gradually
5. **Monitor**: Track loss, perplexity, gradient norms
# Solusi OOM saat training:
✅ 1. Enable gradient checkpointing
model.gradient_checkpointing_enable()
✅ 2. Reduce batch size
per_device_train_batch_size = 1
✅ 3. Increase gradient accumulation
gradient_accumulation_steps = 32
✅ 4. Use quantization
load_in_8bit = True # atau load_in_4bit
✅ 5. Reduce sequence length
max_length = 512 # Start dengan ini
✅ 6. CPU offloading (jika perlu)
device_map = "auto"
offload_folder = "offload"
# Optimasi kecepatan training:
✅ 1. Flash Attention
config.use_flash_attn = True # 2-3x speedup
✅ 2. Compile model (PyTorch 2.0+)
model = torch.compile(model, mode="reduce-overhead")
✅ 3. DataLoader optimization
dataloader = DataLoader(
dataset,
batch_size=batch_size,
num_workers=4, # Parallel data loading
pin_memory=True, # Faster GPU transfer
prefetch_factor=2
)
✅ 4. Mixed precision
use_fp16 = True # atau bf16
✅ 5. Optimize communication (multi-GPU)
find_unused_parameters = False
gradient_as_bucket_view = True
# Jika loss menjadi NaN:
✅ 1. Reduce learning rate
learning_rate = learning_rate * 0.1
✅ 2. Check gradient norms
clip_grad_norm_(model.parameters(), 1.0)
✅ 3. Use BF16 instead of FP16
torch_dtype = torch.bfloat16 # Lebih stable
✅ 4. Add epsilon to RMSNorm
rms_norm_eps = 1e-5 # Increase jika perlu
✅ 5. Check data
# Pastikan tidak ada inf/nan di dataset
assert not torch.isnan(input_ids).any()
assert not torch.isinf(attention_mask).any()
Model ini TIDAK BOLEH digunakan untuk:
Violation consequences: Model access revocation + legal action jika applicable
</div>Model ini dirilis di bawah Apache License 2.0
✅ Anda BEBAS untuk:
⚠️ Dengan syarat:
❌ Tanpa jaminan apapun (use at your own risk)
</div>Full license text: Apache-2.0
Jika Anda menggunakan model ini dalam penelitian, mohon sitasi:
@misc{cacacaca2m,
author = {Lyon},
title = {Caca-caca-2M: Modern Transformer Architecture with Grouped Query Attention},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face Model Hub},
howpublished = {\url{https://huggingface.co/Lyon28/caca-2M-untrained}},
note = {Untrained model with 2,001,216 parameters}
}
APA Style:
Lyon. (2026). Caca-caca-2M: Modern Transformer Architecture with Grouped
Query Attention [Untrained model]. Hugging Face.
https://huggingface.co/Lyon28/caca-2M-untrained
MLA Style:
Lyon. "Caca-caca-2M: Modern Transformer Architecture with Grouped Query Attention."
Hugging Face, 2026, huggingface.co/Lyon28/caca-2M-untrained.
Model ini berdiri di pundak para raksasa! Terima kasih kepada:
<details> <summary><b>🏛️ Klik untuk daftar lengkap acknowledgments</b></summary>Special thanks to Indonesian NLP researchers & practitioners yang telah membangun foundation untuk Indonesian language AI.
</details>Model ini dirilis di bawah Apache License 2.0.
Lihat LICENSE untuk detail lengkap.
Kami sangat terbuka untuk kontribusi! Berikut cara Anda bisa berkontribusi:
Terima kasih kepada komunitas open-source yang telah berkontribusi pada:
Jika model ini berguna, jangan lupa ⭐ repository kami!
<div align="center"> <table> <tr> <td align="center">⭐<br/><b>Star Repo</b><br/><sub>Show your support</sub></td> <td align="center">🔗<br/><b>Share</b><br/><sub>Tell your friends</sub></td> <td align="center">💬<br/><b>Join Discussion</b><br/><sub>Ask questions</sub></td> <td align="center">🤝<br/><b>Contribute</b><br/><sub>Make it better</sub></td> </tr> </table>Model ini menunggu untuk dilatih dan menjadi foundation untuk aplikasi AI Anda.
📥 Download Model • 📖 Read Docs • 💬 Join Community
</div>| Metric | Value |
|---|---|
| 💎 Total Parameters | 2,001,216 |
| 🏗️ Layers | 7 |
| 🎯 Attention Heads | 4 |
| 📖 Max Context | 512 tokens |
| 💾 Size (FP16) | 0.00 GB |
| 💾 Size (INT4) | 0.00 GB |
<br/><br/>
🌟 "Dari nol, untuk semua" 🌟
<sub>Last updated: january 2026</sub>
</div>