Downloads · 30 days
246
100% of all-time downloads
CH3NDev/Nool-Alpha-100M-Chat
Nool-Alpha-100M-Chat is a text generation model from CH3NDev. Use it when you need the model to write or continue text. The card lists the license as mit.
Nool-Alpha-100M-Chat is an efficient bilingual (Indonesian & English) and Python code causal language model based on an innovative hybrid architecture: - GSLA (Grouped-Subspace Latent Attention): Compresses KV-cache b…
Downloads · 30 days
246
100% of all-time downloads
All-time downloads
246
Public
Parameters
150M
599 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors599 MB · 99%
From the Hugging Face model README
Nool-Alpha-100M-Chat is an efficient bilingual (Indonesian & English) and Python code causal language model based on an innovative hybrid architecture:
Official Code Repository: GitHub - Ch3nOff/Nool-Alpha
nool_alphamodel.safetensors: Model weights in safe, zero-copy format.config.json: Complete architectural hyper-parameters.vocab.json, merges.txt, tokenizer.json, etc.import torch
from safetensors.torch import load_file
from nool_alpha.config import NoolAlphaConfig
from nool_alpha.model import NoolAlphaForCausalLM
# Load configuration and weights
config = NoolAlphaConfig.from_dict({
"architectures": [
"NoolAlphaForCausalLM"
],
"model_type": "nool_alpha",
"vocab_size": 50257,
"d_model": 768,
"n_layers": 10,
"num_heads": 12,
"head_dim": 64,
"d_c": 192,
"d_pe": 32,
"rope_theta": 500000.0,
"max_position_embeddings": 2048,
"sliding_window": 512,
"swa_interval": 4,
"shared_ffn_dim": 1536,
"num_experts": 8,
"top_k_experts": 2,
"expert_rank": 96,
"moe_aux_loss_coeff": 0.01,
"highway_alpha_init": 0.05,
"logit_soft_cap": 30.0,
"rms_norm_eps": 1e-06,
"tie_word_embeddings": true,
"checkpoint_step": 900,
"checkpoint_loss": 2.633549153804779
})
model = NoolAlphaForCausalLM(config)
state_dict = load_file("model.safetensors")
model.load_state_dict(state_dict)
model.eval()