Downloads · 30 days
23
18% of all-time downloads
MagistrTheOne/RadonDarkUltima
RadonDarkUltima is a text generation model from MagistrTheOne. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
RadonDarkUltima is an experimental 5TB parameter ultra-large scale Mistral-based transformer model designed for cutting-edge research and development. This model represents the pinnacle of the RADON ecosystem, pushing…
Downloads · 30 days
23
18% of all-time downloads
All-time downloads
125
Public
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json258 KB · 98%
From the Hugging Face model README
RadonDarkUltima is an experimental 5TB parameter ultra-large scale Mistral-based transformer model designed for cutting-edge research and development. This model represents the pinnacle of the RADON ecosystem, pushing the boundaries of what's possible with open-source language models.
This model is in experimental stage and requires massive computational resources. The framework is prepared but actual weights will be uploaded separately.
The model is split into 100 shards for efficient loading:
Each shard is approximately 50GB in size.
⚠️ Note: This repository contains only the model framework. Actual weights will be uploaded separately.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# Load model framework (weights not included)
model = AutoModelForCausalLM.from_pretrained(
"MagistrTheOne/RadonDarkUltima",
torch_dtype=torch.float16,
device_map="auto",
low_cpu_mem_usage=True
)
tokenizer = AutoTokenizer.from_pretrained("MagistrTheOne/RadonDarkUltima")
# Generate text (requires actual weights)
prompt = "Привет! Как дела?"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=100, temperature=0.7)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
RadonDarkUltima (5TB parameters)
├── Mistral Base Architecture
├── Llama 3 Innovations
│ ├── Grouped Query Attention (GQA) - 8:1 ratio
│ ├── RMSNorm Layer Normalization
│ ├── SwiGLU Activation
│ └── Rotary Position Embeddings (RoPE)
├── Flash Attention 2
├── Gradient Checkpointing
├── Sharded Weights (100 shards)
├── FP16 + INT8 Hybrid Quantization
└── Ultra-Large Scale Optimization
This experimental model is designed for:
MagistrTheOne - Creator and lead developer of RADON
Apache 2.0 License
@misc{radon-dark-ultima-2024,
title={RadonDarkUltima: 5TB Parameter Ultra-Large Scale Mistral-based Transformer},
author={MagistrTheOne},
year={2024},
url={https://huggingface.co/MagistrTheOne/RadonDarkUltima}
}
Created with ❤️ by MagistrTheOne
Pushing the boundaries of open-source AI! 🚀
This is an experimental research model requiring massive computational resources. Use responsibly and only for research purposes.