Downloads · 30 days
23
19% of all-time downloads
david-ar/irc-mistral-24b
irc-mistral-24b is a text generation model from david-ar. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<img src="https://raw.githubusercontent.com/axolotl-ai-cloud/axolotl/main/image/axolotl-badge-web.png" alt="Built with Axolotl" width="200" height="32"/ <details<summarySee axolotl config</summary
Downloads · 30 days
23
19% of all-time downloads
All-time downloads
122
Public
Parameters
23.6B
47.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors47.1 GB · 100%
From the Hugging Face model README
axolotl version: 0.8.0.dev0
# Base model configuration
base_model: mistralai/Mistral-Small-24B-Base-2501
model_type: MistralForCausalLM
tokenizer_type: AutoTokenizer
trust_remote_code: true
tokenizer_use_fast: true
# Device mapping for multi-GPU
device_map: "balanced"
# Memory settings
load_in_4bit: true
load_in_8bit: false
bf16: true
low_cpu_mem_usage: true
# Advanced optimizations
flash_attention: true
gradient_checkpointing: true
# Dataset configuration
datasets:
- path: david-ar/synthetic-irc-data
type: completion
# Output directory
output_dir: ./outputs/public-irc-mistral-24b
val_set_size: 0.05 # 75 conversations for validation
dataset_prepared_path: last_run_prepared
# Sequence settings
sequence_len: 4096
sample_packing: true
pad_to_sequence_len: true
train_on_inputs: true
eval_sample_packing: false
# LoRA configuration
adapter: lora
lora_r: 128
lora_alpha: 256
lora_dropout: 0.1
lora_target_modules:
- q_proj
- v_proj
- k_proj
- o_proj
- gate_proj
- down_proj
- up_proj
# Training hyperparameters - adjusted for smaller dataset
micro_batch_size: 1
gradient_accumulation_steps: 16
num_epochs: 4 # Increased from 2, but with careful monitoring
optimizer: adamw_torch
lr_scheduler: cosine
learning_rate: 0.00008 # Same conservative LR
weight_decay: 0.01
warmup_ratio: 0.05
# Performance monitoring
group_by_length: true
shuffle_merged_datasets: true
include_tokens_per_second: true
# Weights & Biases - public project
wandb_project: public-irc-mistral-24b
wandb_entity: davidar
wandb_name: synthetic-irc-data
wandb_log_model: "false"
# Mistral model configuration
is_mistral_derived_model: true
# Early stopping
load_best_model_at_end: true
metric_for_best_model: "loss"
greater_is_better: false
</details><br>
This model is a fine-tuned version of mistralai/Mistral-Small-24B-Base-2501 on the david-ar/synthetic-irc-data dataset, creating a model that generates natural IRC/Discord-style conversations.
This model was trained to replicate authentic IRC (Internet Relay Chat) conversational dynamics, moving away from the typical AI assistant pattern toward more natural, community-style interactions. The model learns from synthetic conversations featuring multiple participants including "Em", an AI character who participates as a community member rather than an assistant.
<username> message contentThe following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 0.9145 | 0.9746 | 24 | 0.9128 |
| 0.6565 | 1.9746 | 48 | 0.8936 |
| 0.4671 | 2.9746 | 72 | 0.9503 |
| 0.3594 | 3.9746 | 96 | 0.9871 |
Note: Best checkpoint at step 48 (lowest validation loss) was used for final model.