Downloads · 30 days
69
7% of all-time downloads
mediusware-ai/intellix
intellix is a text generation model from mediusware-ai. Use it when you need the model to write or continue text. It is set up for peft.
<p align="center" <img src="https://huggingface.co/mediusware-ai/intellix/resolve/main/logo.webp" width="800" </p
Downloads · 30 days
69
7% of all-time downloads
All-time downloads
1K
Public
Repo size
9.4 GB
Likes
4
Public
Click a slice to open those files.
.gguf1.6 GB · 99%
From the Hugging Face model README
Intellix is a high-capacity, fine-tuned large language model (LLM) designed specifically for enterprise-grade applications.
Evaluations were conducted using a proprietary enterprise benchmark suite and real-world business scenarios to ensure the model's readiness for B2B deployment.
Tested on Q8_0 GGUF via optimized local inference.
| Metric | Performance Value |
|---|---|
| Average Throughput | 196.08 tokens/sec |
| Average Latency | 0.68 seconds |
| Peak Throughput | 199.48 tokens/sec |
| Model Footprint | 2.0 GB |
The model was fine-tuned on a massive, curated dataset including:
Data was rigorously cleaned to remove PII (Personally Identifiable Information) and informal/low-quality text, ensuring the model's output remains strictly professional.
The following scenarios were used to validate the model's business intelligence:
mw-intellix was fine-tuned using the Unsloth library for memory-efficient and fast training. The process utilized LoRA (Low-Rank Adaptation) to adapt the base architecture to specialized business domains without compromising the model's general intelligence.
The following hyperparameters were used during the fine-tuning phase:
| Parameter | Value |
|---|---|
| PEFT Type | LoRA |
| LoRA Rank (r) | 16 |
| LoRA Alpha | 16 |
| LoRA Dropout | 0.0 |
| Target Modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Precision | bfloat16 |
| Optimizer | AdamW |
| Learning Rate | 2e-4 |
| Epochs | 3 |
Intellix is highly optimized for local execution using Ollama.
Modelfile in this repository which includes the correct repeat_penalty (1.5) and stop tokens to prevent loops.ollama create intellix -f Modelfile
ollama run intellix
Model Parameters for Stability:
repeat_penalty: 1.5temperature: 0.7stop: ["<|im_start|>", "<|im_end|>", "User:", "Assistant:"]For research or programmatic access, use the transformers library.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "mediusware-ai/intellix"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
# Using the ChatML Template
messages = [
{"role": "system", "content": "You are Intellix, a professional AI assistant developed by Mediusware."},
{"role": "user", "content": "Tell me about Mediusware's US presence."}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, repetition_penalty=1.5)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Designed for Local-First Deployment. When used via Ollama or GGUF, business data never leaves the local infrastructure, ensuring 100% data residency and privacy.
For custom enterprise deployments or inquiries, visit mediusware.com.