Downloads · 30 days
26
2% of all-time downloads
jcordon5/Mistral-7B-cybersecurity-rules
Mistral-7B-cybersecurity-rules is a text generation model from jcordon5. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This model is a fine-tune of Mistral-7B-Instruct-v0.2, via Knowledge Distillation of 0dAI-7.5B. The fine-tuning was conducted using a curated corpus of 950 cybersecurity rules from SIGMA, YARA, and Suricata repositori…
Downloads · 30 days
26
2% of all-time downloads
All-time downloads
1.2K
Public
Parameters
7.2B
14.5 GB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors14.5 GB · 100%
From the Hugging Face model README
This model is a fine-tune of Mistral-7B-Instruct-v0.2, via Knowledge Distillation of 0dAI-7.5B. The fine-tuning was conducted using a curated corpus of 950 cybersecurity rules from SIGMA, YARA, and Suricata repositories for threat and intrusion detection.
Instruct the model to craft a SIGMA rule for detecting potentially malicious commands such as msfvenom and netcat in Audit system logs, or a Suricata rule to spot SSH brute-force attacks, or even a YARA rule to identify obfuscated strings in files — and watch the magic happen! Automate the creation of rules in your cybersecurity systems with this model.
For an in-depth understanding of how this model has been fine-tuned, refer to the associated paper here: [available soon].
You can easily quantize your model for local use on your computer with the help of the llama.cpp or ollama libraries. This process converts your model into a format that is optimized for performance, particularly useful for deployment on devices with limited computational resources.
To perform this quantization using the llama.cpp library (link to llama.cpp), follow the steps below:
First, convert your model's vocabulary to a format suitable for quantization. Use the following command, replacing /path/to/ with the actual path to your model files:
python convert.py /path/to/Mistral-7B-cybersecurity-rules \
--vocab-only \
--outfile /path/to/Mistral-7B-cybersecurity-rules/tokenizer.model \
--vocab-type bpe
This command extracts and converts the vocabulary using the byte pair encoding (BPE) method, saving it to a new file.
Next, prepare the model for quantization by converting it to a half-precision floating-point format (FP16). This step reduces the model size and prepares it for the final quantization to 8-bit integers. Execute the following command:
python convert.py \
--outtype f16 \
--vocab-type bpe \ # Add this line only if you encounter issues with the vocabulary type
--outfile /path/to/Mistral-7B-cybersecurity-rules/ggml-model-f16.gguf
This command outputs a file that has been converted to FP16, which is an intermediary step before applying 8-bit quantization.
Finally, apply 8-bit quantization to the FP16 model file. This step significantly reduces the model's memory footprint, making it suitable for deployment in resource-constrained environments:
quantize /path/to/Mistral-7B-cybersecurity-rules/ggml-model-f16.gguf \
/path/to/Mistral-7B-cybersecurity-rules/mistral-7b-rules-q8_0.gguf \
q8_0
Here, the quantize command converts the FP16 model into an 8-bit quantized model, further compressing the model while retaining its capability to perform its tasks effectively.
This repository is licensed under the Apache License, Version 2.0. You can obtain a copy of the license at Apache License 2.0.
This software is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
This model has been fine-tuned based on the original Mistral-7B-Instruct-v0.2. Significant modifications were made to train it on a cybersecurity corpus for threat and intrusion detection.