Downloads · 30 days
0
tuandunghcmut/cybermetric_processor
cybermetric_processor is a machine learning model from tuandunghcmut. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository contains the processing script used to convert the original CyberMetric JSON files into well-structured Hugging Face datasets with comprehensive documentation.
Downloads · 30 days
0
Access
Public
Updated Sep 27, 2025
Repo size
—
Likes
1
Public
Click a slice to open those files.
.py11.8 KB · 56%
From the Hugging Face model README
This repository contains the processing script used to convert the original CyberMetric JSON files into well-structured Hugging Face datasets with comprehensive documentation.
The script processes 4 different CyberMetric JSON files and uploads them as separate, documented datasets:
All processed datasets are available at: tuandunghcmut
Cryptography:
Q: What is the primary requirement for an Random Bit Generator's (RBG) output to be used for generating cryptographic keys?
A) Length matching target data
B) Computationally indistinguishable from random bits with sufficient entropy ✓
C) Maximum possible length
D) Exact symmetric key length
PCI DSS:
Q: What is the primary purpose of segmentation in the context of PCI DSS?
A) Reduce applicable requirements
B) Limit assessment scope and minimize breach potential ✓
C) Remove PCI DSS applicability
D) Eliminate need for controls
pip install datasets huggingface_hub
huggingface-cli login
# or
hf auth login
git clone https://github.com/cybermetric/CyberMetric.git cyber_metric
python process_cybermetric_with_docs.py --username YOUR_HF_USERNAME
--username: Your Hugging Face username (default: tuandunghcmut)--token: Hugging Face token (optional if already logged in)--data-dir: Path to CyberMetric data directory (default: cyber_metric)Each processed dataset contains these fields:
{
'question': str, # The cybersecurity question text
'option_a': str, # Multiple choice option A
'option_b': str, # Multiple choice option B
'option_c': str, # Multiple choice option C
'option_d': str, # Multiple choice option D
'correct_answer': str, # Correct answer (A, B, C, or D)
'all_options': List[str], # Formatted list of all options
'formatted_question': str # Complete question with options and answer
}
from datasets import load_dataset
# Load any CyberMetric dataset
dataset = load_dataset("tuandunghcmut/cybermetric_500_v1")
# Access a question
sample = dataset['train'][0]
print(f"Question: {sample['question']}")
print(f"Options:")
for i, option_key in enumerate(['option_a', 'option_b', 'option_c', 'option_d'], 1):
print(f" {chr(64+i)}) {sample[option_key]}")
print(f"Correct Answer: {sample['correct_answer']}")
# Use formatted version
print("\\n" + sample['formatted_question'])
| Dataset | Questions | Topics | Difficulty | Use Case |
|---|---|---|---|---|
| 80 V1 | 80 | Core concepts | Professional | Quick evaluation |
| 500 V1 | 500 | Comprehensive | Professional | Standard benchmark |
| 2000 V1 | 2,000 | Extensive | Professional | Comprehensive testing |
| 10000 V1 | 10,180 | Complete | Professional | Training & research |
Total: 12,760 cybersecurity questions 🔐
These datasets are perfect for:
If you use these datasets in your research or applications, please cite:
@misc{cybermetric2024,
title={CyberMetric: Cybersecurity Knowledge Assessment Datasets},
author={CyberMetric Contributors},
year={2024},
publisher={Hugging Face},
url={https://huggingface.co/tuandunghcmut}
}
This processing script is based on the CyberMetric dataset from:
Contributions are welcome! Please feel free to:
This script and processed datasets are released under the same license terms as the original CyberMetric repository.
🔐 Ready for cybersecurity knowledge evaluation and professional training!
Empowering the next generation of cybersecurity professionals through comprehensive, high-quality assessment data. 🛡️