Downloads · 30 days
459
3% of all-time downloads
amahdaouy/DomURLs_BERT
DomURLs_BERT is a feature extraction model from amahdaouy. Use it when you need embeddings to search or compare text. It is set up for transformers.
DomURLsBERT is a pre-trained BERT-based encoder adapted for detecting and classifying suspicious/malicious domains and URLs. DomURLsBERT is pre-trained using the Masked Language Modeling (MLM) objective on a large mul…
Downloads · 30 days
459
3% of all-time downloads
All-time downloads
17.6K
Public
Parameters
111M
442 MB on disk
Likes
3
Public
Click a slice to open those files.
.safetensors442 MB · 100%
From the Hugging Face model README
DomURLs_BERT is a pre-trained BERT-based encoder adapted for detecting and classifying suspicious/malicious domains and URLs. DomURLs_BERT is pre-trained using the Masked Language Modeling (MLM) objective on a large multilingual corpus of URLs, domain names, and Domain Generation Algorithms (DGA) dataset.
Detecting and classifying suspicious or malicious domain names and URLs is fundamental task in cybersecurity. To leverage such indicators of compromise, cybersecurity vendors and practitioners often maintain and update blacklists of known malicious domains and URLs. However, blacklists frequently fail to identify emerging and obfuscated threats. Over the past few decades, there has been significant interest in developing machine learning models that automatically detect malicious domains and URLs, addressing the limitations of blacklists maintenance and updates. In this paper, we introduce DomURLs_BERT, a pre-trained BERT-based encoder adapted for detecting and classifying suspicious/malicious domains and URLs. DomURLs_BERT is pre-trained using the Masked Language Modeling (MLM) objective on a large multilingual corpus of URLs, domain names, and Domain Generation Algorithms (DGA) dataset. In order to assess the performance of DomURLs_BERT, we have conducted experiments on several binary and multi-class classification tasks involving domain names and URLs, covering phishing, malware, DGA, and DNS tunneling. The evaluations results show that the proposed encoder outperforms state-of-the-art character-based deep learning models and cybersecurity-focused BERT models across multiple tasks and datasets. The pre-training dataset, the pre-trained DomURLs_BERT encoder, and the experiments source code are publicly available.
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
@article{ElMahdaouy2026DomURLsBERT,
author = {El Mahdaouy, Abdelkader and
Lamsiyah, Salima and
Janati Idrissi, Meryem and
Alami, Hamza and
Yartaoui, Zakaria and
Berrada, Ismail},
title = {{DomURLs\_BERT}: Pre-trained BERT-based Model for Malicious Domains and URLs Detection and Classification},
journal = {Journal of Network and Systems Management},
year = {2026},
volume = {34},
number = {2},
pages = {36},
doi = {10.1007/s10922-025-10010-9},
url = {https://doi.org/10.1007/s10922-025-10010-9},
issn = {1573-7705},
}
@article{domurlsbert2024,
title={{DomURLs\_BERT}: Pre-trained BERT-based Model for Malicious Domains and URLs Detection and Classification},
author={Abdelkader {El Mahdaouy} and Salima Lamsiyah and Meryem {Janati Idrissi} and Hamza Alami and Zakaria Yartaoui and Ismail Berrada},
journal={arXiv preprint arXiv:2409.09143},
year={2024},
eprint={2409.09143},
archivePrefix={arXiv},
primaryClass={cs.CR},
url={https://arxiv.org/abs/2409.09143},
}
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
Use the code below to get started with the model.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]