Downloads · 30 days
0
annalhq/truthseek
truthseek is a text classification model from annalhq. Use it when you need a label for a piece of text. The card lists the license as mit.
The model is designed to classify emails as spam or not spam, trained on the Enron dataset. It integrates several advanced deep learning techniques to effectively handle the complex patterns inherent in email text. He…
Downloads · 30 days
0
Access
Public
Updated Feb 21, 2025
Repo size
1.8 GB
Likes
1
Public
Click a slice to open those files.
.pkl1.3 GB · 75%
From the Hugging Face model README
The model is designed to classify emails as spam or not spam, trained on the Enron dataset. It integrates several advanced deep learning techniques to effectively handle the complex patterns inherent in email text. Here's how the components work together:
BERT: The first step in the pipeline utilizes BERT, a transformer model pre-trained on a vast text corpus. BERT is used for encoding the input email text, capturing the semantic meaning and contextual relationships between words.
CNN (Convolutional Neural Network)*: After encoding the text with RoBERTa, a CNN is applied to extract features from the encoded representations. This step focuses on identifying local patterns and phrases that are indicative of spam or non-spam content.
BiLSTM (Bidirectional Long Short-Term Memory): To enhance the model's ability to understand context over longer sequences, a BiLSTM is used for sentence-level embeddings. It processes the encoded text in both directions, allowing the model to capture dependencies and relationships between words across the entire email.
Hierarchical Attention Network: Finally, a Hierarchical Attention Network (HAN) is applied to help the model focus on the most important words and sentences within an email. This attention mechanism allows the model to prioritize critical features that distinguish spam from non-spam messages.