Downloads · 30 days
50
17% of all-time downloads
marketeam/Fineweb-Classifier-Marketing
Fineweb-Classifier-Marketing is a text classification model from marketeam. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
<img src="https://huggingface.co/marketeam/Fineweb-Classifier-Marketing/resolve/main/assets/finewebmarketinglogo.png" alt="MarkTeam" style="width: 100%; margin-bottom: 40px;"
Downloads · 30 days
50
17% of all-time downloads
All-time downloads
298
Public
Parameters
306M
2.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.bin1.2 GB · 50%
From the Hugging Face model README
A high-performance regression model for assessing marketing content quality on a 0–5 scale. Trained on 495k marketing documents annotated by Gemma-3-27B-it, optimized for data curation and pretraining dataset filtering.
This model is built for large-scale content quality filtering, not single-document review. Typical applications:
This model predicts marketing content quality by scoring web pages against a five-criterion rubric (Relevance, Competence, Professional, Expert, Exceptional). It uses a frozen Snowflake Arctic Embed v2.0 encoder (305M params) with a trainable regression head (591k params).
The primary model is BMse (Balanced MSE), which dominates alternative training approaches on Spearman correlation, F1@3, and recall@3 simultaneously. It achieves Spearman 0.7953 on a held-out 50k-document evaluation set.
| Component | Details |
|---|---|
| Encoder | Snowflake/snowflake-arctic-embed-m-v2.0 (frozen) |
| Encoder params | 305M (not trainable) |
| Head | Linear(768→768) + ReLU + Linear(768→1) |
| Head params | 591,361 (trainable) |
| Input | [CLS] token embedding (dim=768) |
| Output | Scalar regression score, 0–5 |
| Max token length | 2048 |
Process: sample documents from FineWeb → annotate with Gemma-3-27B-it (3 independent samples per document, majority-vote aggregation) → cache frozen-encoder [CLS] embeddings → train the regression head with Balanced MSE → select the checkpoint with the best held-out Spearman correlation.
| Parameter | Value |
|---|---|
| Loss function | Balanced MSE (noise_var=1.0) |
| Optimizer | Adam, lr=3e-4 (no weight decay) |
| Batch size | 32 |
| Train set | 445,116 docs (natural distribution) |
| Eval set | 50,000 docs, held out from training (fixed split, seed 42) |
| Epochs | 5 (early stopping patience=3 on Spearman) |
| Best checkpoint | Epoch 2 |
| Seed | 42 |
| Framework versions | PyTorch 2.11.0, transformers 4.46.3 |
Source: FineWeb dataset
Sample: 500k documents from Common Crawl
Annotated: 495,116 documents with Gemma-3-27B-it — the full annotated set (including per-sample raw scores) is published at marketeam/FineWeb-Marketing-Annotations
Split:
Distribution: Heavily left-skewed (mean 1.52/5), typical of web crawl data. Top ~20% score ≥3; top ~6% score ≥4.
All usage paths below require transformers in the 4.46.x line (the frozen encoder's own custom code is not compatible with transformers 5.x): pip install "transformers==4.46.3".
from transformers import pipeline
pipe = pipeline(
"text-classification",
model="marketeam/Fineweb-Classifier-Marketing",
trust_remote_code=True,
)
document = (
"""
Marketeam.ai is more than a tool; it’s your strategic partner.
Powered by a proprietary marketing LLM and autonomous AI agents,
it integrates seamlessly into your workflow to boost precision,
efficiency, and impact. Whether you're a solo marketer or part of a larger team,
Marketeam.ai expands your capabilities and helps you tackle any challenge with confidence.
"""
)
result = pipe(document)
print(round(result["score"]))
Output:
2
trust_remote_code=True is required because this repo ships a custom PreTrainedModel/pipeline (frozen encoder + regression head is not a stock transformers architecture). Review modeling_marketing_classifier.py, configuration_marketing_classifier.py, and pipeline_marketing_classifier.py in this repo before enabling it, per standard trust_remote_code practice.
trust_remote_code)import torch
from transformers import AutoTokenizer, AutoModel
from model import MarketingClassifier
# Load the model and encoder
model = MarketingClassifier()
state_dict = torch.load("pytorch_model.bin", map_location="cpu", weights_only=True)
model.load_state_dict(state_dict)
model.eval()
# Load the embedding encoder
tokenizer = AutoTokenizer.from_pretrained("Snowflake/snowflake-arctic-embed-m-v2.0")
encoder = AutoModel.from_pretrained("Snowflake/snowflake-arctic-embed-m-v2.0")
# Score a document
text = "Marketeam.ai is more than a tool..."
inputs = tokenizer(text, return_tensors="pt", max_length=2048, truncation=True)
with torch.no_grad():
embeddings = encoder(**inputs).last_hidden_state[:, 0, :] # [CLS] token
score = model(embeddings).item()
print(f"Quality score: {score:.2f}")
If you already have [CLS] embeddings, skip the encoder:
from model import load_head_only, HIDDEN_SIZE
head = load_head_only("pytorch_model.bin", HIDDEN_SIZE)
head.eval()
with torch.no_grad():
scores = head(embeddings_tensor).squeeze(-1)
Score 5 is sparse. Only 1 document in the training set reached score 5. Percentile-based filtering is robust to this; absolute thresholds may need adjustment.
Gemma-calibrated. This classifier reflects Gemma-3-27B's annotation patterns. Different annotators produce different distributions.
English-only. Trained on English marketing content from Common Crawl. Performance on other languages is unknown.
Single score output. No per-criterion breakdown. To get C1–C5 scores, use the annotation model directly.
Frozen encoder. Only the MLP head (591k params) is trained. The embedding encoder is fixed.
Evaluated on a held-out 50k-document split. Best checkpoint selected by eval Spearman.
| Metric | Value |
|---|---|
| Spearman | 0.7953 |
| F1@3 | 0.691 |
| Precision@3 | 0.717 |
| Recall@3 | 0.666 |
| Pred std | 1.88 |
| % predicted ≥3 | 18.7% (true ≥3 = 20.2%) |
Spearman correlation directly measures how well the model ranks documents by quality, which is critical for data filtering. F1@3 is order-invariant and unsuitable for ordinal (0–5) scoring. BMse's 0.7953 Spearman indicates strong correlation between predicted and ground-truth quality scores.
For scoring at FineWeb scale (all CC-MAIN dumps or a specific one), use the batch inference path in infer.py, which streams a dump, scores in batches, and writes (id, dump, score, int_score, percentile) parquet shards per dump — see scripts/run_inference.sh for the CLI invocation. Percentile rank is computed per-dump so it stays stable when dumps are processed independently.
For ad hoc or low-volume scoring, the pipeline(...) one-liner above is sufficient.
See requirements.txt for all dependencies. Key packages: