Downloads · 30 days
19
19% of all-time downloads
PushkarKumar/veritas_ai_new
veritas_ai_new is a text classification model from PushkarKumar. Use it when you need a label for a piece of text. It is set up for transformers. The card lists the license as apache-2.0.
A binary text-classification model that fine-tunes allenai/longformer-base-4096 to classify long-form news articles as REAL or FAKE, trained on a subsampled ISOT Fake News Dataset.
Downloads · 30 days
19
19% of all-time downloads
All-time downloads
101
Public
Parameters
149M
595 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors595 MB · 99%
From the Hugging Face model README
A binary text-classification model that fine-tunes allenai/longformer-base-4096 to classify long-form news articles as REAL or FAKE, trained on a subsampled ISOT Fake News Dataset.
allenai/longformer-base-40960 = REAL, 1 = FAKElongformer-base-4096 with a newly initialized 2-class classifier headtransformers (Trainer API)True.csv (REAL), Fake.csv (FAKE)label column: 0 for REAL (True.csv), 1 for FAKE (Fake.csv)title and text into full_textrandom_state=42df_small)labelAutoTokenizer.from_pretrained("allenai/longformer-base-4096")padding="max_length"truncation=Truemax_length=1024global_attention_mask as a Python list of length len(inputs["input_ids"]) with the first element set to 1 and the rest 0, then attached as inputs["global_attention_mask"](batch_size, seq_len) tensor mask used at inference timeModel init
model = AutoModelForSequenceClassification.from_pretrained(
"allenai/longformer-base-4096",
num_labels=2,
)
TrainingArguments
evaluation_strategy = "epoch"save_strategy = "epoch"learning_rate = 2e-5per_device_train_batch_size = 1per_device_eval_batch_size = 1gradient_accumulation_steps = 4num_train_epochs = 1weight_decay = 0.01fp16 = Truegradient_checkpointing = Trueload_best_model_at_end = Truepush_to_hub = Falsereport_to = "none"Trainer
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized_datasets["train"],
eval_dataset=tokenized_datasets["test"],
tokenizer=tokenizer,
)
TrainOutput.training_loss: 0.017658273408750813No accuracy, precision, recall, or F1 metrics were computed in the training script; evaluation is currently reported only via loss on the held-out test split.
Minimal example for using the model from the Hub:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
model_name = "[PushkarKumar/veritas_ai_new](https://huggingface.co/PushkarKumar/veritas_ai_new/)"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
model.eval()
def classify(text: str):
inputs = tokenizer(
text,
padding="max_length",
truncation=True,
max_length=1024,
return_tensors="pt",
)
global_attention_mask = torch.zeros(
inputs["input_ids"].shape,
dtype=torch.long,
)
global_attention_mask[:, 0] = 1
inputs["global_attention_mask"] = global_attention_mask
with torch.no_grad():
outputs = model(**inputs)
probs = torch.softmax(outputs.logits, dim=1)
label_id = int(torch.argmax(probs))
labels = {0: "REAL", 1: "FAKE"}
return labels[label_id], float(probs[label_id])