Downloads · 30 days
21
100% of all-time downloads
Mudi12137/WaterBERT
WaterBERT is a fill-mask model from Mudi12137. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
WaterBERT is a BERT-base encoder for wastewater- and water-treatment literature. It was obtained by continuing masked-language-model (MLM) pretraining of SciBERT (scivocab, uncased) on a domain corpus of water-treatme…
Downloads · 30 days
21
100% of all-time downloads
All-time downloads
21
Public
Parameters
110M
440 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors440 MB · 100%
From the Hugging Face model README
WaterBERT is a BERT-base encoder for wastewater- and water-treatment literature. It was
obtained by continuing masked-language-model (MLM) pretraining of
SciBERT (scivocab, uncased) on a
domain corpus of water-treatment research text. The vocabulary is SciBERT's (31,090 tokens);
only the weights were adapted.
It is the shared backbone of WaterBERT-NER and WaterBERT-RE, and is intended as a starting point for fine-tuning on water-treatment NLP tasks (entity recognition, relation extraction, classification, retrieval).
from transformers import pipeline
fill = pipeline("fill-mask", model="Mudi12137/WaterBERT")
fill("The Fenton process generates hydroxyl radicals from hydrogen peroxide and [MASK] ions.")
# top prediction: "iron"
For fine-tuning, load it like any BERT checkpoint:
from transformers import AutoModelForTokenClassification, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Mudi12137/WaterBERT")
model = AutoModelForTokenClassification.from_pretrained("Mudi12137/WaterBERT", num_labels=13)
| Setting | Value |
|---|---|
| Initialisation | allenai/scibert_scivocab_uncased |
| Objective | masked language modelling |
| Epochs / steps | 30 / 940,440 (this checkpoint is the final step) |
| Effective batch size | 160 (80 per device x 2 gradient-accumulation steps, 1 GPU) |
| Optimiser | AdamW, learning rate 1e-4, linear decay, warmup ratio 0.048, weight decay 0.01 |
| Precision | bf16 |
| Architecture | BERT-base: 12 layers, hidden size 768, 110M parameters, max length 512 |
The pretraining corpus and its construction are described in the accompanying paper.
Apache-2.0, the same as the SciBERT base model.
See CITATION.cff in the project repository.