Downloads · 30 days
17
3% of all-time downloads
EMBO/sd-smallmol-roles
sd-smallmol-roles is a token classification model from EMBO. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as agpl-3.0.
This model is a RoBERTa base model that was further trained using a masked language modeling task on a compendium of english scientific textual examples from the life sciences using the BioLang dataset. It has then be…
Downloads · 30 days
17
3% of all-time downloads
All-time downloads
623
Public
Repo size
993 MB
Likes
0
Public
Click a slice to open those files.
.bin496 MB · 99%
From the Hugging Face model README
This model is a RoBERTa base model that was further trained using a masked language modeling task on a compendium of english scientific textual examples from the life sciences using the BioLang dataset. It has then been fine-tuned for token classification on the SourceData sd-nlp dataset with the SMALL_MOL_ROLES configuration to perform pure context-dependent semantic role classification of bioentities.
The intended use of this model is to infer the semantic role of small molecules with regard to the causal hypotheses tested in experiments reported in scientific papers.
To have a quick check of the model:
from transformers import pipeline, RobertaTokenizerFast, RobertaForTokenClassification
example = """<s>The <mask> overexpression in cells caused an increase in <mask> expression.</s>"""
tokenizer = RobertaTokenizerFast.from_pretrained('roberta-base', max_len=512)
model = RobertaForTokenClassification.from_pretrained('EMBO/sd-smallmol-roles')
ner = pipeline('ner', model, tokenizer=tokenizer)
res = ner(example)
for r in res:
print(r['word'], r['entity'])
The model must be used with the roberta-base tokenizer.
The model was trained for token classification using the EMBO/sd-nlp dataset which includes manually annotated examples.
The training was run on a NVIDIA DGX Station with 4XTesla V100 GPUs.
Training code is available at https://github.com/source-data/soda-roberta
per_device_train_batch_size: 16per_device_eval_batch_size: 16learning_rate: 0.0001weight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08max_grad_norm: 1.0On 7178 example of test set with sklearn.metrics:
precision recall f1-score support
CONTROLLED_VAR 0.76 0.90 0.83 2946
MEASURED_VAR 0.60 0.71 0.65 852
micro avg 0.73 0.86 0.79 3798
macro avg 0.68 0.80 0.74 3798
weighted avg 0.73 0.86 0.79 3798
{'test_loss': 0.011743436567485332, 'test_accuracy_score': 0.9951612532624371, 'test_precision': 0.7261345852895149, 'test_recall': 0.8551869404949973, 'test_f1': 0.7853947527505744, 'test_runtime': 58.0378, 'test_samples_per_second': 123.678, 'test_steps_per_second': 1.947}