Downloads · 30 days
0
vraj1/bert-astronomy-tokenizer
bert-astronomy-tokenizer is a machine learning model from vraj1. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
WordPiece tokenizer (30k vocab) shared across all astronomy models
Downloads · 30 days
0
Access
Public
Updated Dec 1, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.json1.3 MB · 100%
From the Hugging Face model README
WordPiece tokenizer (30k vocab) shared across all astronomy models
[PAD], [UNK], [CLS], [SEP], [MASK]from transformers import PreTrainedTokenizerFast
tokenizer = PreTrainedTokenizerFast.from_pretrained("vraj1/bert-astronomy-tokenizer")
# Tokenize text
text = "The Hubble telescope orbits Earth."
tokens = tokenizer.tokenize(text)
print(tokens)
# Output: ['the', 'hub', '##ble', 'telescope', 'orbit', '##s', 'earth', '.']
This tokenizer is part of a research project studying the effect of corpus composition on language model performance.
Project: Effect of Corpus on Language Model Performance
Institution: [Your University]
Course: NLP - Master's Computer Science
Date: November 2024