Downloads · 30 days
109
11% of all-time downloads
thellert/accphysbert_cased
accphysbert_cased is a feature extraction model from thellert. Use it when you need embeddings to search or compare text. It is set up for sentence-transformers. The card lists the license as cc-by-4.0.
AccPhysBERT is a specialized sentence-embedding model fine-tuned for accelerator physics, capturing semantic nuances in this technical domain. It delivers state-of-the-art performance in tasks such as semantic search,…
Downloads · 30 days
109
11% of all-time downloads
All-time downloads
983
Public
Parameters
109M
438 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors438 MB · 100%
From the Hugging Face model README
AccPhysBERT is a specialized sentence-embedding model fine-tuned for accelerator physics, capturing semantic nuances in this technical domain. It delivers state-of-the-art performance in tasks such as semantic search, citation classification, reviewer matching, and clustering of accelerator-physics literature.
Developed by: Thorsten Hellert, João Montenegro, Marco Venturini, Andrea Pollastro
Funded by: US Department of Energy, Lawrence Berkeley National Laboratory
Model Type: Sentence embedding (BERT-based, SimCSE fine-tuned)
Language: English
License: CC BY 4.0
Paper: Domain-specific text embedding model for accelerator physics, Phys. Rev. Accel. Beams 28, 044601 (2025)
https://doi.org/10.1103/PhysRevAccelBeams.28.044601
Core Corpus:
Annotation Sources:
| Task | Metric | Score |
|---|---|---|
| Citation Classification | Cosine Accuracy | 91.0% |
| Category Clustering | V‑measure (main/sub) | 63.7 / 77.2 |
| Information Retrieval | nDCG@10 | 66.3 |
AccPhysBERT outperforms BERT, SciBERT, and large general-purpose embedding models in all accelerator-specific benchmarks.
from transformers import AutoTokenizer, AutoModel
import torch
tokenizer = AutoTokenizer.from_pretrained("thellert/accphysbert")
model = AutoModel.from_pretrained("thellert/accphysbert")
text = "We report on beam instabilities observed in the LCLS-II injector."
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
# Use mean pooling (excluding [CLS] and [SEP])
token_embeddings = outputs.last_hidden_state[:, 1:-1, :]
sentence_embedding = token_embeddings.mean(dim=1)
If you use AccPhysBERT, please cite:
@article{Hellert_2025,
title = {Domain-specific text embedding model for accelerator physics},
author = {Hellert, Thorsten and Montenegro, João and Venturini, Marco and Pollastro, Andrea},
journal = {Physical Review Accelerators and Beams},
volume = {28},
number = {4},
pages = {044601},
year = {2025},
publisher = {American Physical Society},
doi = {10.1103/PhysRevAccelBeams.28.044601},
url = {https://doi.org/10.1103/PhysRevAccelBeams.28.044601}
}
Thorsten Hellert
Lawrence Berkeley National Laboratory
📧 [email protected]
This model builds on PhysBERT and was trained using NERSC resources. Thanks to Alex Hexemer, Fernando Sannibale, and Antonin Sulc for their support and discussions.