Downloads · 30 days
9
21% of all-time downloads
PeytonT/paper-fulltext-embedding
paper-fulltext-embedding is a feature extraction model from PeytonT. Use it when you need embeddings to search or compare text. It is set up for peft.
Produces embeddings over full paper text for retrieval and clustering tasks.
Downloads · 30 days
9
21% of all-time downloads
All-time downloads
43
Public
Repo size
1.2 MB
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 MB · 55%
From the Hugging Face model README
Produces embeddings over full paper text for retrieval and clustering tasks.
allenai/scibert_scivocab_uncasedencoderM6T3_paper_textThis model is part of the Repository Library stack, a research system for indexing, retrieving, aligning, and reasoning over scientific papers, structured paper content, repositories, and cross-domain links between them.
https://huggingface.co/PeytonT/paper-fulltext-embeddinghttps://huggingface.co/collections/PeytonT/research-library-6a49c589ef4d763f7539b50dhttps://github.com/peytontolbert/research_libraryhttps://github.com/peytontolbert/research_library/blob/main/models/experiments/m6_paper_fulltext_embedding.jsonhttps://github.com/peytontolbert/research_library/tree/main/modelsThe training inputs for this package were assembled from the following Repository Library data sources:
local/paper_text_2m_dedup_v1paper_text_parquet: full-text paper corpus records prepared for model training.paper_text_parquettitle, abstract, textfulltext_embedding[0.8, 0.1, 0.1]04bf16contrastive0.0001512128peft_lora1000ddp0recall_at_10, ndcg_at_10from transformers import AutoModel, AutoTokenizer
from peft import PeftModel
repo_id = "PeytonT/paper-fulltext-embedding"
base_id = "allenai/scibert_scivocab_uncased"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
base = AutoModel.from_pretrained(base_id)
model = PeftModel.from_pretrained(base, repo_id)
https://github.com/peytontolbert/research_libraryhttps://huggingface.co/collections/PeytonT/research-library-6a49c589ef4d763f7539b50dPeytonT