Downloads · 30 days
20
30% of all-time downloads
PeytonT/repo-embedding
repo-embedding is a feature extraction model from PeytonT. Use it when you need embeddings to search or compare text. It is set up for transformers.
Produces repository-level dense representations for retrieval and alignment.
Downloads · 30 days
20
30% of all-time downloads
All-time downloads
67
Public
Parameters
22.7M
90.9 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors90.9 MB · 99%
From the Hugging Face model README
Produces repository-level dense representations for retrieval and alignment.
sentence-transformers/all-MiniLM-L6-v2encoderR1T4_repoThis model is part of the Repository Library stack, a research system for indexing, retrieving, aligning, and reasoning over scientific papers, structured paper content, repositories, and cross-domain links between them.
https://huggingface.co/PeytonT/repo-embeddinghttps://huggingface.co/collections/PeytonT/research-library-6a49c589ef4d763f7539b50dhttps://github.com/peytontolbert/research_libraryhttps://github.com/peytontolbert/research_library/blob/main/models/experiments/r1_repo_embedding.jsonhttps://github.com/peytontolbert/research_library/tree/main/modelsThe training inputs for this package were assembled from the following Repository Library data sources:
github_repos: repository graph and code chunk data exported from the Repository Library repo pipeline.github_reposrepo_querysource_chunk[0.9, 0.1, 0.0]40008bf16contrastive5e-05256256full_finetune1000ddp0recall_at_10, ndcg_at_10from transformers import AutoModel, AutoTokenizer
repo_id = "PeytonT/repo-embedding"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModel.from_pretrained(repo_id)
https://github.com/peytontolbert/research_libraryhttps://huggingface.co/collections/PeytonT/research-library-6a49c589ef4d763f7539b50dPeytonT