Downloads · 30 days
30
27% of all-time downloads
Shuu12121/CodeModernBERT-Crow-v1-Pre
CodeModernBERT-Crow-v1-Pre is a machine learning model from Shuu12121. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
CodeModernBERT-Crow-v1-Pre is a pretrained language model based on the ModernBERT architecture, specifically adapted for source code and docstring style natural language. It supports multiple programming languages and…
Downloads · 30 days
30
27% of all-time downloads
All-time downloads
113
Public
Parameters
153M
610 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors610 MB · 99%
From the Hugging Face model README
CodeModernBERT-Crow-v1-Pre is a pretrained language model based on the ModernBERT architecture, specifically adapted for source code and docstring style natural language. It supports multiple programming languages and is trained using large-scale code datasets curated from open-source repositories.
License: Apache-2.0
Supported Languages: Python, JavaScript, TypeScript, Java, Go, Rust, PHP, Ruby, C++, C, SQL
Datasets:
Pipeline tag: fill-mask
This model is a pretraining checkpoint, designed for further fine-tuning on downstream tasks such as semantic code search, bug detection, or code summarization.
The model was pretrained on large-scale multilingual code corpora with the following goals:
A custom BPE tokenizer was trained for code and docstrings.
Vocabulary size: 50,368
Special tokens: Standard Hugging Face special tokens + custom tokens for code/document structure.
Training process:
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("Shuu12121/CodeModernBERT-Crow-v1-Pre")
model = AutoModelForMaskedLM.from_pretrained("Shuu12121/CodeModernBERT-Crow-v1-Pre")
inputs = tokenizer("def add(a, b): return a + b", return_tensors="pt")
outputs = model(**inputs)
The model can be fine-tuned for: