Downloads · 30 days
95
24% of all-time downloads
BidirLM/BidirLM-1B-Base
BidirLM-1B-Base is a fill-mask model from BidirLM. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
BidirLM-1B-Base is the intermediate MNTP-adapted checkpoint of the BidirLM family. It is obtained by converting Gemma3-1B from causal to bidirectional attention and training with Masked Next Token Prediction (MNTP) on…
Downloads · 30 days
95
24% of all-time downloads
All-time downloads
404
Public
Parameters
1000M
2.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2 GB · 99%
From the Hugging Face model README
BidirLM-1B-Base is the intermediate MNTP-adapted checkpoint of the BidirLM family. It is obtained by converting Gemma3-1B from causal to bidirectional attention and training with Masked Next Token Prediction (MNTP) on 30B tokens from a multi-domain corpus (FineWeb-Edu, FineWeb2-HQ, FineMath, Stack V2), then merged 50/50 with the original Gemma3-1B weights.
For general embeddings and downstream fine-tuning, use BidirLM/BidirLM-1B-Embedding which adds contrastive training on top of this checkpoint.
from transformers import AutoTokenizer, AutoModel, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("BidirLM/BidirLM-1B-Base", trust_remote_code=True)
# Base encoder
model = AutoModel.from_pretrained("BidirLM/BidirLM-1B-Base", trust_remote_code=True)
# Masked language model
mlm = AutoModelForMaskedLM.from_pretrained("BidirLM/BidirLM-1B-Base", trust_remote_code=True)
transformers>=5.0
This model requires trust_remote_code=True.
Note: This model was trained with
transformers==4.57.6(transformers 4.x). The version onmainwas patched to work withtransformers>=5.0. For the original (pre-patch) version, which is compatible withtransformers>=4.57.6,<5.0.0, use thetransformers-v4branch:from transformers import AutoModel model = AutoModel.from_pretrained( "BidirLM/BidirLM-1B-Base", trust_remote_code=True, revision="transformers-v4", )
@misc{boizard2026bidirlmtextomnimodalbidirectional,
title={BidirLM: From Text to Omnimodal Bidirectional Encoders by Adapting and Composing Causal LLMs},
author={Nicolas Boizard and Théo Deschamps-Berger and Hippolyte Gisserot-Boukhlef and Céline Hudelot and Pierre Colombo},
year={2026},
eprint={2604.02045},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2604.02045},
}