Downloads · 30 days
17
49% of all-time downloads
novelcore/gem-roberta-bilingual
gem-roberta-bilingual is a fill-mask model from novelcore. Use it when you need the model to fill a missing word. It is set up for transformers. The card lists the license as apache-2.0.
TGEM-RoBERTa Legal Bilingual is a RoBERTa-base model pre-trained from scratch on a comprehensive 26GB bilingual corpus of Greek and English legal, parliamentary, and governmental text. This model represents the first…
Downloads · 30 days
17
49% of all-time downloads
All-time downloads
35
Public
Parameters
125M
499 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors499 MB · 99%
From the Hugging Face model README
TGEM-RoBERTa Legal Bilingual is a RoBERTa-base model pre-trained from scratch on a comprehensive 26GB bilingual corpus of Greek and English legal, parliamentary, and governmental text. This model represents the first large-scale bilingual legal language model combining Greek and English legal domains, enabling cross-lingual legal understanding and applications.
The model employs the RoBERTa architecture optimized for legal text understanding across both languages, with dynamic masking and focused Masked Language Modeling (MLM) training. The bilingual approach allows the model to leverage legal concepts and terminology from both the Greek and Anglo-American legal traditions.
This model builds upon legal datasets including portions of the Pile of Law collection from Hugging Face, combined with comprehensive Greek legal corpora to create a unique bilingual legal language resource.
You can use this model directly with the fill-mask pipeline:
from transformers import pipeline
# Load the model
fill_mask = pipeline(
"fill-mask",
model="novelcore/gem-roberta-bilingual",
tokenizer="novelcore/gem-roberta-bilingual"
)
# Example in Greek
text_gr = "Ο κ. Μητσοτάκης <mask> ότι η κυβέρνηση σέβεται πλήρως τις αποφάσεις του Συμβουλίου της Επικρατείας."
predictions_gr = fill_mask(text_gr)
print("Greek predictions:", predictions_gr)
For downstream tasks:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
# For bilingual legal document classification
tokenizer = AutoTokenizer.from_pretrained("novelcore/gem-roberta-bilingual")
model = AutoModelForSequenceClassification.from_pretrained("novelcore/gem-roberta-bilingual")
# Process texts in both languages
greek_text = "Το Συνταγματικό Δικαστήριο αποφάσισε..."
english_text = "The Constitutional Court decided..."
The model was pre-trained on a comprehensive 26GB bilingual corpus comprising 60.3% Greek legal content (13.85GB) and 39.7% English legal content (9.12GB), creating a balanced exposure to both legal traditions.
| Dataset | Size (GB) | Context | Rationale |
|---|---|---|---|
| FEK - Greek Government Gazette | 11.0 | Legal/Regulatory | Official government publications, regulatory language |
| Greek Parliament Proceedings | 2.9 | Legal/Parliamentary | Legislative discourse, policy language |
| Political Reports of Supreme Court | 1.2 | Legal/Judicial | High-level judicial reasoning, precedents |
| Eur-Lex (Greek Content) | 0.92 | Legal/EU | EU legal documents, multilingual legal terminology |
| Europarl (Greek Content) | 0.38 | Legal/Parliamentary | Parliamentary proceedings, EU legislative language |
| Raptarchis Legal Dictionary | 0.35 | Legal/Reference | Legal terminology, definitions |
| Dataset | Size (GB) | Context | Greek Equivalent |
|---|---|---|---|
| CourtListener Opinions | 4.2 | Legal/Judicial | Supreme Court Reports |
| EDGAR (SEC Filings) | 2.1 | Legal/Corporate | Corporate regulatory compliance |
| Eur-Lex (English) | 1.1 | Legal/EU | Direct parallel to Greek Eur-Lex |
| US Bills | 1.0 | Legal/Legislative | Parliamentary proceedings equivalent |
| CFR (Code of Federal Regulations) | 0.48 | Legal/Regulatory | Federal regulatory framework |
| Europarl (English) | 0.24 | Legal/Parliamentary | Direct parallel to Greek Europarl |
| Federal Register | 0.12 | Legal/Regulatory | Government gazette equivalent (FEK) |
Note: English legal datasets partially sourced from the Pile of Law collection on Hugging Face.
The 60:40 Greek-to-English ratio was designed to:
The model uses the RoBERTa-base architecture with the following configuration:
The text was tokenized using a custom ByteLevelBPE tokenizer trained from scratch on the bilingual Greek-English legal corpus. The tokenizer uses a vocabulary of 50,264 tokens optimized for both Greek and English legal terminology, enabling effective cross-lingual representation.
The data was processed into fixed-size chunks of 512 tokens, respecting document boundaries to ensure contextual coherence across both languages.
The model was pre-trained from scratch for 150,000 steps on 8x NVIDIA H100 GPUs, using BFloat16 (bf16) mixed-precision for stability and speed. The training took approximately 25 hours and 7 minutes to complete.
The key hyperparameters used were:
per_device_train_batch_size: 128, gradient_accumulation_steps: 2)The model achieved the following performance metrics: