Downloads · 30 days
0
BramVanroy/kenlm_sonar
kenlm_sonar is a machine learning model from BramVanroy. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository contains KenLM models (n=5) for Dutch, based on the SONAR corpus - sentence-segmented (one sentence per line). Models are provided on tokens, part-of-speech, dependency labels, and lemmas, as processed…
Downloads · 30 days
0
Access
Public
Updated Apr 16, 2024
Repo size
194 GB
Likes
1
Public
Click a slice to open those files.
.arpa61.8 GB · 66%
From the Hugging Face model README
This repository contains KenLM models (n=5) for Dutch, based on the SONAR corpus - sentence-segmented (one sentence per line). Models are provided on tokens, part-of-speech, dependency labels, and lemmas, as processed with spaCy nl_core_news_sm:
More noisy SONAR components (WRPEA, WRPED, WRUEA, WRUED, WRUEB) were excluded.
Both regular .arpa files as well as more efficient KenLM binary files (.arpa.bin) are provided. You probably want to use the binary versions.
Make sure to install dependencies:
pip install huggingface_hub
pip install https://github.com/kpu/kenlm/archive/master.zip
# If you want to use spaCy preprocessing
pip install spacy
python -m spacy download nl_core_news_sm
We can then use the Hugging Face hub software to download and cache the model file that we want, and directly use it with KenLM.
import kenlm
from huggingface_hub import hf_hub_download
model_file = hf_hub_download(repo_id="BramVanroy/kenlm_sonar", filename="kenlm_sonar_token.arpa.bin")
model = kenlm.Model(model_file)
text = "Ik eet graag koekjes !" # pre-tokenized
model.perplexity(text)
# 148.21996373689134
It is recommended to use spaCy as a preprocessor to automatically use the same tagsets and tokenization as were used when creating the LMs.
import kenlm
import spacy
from huggingface_hub import hf_hub_download
model_file = hf_hub_download(repo_id="BramVanroy/kenlm_sonar", filename="kenlm_sonar_pos.arpa.bin") # pos file
model = kenlm.Model(model_file)
nlp = spacy.load("nl_core_news_sm")
text = "Ik eet graag koekjes!"
pos_sequence = " ".join([token.pos_ for token in nlp(text)])
# 'PRON VERB ADV NOUN PUNCT'
model.perplexity(pos_sequence)
# 6.916279238079976
bin/lmplz -o 5 -S 75% -T ../data/tmp/ < ../data/processed_sonar_token_dedup.txt > ../data/kenlm_sonar_token.arpa
For class-based LMs (POS and DEP), the --discount_fallback was used and the parsed data was not deduplicated (but it was deduplicated on the sentence-level for token and lemma models).