Downloads · 30 days
81
100% of all-time downloads
langtech-languagemodeling/bsc-edu-annotator
bsc-edu-annotator is a machine learning model from langtech-languagemodeling. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
A multilingual text-embedding regression model for assigning a continuous educational-value score to web documents. The model is designed for quality-based data filtering, producing scores in the range [0, 4].
Downloads · 30 days
81
100% of all-time downloads
All-time downloads
81
Public
Parameters
568M
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
From the Hugging Face model README
A multilingual text-embedding regression model for assigning a continuous educational-value score to web documents. The model is designed for quality-based data filtering, producing scores in the range [0, 4].
The BSC-EDU classifier is a lightweight multilingual regression model based on a text-embedding encoder. It is fine-tuned to predict a continuous educational-value score for text documents. The resulting score can be used to rank or filter documents according to their predicted educational value.
The model was trained using examples in Spanish, Catalan, and Basque, with 500,000 examples per language. The first 512 tokens from each example were used. The model's predictions are continuous rather than restricted to discrete classes, allowing flexible score thresholds for downstream data filtering.
BibTeX:
@inproceedings{bsc-edu,
title = {BSC-EDU},
note = {}
}