Skip to content

Icelandic-lt

convbert-small-igc-is

Icelandic-lt/convbert-small-igc-is

convbert-small-igc-is is a feature extraction model from Icelandic-lt. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as cc-by-4.0.

This model was pretrained on the Icelandic Gigaword Corpus, which contains approximately 1.69B tokens, using default settings. The model uses a Unigram tokenizer with a vocabulary size of 96,000.

Downloads · 30 days

22

20% of all-time downloads

All-time downloads

109

Public

Parameters

21.5M

259 MB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.h586.4 MB · 33%

Parameter types

How the weights are stored.

F3221.5M · 100%

Datasets

Task
Feature Extraction
Library
transformers
Type
convbert
License
cc-by-4.0
Languages
is
Created
May 27, 2024
Updated
Feb 21, 2025