Downloads · 30 days
0
CALDISS-AAU/da-reported-speech-e5
da-reported-speech-e5 is a text classification model from CALDISS-AAU. Use it when you need a label for a piece of text. It is set up for setfit. The card lists the license as mit.
Downloads · 30 days
0
Access
Public
Updated Apr 1, 2025
Repo size
2.3 GB
Likes
2
Public
Click a slice to open those files.
.safetensors2.2 GB · 99%
From the Hugging Face model README
This fine-tuned few-shot model
This modelcard aims to be a base template for new models. It has been generated using this raw template.
Base model: intfloat/multilingual-e5-large Language: Danish (da) Task: Reported Speech Detection Training data: Danish jobcenter conversation transcripts
This model is a few-shot classifier fine-tuned on transcribed interviews from a job center in Denmark. It is designed for binary classification of reported speech, identifying sentences where a speaker references or quotes another person.
To support real-world usage, this model is integrated into a two-part processing pipeline that allows users to analyze interview documents and highlight relevant sentences.
This model is used in a document processing pipeline that performs the following tasks:
Additionally, a GUI-based wrapper (built with Gooey) provides a user-friendly .exe program, allowing non-technical users to process documents efficiently. For a more in-depth view for the GUI, please read the Github page provided further down.
The model is trained and evaluated on text snippets of "reported speech" in Danish interviews between citizens and job counselors. It is intended to identify "reported speech" in similar text documents of that genre. It is assumed unsuitable for general classification of "reported speech".
Inteded users inludes researchers or analysts working with danish conversational data or transcripts specifically interested in reported speech as a phenomenon.
Following group (but not excluded to) may find it useful:
Social Scientists & political scientist:
Linguists & NLP researchers:
The model is trained on Danish job center interviews, so performance may vary on other types of texts.
Binary classification is based on reported speech detection, but edge cases may exist.
While based on a multilingual model, this fine-tuned version is specifically optimized for Danish. Performance may be unreliable in other languages.
The model assumes transcripts. Messy, formal, or highly unstructured text (e.g., speech-to-text outputs with errors) may reduce accuracy.
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model.
Use the code below to get started with the model.
from transformers import AutoModel, AutoTokenizer
model_name = "your-huggingface-username/danish-rep-speech-e5"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)
text = "Han sagde: 'Jeg kommer i morgen.'"
inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)
# Extract the embedding
embedding = outputs.last_hidden_state[:, 0, :].detach().numpy()
Training data consits of 55 transcripts of conversations between a citizen and a social worker collected from a danish jobcenter. Data is therefore sensitive and not attached in this model card. Data was further evaluated to be balanced and containing a 50/50 split between both tags.
Pretraining & Base Model:
This model is fine-tuned on top of intfloat/multilingual-e5-large, a transformer-based model optimized for embedding-based retrieval. The base model was pretrained using contrastive learning and large-scale multilingual datasets, making it well-suited for semantic similarity and classification tasks. Fine-Tuning Details
Training Dataset:
The model was fine-tuned using labelled transcribed interviews from a Danish job center.
Due to the sensitive nature of the data, it is not publicly available.
Objective:
The model was trained for binary classification of reported speech.
Labels indicate whether a sentence contains reported speech (reported-speech, not reported-speech).
Training Configuration:
Few-shot learning approach with domain-specific samples.
Batch size: 32.
Body Learning rate: 1.0770502781075495e-06
Solver: lbfgs.
Number of epochs: 6
Max Iterations: 279
Evaluation metric: Accuracy & F1-score.
Technical Implementation
Tokenization performed using the SentencePiece-based tokenizer from intfloat/multilingual-e5-large.
Fine-tuning was done using PyTorch and the Hugging Face Trainer API.
The model is optimized for batch inference rather than real-time processing.
📌 For more details on the architecture, refer to the base model: multilingual-e5-large.
metrics:
- type: accuracy
value: 0.9724770642201835
name: Accuracy
- type: precision
value: 0.9557522123893806
name: Precision
- type: recall
value: 0.9908256880733946
name: Recall
- type: f1
value: 0.972972972972973
name: F1
[More Information Needed]
The model was evaluated using standard classification metrics to measure its performance. Evaluation Metrics
Accuracy: Measures the overall correctness of predictions.
F1-Score: Balances precision and recall, ensuring that both false positives and false negatives are considered.
Precision: Measures how many of the predicted reported speech sentences are actually correct.
Results:
Not Reported Speech: Precision: 0.959 Recall: 0.924 F1-Score: 0.941 Recall: 0.942
Reported Speech: Precision: 0.927 Recall: 0.961 F1: 0.943
Accuracy: 0.942
Ucloud-cloud infrastructure available at the danish universities
BibTeX:
@article{https://doi.org/10.48550/arxiv.2209.11055, doi = {10.48550/ARXIV.2209.11055}, url = {https://arxiv.org/abs/2209.11055}, author = {Tunstall, Lewis and Reimers, Nils and Jo, Unso Eun Seo and Bates, Luke and Korat, Daniel and Wasserblat, Moshe and Pereg, Oren}, keywords = {Computation and Language (cs.CL), FOS: Computer and information sciences, FOS: Computer and information sciences}, title = {Efficient Few-Shot Learning Without Prompts}, publisher = {arXiv}, year = {2022}, copyright = {Creative Commons Attribution 4.0 International} }