Downloads · 30 days
12
33% of all-time downloads
smkrv/xlm-roberta-multihead-coreml
xlm-roberta-multihead-coreml is a text classification model from smkrv. Use it when you need a label for a piece of text. It is set up for coreml. The card lists the license as mit.
Fine-tuned CoreML version of XLM-RoBERTa-base with three classification heads for on-device multilingual text analysis on Apple Silicon. Performs sentiment analysis, multi-label tagging, and named entity recognition i…
Downloads · 30 days
12
33% of all-time downloads
All-time downloads
36
Public
Repo size
1.7 GB
Likes
0
Public
Click a slice to open those files.
.bin833 MB · 99%
From the Hugging Face model README
Fine-tuned CoreML version of XLM-RoBERTa-base with three classification heads for on-device multilingual text analysis on Apple Silicon. Performs sentiment analysis, multi-label tagging, and named entity recognition in a single forward pass.
.mlpackage (mlprogram)input_ids + attention_mask, int32)Single-label classification: positive, neutral, risk, toxic
Multi-label phrase tagging: stress_signal, confidence, emotional_state, trust_indicator, defensiveness, active_listening, rapport_building, conflict_signal, cooperation, clarification, pressure_tactic, concession, information_sharing, commitment, deadline_mention, deception_signal, manipulation, power_dynamic, agreement, problem_solving
Named entity recognition: O, B-PER, I-PER, B-ORG, I-ORG, B-MONEY, I-MONEY, B-DATE, I-DATE
| Head | Metric | Value |
|---|---|---|
| Sentiment | Accuracy | 76.6% |
| Tags | F1 (macro) | 57.0% |
| NER | Accuracy | 94.8% |
| Combined | Val Loss | 0.492 |
| File | Size | Description |
|---|---|---|
XLMRobertaMultiHead.mlpackage/ | 529 MB | FP16 model |
XLMRobertaMultiHead_INT8.mlpackage/ | 266 MB | INT8 quantized (recommended) |
label_definitions.json | 1 KB | Label mappings for all heads |
config.json | — | Model configuration |
import CoreML
let model = try MLModel(contentsOf: modelURL)
// Prepare inputs (use XLM-RoBERTa tokenizer)
let inputArray = try MLMultiArray(shape: [1, 128], dataType: .int32)
let maskArray = try MLMultiArray(shape: [1, 128], dataType: .int32)
// ... fill with tokenized text ...
let input = try MLDictionaryFeatureProvider(dictionary: [
"input_ids": MLFeatureValue(multiArray: inputArray),
"attention_mask": MLFeatureValue(multiArray: maskArray)
])
let output = try model.prediction(from: input)
let sentimentLogits = output.featureValue(for: "sentiment_logits")!.multiArrayValue!
let tagLogits = output.featureValue(for: "tag_logits")!.multiArrayValue!
let nerLogits = output.featureValue(for: "ner_logits")!.multiArrayValue!
This model uses the standard XLM-RoBERTa tokenizer from FacebookAI/xlm-roberta-base. CoreML does not include the tokenizer — use tokenizers library or bundle tokenizer.json separately.
Base model XLM-RoBERTa by Facebook AI. Fine-tuning on Russian/English datasets and CoreML conversion by @smkrv.