Downloads · 30 days
5
11% of all-time downloads
synturk/sentagram
sentagram is a token classification model from synturk. Use it when you need labels on individual words, such as names. It is set up for transformers. The card lists the license as apache-2.0.
The SENTAGRAM model is a BERT-based model fine-tuned on a custom Turkish grammar dataset. It is designed to analyze and classify grammatical elements within Turkish sentences, such as subjects, predicates, objects, an…
Downloads · 30 days
5
11% of all-time downloads
All-time downloads
45
Public
Parameters
184M
1.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors735 MB · 99%
From the Hugging Face model README
The SENTAGRAM model is a BERT-based model fine-tuned on a custom Turkish grammar dataset. It is designed to analyze and classify grammatical elements within Turkish sentences, such as subjects, predicates, objects, and adjuncts. The model is built on the BERTürk architecture, specifically adapted to understand and process the intricacies of Turkish grammar. For more information visit GitHub repository of project.
The model was evaluated on the SYNTÜRK SENTAGRAM dataset with the following results:
| Precision | Recall | F1 Score | Accuracy |
|---|---|---|---|
| 0.911349 | 0.911826 | 0.911588 | 0.935395 |
These metrics demonstrate the model's effectiveness in correctly identifying and classifying grammatical elements in Turkish sentences.
You can load and use the model with Hugging Face's transformers library:
from transformers import AutoTokenizer, AutoModelForTokenClassification
# Load the tokenizer and model
tokenizer = AutoTokenizer.from_pretrained("synturk/sentagram-berturk")
model = AutoModelForTokenClassification.from_pretrained("synturk/sentagram-berturk")
# Example sentence
sentence = "SYNTÜRK yarışmayı kazandı."
# Tokenize and predict
inputs = tokenizer(sentence, return_tensors="pt")
outputs = model(**inputs)
predictions = torch.argmax(outputs.logits, dim=2)
# Decode the predictions
tokens = tokenizer.convert_ids_to_tokens(inputs["input_ids"][0])
predicted_labels = [model.config.id2label[p.item()] for p in predictions[0]]
print(list(zip(tokens, predicted_labels)))
We plan to enhance the model by integrating additional grammatical features, such as semantic roles and more complex sentence structures. This will further improve its ability to process and understand the nuances of the Turkish language.
This model is licensed under the Apache 2.0 License.
If you use this model in your research or applications, please cite it as follows:
@model{synturk-sentagram,
author = {SYNTÜRK Team},
title = {SENTAGRAM Model},
year = {2024},
publisher = {Hugging Face},
url = {https://huggingface.co/synturk/sentagram},
}
For more information or questions, please contact the SYNTÜRK Team through our GitHub repository.
Follow SYNTÜRK Team on,