Downloads · 30 days
0
koushikkanch/midtermnexttokenprediction
midtermnexttokenprediction is a machine learning model from koushikkanch. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This project implements a neural network-based language model for next token prediction in both English and Czech. The model uses an LSTM architecture and is trained on a combined dataset of English and Czech text.
Downloads · 30 days
0
Access
Public
Updated Oct 11, 2024
Repo size
31.1 MB
Likes
0
Public
Click a slice to open those files.
.ipynb74 KB · 94%
From the Hugging Face model README
This project implements a neural network-based language model for next token prediction in both English and Czech. The model uses an LSTM architecture and is trained on a combined dataset of English and Czech text.
The perplexity score of 1900.06 indicates that the model is performing reasonably well for a bilingual task, but there's room for improvement. Here's what this score means:
Bilingual Complexity: A perplexity of 1900.06 for a bilingual model is actually quite reasonable. Bilingual models typically have higher perplexity than monolingual models due to the increased complexity of handling two languages simultaneously.
Translation Challenges: The perplexity score suggests that while the model has learned patterns in both languages, it may struggle with precise translations or generating highly coherent text, especially in the less represented language (likely Czech in this case).
Comparison to Monolingual Models: For context, state-of-the-art monolingual models can achieve perplexities below 20, but these are much larger models trained on vast amounts of data.
Implications for Text Generation: With this perplexity, the model can generate text that follows general patterns of both languages but may produce some nonsensical or incorrect phrases, especially when attempting to switch between languages or translate.
The perplexity of 89.06 correlates with the observed issues in translation quality:
Vocabulary Limitations: The model may not have a comprehensive grasp of vocabulary in both languages, leading to incorrect word choices.
Contextual Understanding: A higher perplexity indicates that the model sometimes struggles to predict the next token accurately, which can result in contextually inappropriate words or phrases in the generated text.
Grammar and Structure: The model may not have fully captured the grammatical structures of both languages, especially Czech, which has a more complex grammar than English.
Language Mixing: In bilingual settings, the model might inadvertently mix elements from both languages, leading to nonsensical translations.
Data Imbalance: If one language (likely English) was more represented in the training data, the model's performance on the other language (Czech) could be compromised.
While the current model shows promise in bilingual text generation, the perplexity score of 1900.06 indicates that there's significant room for improvement, especially in translation accuracy and coherence. Future iterations of this project should focus on reducing perplexity to enhance the quality of generated text in both languages.