Downloads · 30 days
10
20% of all-time downloads
nbrahme/IndusQ
IndusQ is a machine learning model from nbrahme. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as osl-3.0.
Project Indus LLM is a groundbreaking open-source language model tailored for Hindi and its dialects, designed to enhance natural language processing and generation across diverse Indian linguistic applications.
Downloads · 30 days
10
20% of all-time downloads
All-time downloads
49
Public
Parameters
1.2B
4.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.4 GB · 100%
From the Hugging Face model README
Project Indus LLM is a groundbreaking open-source language model tailored for Hindi and its dialects, designed to enhance natural language processing and generation across diverse Indian linguistic applications.
Project Indus LLM aims to provide a robust language model for Indian languages, starting with Hindi and its dialects. This open-source foundational model, hosted on Hugging Face, is tailored for easy integration and further development by researchers and developers focusing on Indian linguistic diversity.
<!-- Provide a longer summary of what this model is/does. -->The model is a pretrained model in Hindi and dialects which is instruct tuned.
Uses include question and answeting and conversation in Hindi and Dialects. The model would be reward tuned to be used across various industries
Project Indus can be directly used for generating text, simulating conversation, and other text generation tasks without additional training.
Uses include question and answeting and conversation in Hindi and Dialects. The model would be reward tuned to be used across various industries
Project Indus is not designed for high-stakes decision-making tasks such as medical diagnosis or legal advice, nor can it be used for fill-in-the-blank exercises, multiple Q&A, and similar applications at the moment.
Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Predictions generated by the model may include disturbing and harmful stereotypes across protected classes; identity characteristics; and sensitive, social, and occupational groups. We have taken care across various biases by trying to remove them from training data. However since the model is a generative model, it would tend to produce hallucinations. Any disturbing or harmful sterotype produced by the model is purely un-intentional and coincidental.
It is recommended to avoid biases and negative connotations in the model, and regular updates along with community feedback are crucial for addressing any emergent bias or misuse scenarios.
The model was trained on a curated dataset comprising various sources of Hindi text, including literature, news articles, and web content.
The Project Indus LLM was trained on a diverse and extensive dataset comprising various sources of Hindi text and its dialects. The data collection and curation process was meticulously designed to cater to the linguistic diversity and complexity of Indian languages, particularly focusing on Hindi and its 37 dialects.
Data was collected in three main buckets:
Open-Source Hindi Data: This included publicly available sources from the internet across different categories such as news, and non-news. Automated scripts were used to scrape and extract text from web pages. Here are some of the sources:
Translated Data: A portion of the Pile dataset, which is a large English dataset used for training AI models, was translated into Hindi using three different translation models. IndicTrans2 (AI4Bharat) was selected as the best model for this purpose based on its accuracy and efficiency.
Dialects: Data collection for dialects presented a unique challenge due to the limited material available on the internet. Data for major dialects like Maithili, Bhojpuri, Magahi, and Braj Bhasha was collected from multiple sources, including fieldwork where representatives collected old books and other texts, which were then digitized and converted into text data.
Training involved extensive preprocessing to clean and standardize the text, followed by supervised learning on a high-performance computing setup.
Below is a table summarizing the datasets used for pre-training and fine-tuning the model:
| Phase | Data Source | Tokens | Notes |
|---|---|---|---|
| Pre-training | Cleaned dataset of Hindi and dialects | 22 billion | Utilized advanced tokenization |
| Fine-tuning | Custom datasets tailored for Indian languages | Varied | Focus on cultural, political, and social contexts |
The collected data underwent several stages of cleaning and preprocessing to ensure high quality and usability for training:
The final dataset used for training consisted of:
This diverse and extensive training data foundation allowed Project Indus LLM to develop robust capabilities for understanding and generating Hindi text, making it a powerful tool for applications requiring Indian language processing.
Project Indus LLM has been evaluated using the Indic LLM Leaderboard, which employs the indic_eval evaluation framework specifically designed for assessing models on Indian language tasks. This framework provides a comprehensive view of model performance across a variety of benchmarks tailored to Indian languages.
These results highlight the model's capabilities in understanding and generating Hindi language text under controlled testing conditions. The standard error values indicate the variance observed during the evaluation, providing insights into the consistency of the model's performance across different evaluation runs.
Additionally, Project Indus LLM has been evaluated on the Open LLM Leaderboard, which provides another layer of benchmarking by comparing the model's performance against other state-of-the-art language models. Below are the summarized results from the Open LLM Leaderboard:
These benchmark results can be explored further on Hugging Face Open LLM Leaderboard.
The evaluation metrics acc (accuracy) and acc_norm (normalized accuracy) are used to quantify the model's performance. The tasks are differentiated by their difficulty and the specific dataset used, such as the ARC Challenge and ARC Easy sets, both adapted to Hindi language conditions to ensure relevant assessment. This structured evaluation ensures that the Indus LLM not only performs well in generalized text generation tasks but also in more specialized, context-specific scenarios pertinent to the Indian linguistic framework.
Project Indus demonstrates competitive performance, particularly in text generation tasks, as evidenced by its scores on standardized benchmarks.
Project Indus LLM is based on a GPT-2.0-like architecture, tailored to handle the complexities of the Hindi language and its dialects. This model was designed to serve as a foundational model that can be fine-tuned for various applications, making it highly versatile and adaptable to different domains within the Indian context.
The objective of this model is to provide a robust tool for text generation and understanding in Hindi and its dialects, supporting the development of applications that require natural language processing in these languages. It also aims to bridge the gap in technology where Indian languages are underrepresented, providing a platform for further linguistic research and technological inclusion.
The pre-training and fine-tuning of Project Indus LLM were conducted on high-performance computing infrastructure provided by the Centre for Development of Advanced Computing (CDAC). This setup included:
Inference performance was tested on GPU as well as CPU.
The software environment was crucial for efficiently training and running the model. Key components included:
The detailed citation information will help in acknowledging the work and efforts of the team behind Project Indus LLM when it is used or referenced in academic or professional settings.
@article{malhotra2024projectindus,
title={Project Indus: A Foundational Model for Indian Languages},
author={Malhotra, Nikhil and Brahme, Nilesh and Mishra, Satish and Sharma, Vinay},
journal={Tech Mahindra Makers Lab},
year={2024},
url={https://www.techmahindra.com/en-in/innovation/the-indus-project/}
}
APA: Malhotra, N., Brahme, N., Mishra, S., & Sharma, V. (2024). Project Indus: A Foundational Model for Indian Languages. Tech Mahindra Makers Lab. Available at https://www.techmahindra.com/en-in/innovation/the-indus-project/
This glossary section explains key terms used throughout the model documentation and technical details, helping users unfamiliar with certain concepts to better understand the content.
For further details on Project Indus LLM, including additional documentation, tutorials, and community discussions, visit the following resources:
The model card and documentation for Project Indus LLM were collaboratively authored by:
For inquiries, support, or further information regarding Project Indus LLM, please reach out through the following channels:
To begin using Project Indus LLM for your projects, follow these steps to set up and run the model:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("nickmalhotra/ProjectIndus")
tokenizer = AutoTokenizer.from_pretrained("nickmalhotra/ProjectIndus")
# Example inference
def format_template(user_prompt):
messages = [
{"role": "user", "content": user_prompt},
]
response = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt")
return response
user_prompt = """भारत के वर्तमान प्रधानमंत्री कौन हैं?"""
input_ids = format_template(user_prompt)
# Generate text using the model
output = model.generate(input_ids,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.eos_token_id,
max_length=1024,
num_beams=5,
do_sample=True,
early_stopping=True,
temperature=0.7,
top_k=50,
top_p=0.95,
repetition_penalty=1.2,
no_repeat_ngram_size=3,
num_return_sequences=1,
)
print(tokenizer.decode(output[0], skip_special_tokens=False))
Project Indus LLM is trained with single instruction tuning, which may result in hallucinations—instances where the model generates plausible but inaccurate information. Users should exercise caution, especially in scenarios requiring high factual accuracy.
Project Indus LLM is designed as a foundational model suitable for further development and fine-tuning. Users are encouraged to adapt and refine the model to meet specific requirements of their applications.
This disclaimer aims to provide users with a clear understanding of the model's capabilities and limitations, facilitating its effective application and development.