Downloads · 30 days
0
shkbd/sheikh-bangla-coder-350m
sheikh-bangla-coder-350m is a machine learning model from shkbd. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A project to build a specialised 350 million parameter causal language model focused on Bangla text generation and coding assistance. The goal of this repository is to provide a full end‑to‑end pipeline for collecting…
Downloads · 30 days
0
Access
Public
Updated Sep 24, 2025
Repo size
20.4 KB
Likes
0
Public
Click a slice to open those files.
.py30.9 KB · 52%
From the Hugging Face model README
A project to build a specialised 350 million parameter causal language model focused on Bangla text generation and coding assistance. The goal of this repository is to provide a full end‑to‑end pipeline for collecting data, training a custom tokenizer, preparing a dataset that mixes Bangla and English code, training a transformer model, evaluating it on both natural language and code generation tasks, and finally packaging and uploading the resulting artefacts to HuggingFace.
Because this environment does not have network access to download packages or scrape live websites, the scripts in this repository are written as templates and are not executed here. When run in a proper environment with Python dependencies installed (transformers, datasets, tokenizers, torch, sentencepiece, beautifulsoup4, etc.), the scripts will perform the following tasks:
datasets.Dataset for training the language model. Special tokens such as <code>/</code> and <bangla>/</bangla> mark code blocks and Bangla passages.transformers.Trainer or a custom loop with fully sharded data parallel (FSDP) and mixed precision. The model is trained in phases: base language modelling, instruction tuning, and fine‑tuning for code‑switching.The repository follows a clear directory structure:
sheikh-bangla-coder/
├── data/
│ ├── raw/ # unprocessed scraped text and code
│ ├── processed/ # cleaned and deduplicated corpus
│ └── tokenizer/ # tokenizer training files and vocabulary
├── training/
│ ├── scripts/ # Python scripts for training and preprocessing
│ ├── configs/ # model and training configuration files
│ └── checkpoints/ # model checkpoints
├── evaluation/ # evaluation scripts and metrics
└── deployment/
└── model_card.md # documentation and HuggingFace metadata
To get started, create a Python virtual environment, install the required dependencies, and adapt the scraping logic for your environment. Once the data has been collected and processed, run the tokenizer training script followed by the model training script. Finally, evaluate your model and use the deployment script to upload it to HuggingFace.