Downloads · 30 days
61
7% of all-time downloads
pranavupadhyaya52/Wiki-SmartBotLM-Instruct
Wiki-SmartBotLM-Instruct is a feature extraction model from pranavupadhyaya52. Use it when you need embeddings to search or compare text. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
61
7% of all-time downloads
All-time downloads
870
Public
Parameters
276M
1.7 GB on disk
Likes
13
Public
Click a slice to open those files.
.safetensors552 MB · 99%
From the Hugging Face model README

A 270M parameter decoder-only Small Language Model (SLM) pretrained on English Wikipedia and instruction-tuned for general-purpose conversational AI.
Wiki-SmartBotLM-Instruct is a compact autoregressive language model developed entirely from scratch as an independent research project. The model was designed to explore the complete lifecycle of modern language model development—from pretraining through instruction tuning—using open datasets and modern transformer architectures.
Unlike models that rely on continued pretraining of existing foundation models, Wiki-SmartBotLM-Instruct was initialized with random weights and trained using a multi-stage pipeline.
The primary goals of this project were:
| Property | Value |
|---|---|
| Model Name | Wiki-SmartBotLM-Instruct |
| Architecture | Decoder-only Transformer |
| Parameters | 270 Million |
| Context Length | 8192 Tokens |
| Training Framework | PyTorch |
| Precision | BF16 / FP16 |
| License | MIT |
Wiki-SmartBotLM-Instruct incorporates several modern transformer improvements including:
The model was trained in two stages.
The model was pretrained on the English Wikipedia corpus using next-token prediction.
Objectives:
Following pretraining, the model was instruction tuned using an open instruction-following dataset.
The objective of this stage was to improve:
Wiki-SmartBotLM-Instruct can perform a variety of natural language tasks including:
The model was evaluated on the following text-based benchmarks:
| Benchmark | WikiSmartBotLM | Gemma 3 270M IT | Gap |
|---|---|---|---|
| BoolQ | 37.8 | 61.4 (PT) ([Google AI for Developers][1]) | -23.6 |
| ARC Challenge | 24.2 | 28.2 ([Google AI for Developers][1]) | -4.0 |
| HellaSwag | 26.3 | 37.7 ([Google AI for Developers][1]) | -11.4 |
| PIQA | 53.5 | 66.2 ([Google AI for Developers][1]) | -12.7 |
| Winogrande | 50.2 | 52.3 ([Google AI for Developers][1]) | -2.1 |
It shows despite the training difference between Gemma3 270M and WikiSmartBotLM, It does come close to Gemma 3 in some benchmarks.
Prompt
Explain what artificial intelligence is.
Response
Artificial intelligence (AI) is a branch of computer science focused on building systems capable of performing tasks that typically require human intelligence. These tasks include reasoning, learning, planning, perception, natural language understanding, and decision making. Modern AI encompasses techniques such as machine learning, deep learning, computer vision, and reinforcement learning, and is widely used in applications ranging from recommendation systems to autonomous vehicles.
Wiki-SmartBotLM-Instruct is intended for:
As a compact language model, Wiki-SmartBotLM-Instruct has several limitations.
Wiki-SmartBotLM-Instruct is intended for research and educational purposes.
Although instruction tuning improves response quality, users should independently verify information before relying on generated content in high-stakes domains such as healthcare, finance, or legal advice.
Future versions of Wiki-SmartBotLM aim to include:
If you use Wiki-SmartBotLM-Instruct in your research, please cite:
@misc{wikismartbotlm2026,
title={Wiki-SmartBotLM-Instruct: A 270M Parameter Decoder-Only Small Language Model},
author={Pranav Upadhyaya},
year={2026},
howpublished={Hugging Face Model Repository}
}
This project was made possible by the open-source AI community.
Special thanks to:
Wiki-SmartBotLM-Instruct was developed as an independent research initiative to better understand modern language model development. From random initialization to pretraining and instruction tuning, every stage of the model was built to explore the practical engineering challenges of creating compact, efficient, and accessible language models.
Feedback, issues, and contributions are always welcome.