Downloads · 30 days
0
Ganesh-Nadkarni/nl2sql-comparative-study-conit2026
nl2sql-comparative-study-conit2026 is a table question answering model from Ganesh-Nadkarni. Use it for the table question answering task on the model card, and read the license before you ship it in a product. It is set up for transformers.
A research project on Natural Language to SQL (NL2SQL) that explores multiple approaches for converting natural-language questions into SQL queries.
Downloads · 30 days
0
Access
Public
Updated Sep 4, 2026
Repo size
739 MB
Likes
0
Public
Click a slice to open those files.
.pkl363 MB · 49%
From the Hugging Face model README
A research project on Natural Language to SQL (NL2SQL) that explores multiple approaches for converting natural-language questions into SQL queries.
The project implements and evaluates four approaches:
The trained model artifacts are provided for research, experimentation, and reproducibility.
The goal of the project is to investigate different machine-learning and deep-learning approaches for translating natural-language database questions into SQL queries.
For example:
Natural Language:
What are the names of all students?
Generated SQL:
SELECT name FROM students;
The project compares traditional rule-based and machine-learning approaches with neural sequence-to-sequence architectures.
A rule-based NLP approach is included as a baseline.
It uses predefined patterns and templates to map natural-language questions to SQL queries.
| Evaluation Set | Exact Match | Template Match |
|---|---|---|
| Train | 95.5% | 97.5% |
| Familiar | 53.0% | 86.0% |
| Unseen | 53.0% | 87.0% |
The rule-based system serves primarily as a baseline for comparison.
A traditional machine-learning approach using TF-IDF vectorization followed by a Random Forest classifier.
RandomForestClassifiermodels/random_forest.pkl
The serialized model contains the TF-IDF vectorizer, classifier, label encoder, training examples, SQL templates, and related prediction artifacts.
A neural sequence-to-sequence model implemented using TensorFlow/Keras.
Natural Language Input
↓
Embedding
↓
Bidirectional LSTM Encoder
↓
Attention Mechanism
↓
LSTM Decoder
↓
Generated SQL
| Metric | Result |
|---|---|
| Final Training Accuracy | 83.54% |
| Final Validation Accuracy | 48.14% |
| Training Epochs | 12 |
The difference between training and validation accuracy indicates challenges in generalization to unseen examples.
models/lstm/
├── lstm_weights.weights.h5
├── lstm_meta.json
├── lstm_meta.pkl
├── history.json
├── nl_vocab.pkl
└── sql_vocab.pkl
A fine-tuned T5 Transformer model for Natural Language to SQL translation.
The model is based on the T5 architecture and is trained for sequence-to-sequence text generation.
Natural Language Question
↓
T5 Encoder
↓
Transformer
↓
T5 Decoder
↓
SQL Query
The final trained T5 model is stored in:
models/t5/
├── config.json
├── generation_config.json
├── model.safetensors
├── tokenizer.json
└── tokenizer_config.json
The model.safetensors file contains the trained model weights.
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
model_path = "Ganesh-Nadkarni/nl2sql-research"
tokenizer = AutoTokenizer.from_pretrained(
model_path,
subfolder="models/t5"
)
model = AutoModelForSeq2SeqLM.from_pretrained(
model_path,
subfolder="models/t5"
)
Generate SQL:
question = "What are the names of all students?"
inputs = tokenizer(
question,
return_tensors="pt"
)
outputs = model.generate(
**inputs,
max_length=128
)
sql = tokenizer.decode(
outputs[0],
skip_special_tokens=True
)
print(sql)
The exact input formatting may depend on the preprocessing/prompt format used during T5 training.
| Approach | Type | Main Technique |
|---|---|---|
| Rule-Based | Baseline | NLP Rules & Templates |
| Random Forest | Machine Learning | TF-IDF + Random Forest |
| LSTM | Deep Learning | BiLSTM + Attention + Seq2Seq |
| T5 | Transformer | T5 Seq2Seq |
The project demonstrates the progression from rule-based methods and traditional machine learning to neural sequence-to-sequence and transformer-based approaches.
The project uses the Spider Text-to-SQL dataset for Natural Language to SQL research.
Spider is designed for evaluating systems that translate natural-language questions into SQL queries across databases and schemas.
The task can be represented as:
Natural Language Question
+
Database Schema
↓
SQL Query
nl2sql_research/
│
├── models/
│ ├── lstm_model.py
│ ├── random_forest_model.py
│ ├── rule_based.py
│ └── t5_model.py
│
├── trained_models/
│ ├── random_forest.pkl
│ ├── rule_based.pkl
│ ├── nl_vocab.pkl
│ ├── sql_vocab.pkl
│ │
│ ├── lstm/
│ │ ├── lstm_weights.weights.h5
│ │ ├── lstm_meta.json
│ │ ├── lstm_meta.pkl
│ │ ├── history.json
│ │ ├── nl_vocab.pkl
│ │ └── sql_vocab.pkl
│ │
│ └── t5_final/
│ ├── config.json
│ ├── generation_config.json
│ ├── model.safetensors
│ ├── tokenizer.json
│ └── tokenizer_config.json
│
├── data/
│ └── splits/
│
└── utils/
The trained artifacts in this repository are organized as:
models/
├── random_forest.pkl
├── nl_vocab.pkl
├── sql_vocab.pkl
│
├── lstm/
│ ├── lstm_weights.weights.h5
│ ├── lstm_meta.json
│ ├── lstm_meta.pkl
│ ├── history.json
│ ├── nl_vocab.pkl
│ └── sql_vocab.pkl
│
└── t5/
├── config.json
├── generation_config.json
├── model.safetensors
├── tokenizer.json
└── tokenizer_config.json
This work is associated with a research paper published at IEEE CONIT 2026.
DOI:
https://doi.org/10.1109/CONIT69683.2026.11621699
IEEE Xplore:
https://ieeexplore.ieee.org/document/11621699
The trained models and supporting artifacts are provided to support research, experimentation, and reproducibility.
These models are research implementations and should not be considered production-ready SQL generation systems.
Important limitations include:
Generated SQL should not be executed directly on production databases without validation.
Recommended safeguards include:
This repository is intended for:
Potential improvements include:
Ganesh Nadkarni
GitHub:
https://github.com/GANESH-NADKARNI
Please refer to the original dataset and publication terms when using the dataset, research material, or derived artifacts.
The IEEE publication should be accessed through the official IEEE Xplore/DOI link provided above.
If you use this repository or build upon this work, please refer to the associated research publication:
DOI: 10.1109/CONIT69683.2026.11621699
IEEE Xplore: