Downloads · 30 days
4
8% of all-time downloads
dhvanit2026/AgriGpt
AgriGpt is a machine learning model from dhvanit2026. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
A fine-tuned Qwen2.5-3B-Instruct model specialized for Indian and Gujarat agriculture, trained on Anand Agricultural University (AAU) knowledge. This is the merged single-file model (no LoRA adapter needed at inferenc…
Downloads · 30 days
4
8% of all-time downloads
All-time downloads
50
Public
Parameters
3.1B
6.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors6.2 GB · 100%
From the Hugging Face model README
A fine-tuned Qwen2.5-3B-Instruct model specialized for Indian and Gujarat agriculture, trained on Anand Agricultural University (AAU) knowledge. This is the merged single-file model (no LoRA adapter needed at inference time).
| Attribute | Value |
|---|---|
| Base Model | unsloth/Qwen2.5-3B-Instruct-bnb-4bit |
| Architecture | Qwen2ForCausalLM |
| Fine-tuning | LoRA (r=32, alpha=32, dropout=0) → merged & unloaded |
| Parameter Count | ~3.1B |
| Weight Format | FP16 safetensors |
| Weight Size | ~6.2 GB (model.safetensors) |
| Vocab Size | 151,936 |
| Languages | English + Gujarati |
| Domain | Agriculture (AAU crop varieties, FAQs, crop management, plant diseases) |
| Chat Template | Qwen2.5 (ChatML: `< |
qwen2Fine-tuned with Unsloth + TRL SFTTrainer on a curated agriculture instruction dataset:
combined_agri_dataset (~2,200 instruction examples)
dataset/train_full_model.pyq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_projAfter training, the LoRA weights were merged into the base model and unloaded, producing this standalone model directory (model_c_1).
Load with standard Hugging Face Transformers (no adapter or trust_remote_code flags needed beyond trust_remote_code=True):
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "model_c_1"
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.float16 if torch.cuda.is_available() else torch.float32,
device_map="auto" if torch.cuda.is_available() else None,
trust_remote_code=True,
)
model.eval()
messages = [
{"role": "system", "content": "You are an expert AI Agricultural Assistant for Indian and Gujarat agriculture (AAU data)."},
{"role": "user", "content": "Provide complete agricultural information for the AAU crop variety C-10-2."},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.7,
top_p=0.8,
top_k=20,
repetition_penalty=1.05,
pad_token_id=tokenizer.eos_token_id,
eos_token_id=tokenizer.eos_token_id,
)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
model_c_1/
├── README.md # This file
├── config.json # Model architecture config
├── generation_config.json
├── model.safetensors # Merged FP16 weights (~6.2 GB)
├── chat_template.jinja # Qwen2.5 ChatML template
├── tokenizer.json
└── tokenizer_config.json
config.py → MODEL_DIR = model_c_1.