Downloads · 30 days
13
6% of all-time downloads
SustcZhangYX/EnvGPT
EnvGPT is a machine learning model from SustcZhangYX. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
<div align="center" <img src="LOGO.PNG" width="450px" <h1 align="center"<font face="Arial"EnvGPT: Leveraging a Large Language Model for Environmental Science</font</h1
Downloads · 30 days
13
6% of all-time downloads
All-time downloads
219
Public
Parameters
8B
48.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
EnvGPT is the first domain-specific large language model tailored for environmental science tasks.
Environmental science presents unique challenges for LLMs due to its interdisciplinary nature. EnvGPT was developed to address these challenges by leveraging a domain-specific environmental science instruction dataset and benchmark.
The model was fine-tuned on this environmental science-specific instruction dataset, ChatEnv, through Supervised Fine-Tuning (SFT). The dataset contains a total token count of 107,197,329, highlighting its depth and comprehensiveness for environmental science tasks.
Download the model: EnvGPT
git lfs install
git clone https://huggingface.co/SustcZhangYX/EnvGPT
Here is a Python code snippet that demonstrates how to load the tokenizer and model and generate text using EnvGPT.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline
# 1. Set your local EnvGPT model path here
model_path = "YOUR_LOCAL_MODEL_PATH"
# 2. Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.bfloat16,
device_map="auto",
)
# 3. Build chat messages
messages = [
{"role": "system", "content": "You are an expert assistant in environmental science, EnvGPT. You are a helpful assistant."},
{"role": "user", "content": "What is the definition of environmental science?"},
]
# 4. Format the prompt using the chat template
# add_generation_prompt=True appends the assistant start token (e.g., <|assistant|>)
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
# 5. Initialize the text-generation pipeline
text_gen = pipeline(
"text-generation",
model=model,
tokenizer=tokenizer,
device_map="auto",
torch_dtype=torch.bfloat16,
return_full_text=False, # Only return the newly generated text
)
# 6. Generate the response
# do_sample=True enables sampling (stochastic decoding)
# top_p=0.6 applies nucleus sampling
# temperature=0.8 controls randomness
# max_new_tokens=4096 allows up to 4096 new tokens
outputs = text_gen(
text,
max_new_tokens=4096, # Up to 4096 new tokens
do_sample=True, # Enable sampling instead of greedy decoding
top_p=0.6, # Nucleus sampling parameter
temperature=0.8, # Sampling temperature
)
# 7. Print the assistant’s reply (without the original prompt)
print(outputs[0]["generated_text"])
This code demonstrates how to load the tokenizer and model from your local path, define environmental science-specific prompts, and generate responses using sampling techniques like top-p and temperature.
EnvGPT is fine-tuned based on the open-sourced LLaMA. We thank Meta AI for their contributions to the community.
This project is intended solely for academic research and exploration. Please note that, like all large language models, this model may exhibit limitations, including potential inaccuracies or hallucinations in generated outputs.
If you find our work helpful, please consider citing our research: "Fine-Tuning Large Language Models for Interdisciplinary Environmental Challenges":
@article{ZHANG2025100608,
title = {Fine-Tuning Large Language Models for Interdisciplinary Environmental Challenges},
journal = {Environmental Science and Ecotechnology},
pages = {100608},
year = {2025},
issn = {2666-4984},
doi = {https://doi.org/10.1016/j.ese.2025.100608},
url = {https://www.sciencedirect.com/science/article/pii/S2666498425000869},
author = {Yuanxin Zhang and Sijie Lin and Yaxin Xiong and Nan Li and Lijin Zhong and Longzhen Ding and Qing Hu}
}