Downloads ยท 30 days
21
35% of all-time downloads
shellwork/ChatParts-qwen2.5-14b
ChatParts-qwen2.5-14b is a question answering model from shellwork. Use it when the input is a question plus a passage. The card lists the license as apache-2.0.
๐ค XJTLU-Software RAG GitHub Repository โข ๐ ChatParts Dataset
Downloads ยท 30 days
21
35% of all-time downloads
All-time downloads
60
Public
Parameters
14.8B
29.5 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors29.5 GB ยท 100%
From the Hugging Face model README
๐ค XJTLU-Software RAG GitHub Repository โข ๐ ChatParts Dataset
shellwork/ChatParts-qwen2.5-14b is a specialized dialogue model fine-tuned from Qwen2.5-14B-Instruct by the XJTLU-Software iGEM Competition team. This model is tailored for the synthetic biology domain, aiming to assist competition participants and researchers in efficiently collecting and organizing relevant information. It serves as the local model component of the XJTLU-developed Retrieval-Augmented Generation (RAG) software, enhancing search and summarization capabilities within synthetic biology data.
The model is trained on a comprehensive synthetic biology-specific dataset curated from multiple authoritative sources:
In total, the dataset comprises over 200,000 question-answer pairs, meticulously assembled to cover a wide spectrum of synthetic biology topics. For more detailed information about the dataset, please visit our training data repository.
This repository supports usage with the transformers library. Below is a straightforward example of how to deploy the shellwork/ChatParts-qwen2.5-14b model using transformers.
Transformers Library: Ensure you have transformers version >= 4.43.0 installed. You can update your installation using:
pip install --upgrade transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load the tokenizer and model
model_name = "shellwork/ChatParts-qwen2.5-14b"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
# Define the prompt and messages
prompt = "Give me a short introduction to synthetic biology."
messages = [
{"role": "system", "content": "You are ChatParts, a model specialized in synthetic biology created by XJTLU-Software."},
{"role": "user", "content": prompt}
]
# Apply chat template
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
# Tokenize the input
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
# Generate the response
generated_ids = model.generate(
**model_inputs,
max_new_tokens=512
)
# Extract the generated tokens
generated_ids = [
output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]
# Decode the response
response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(response)
Import Libraries: Import the necessary libraries including torch, modelscope, and transformers.
Load Model and Tokenizer: Use AutoModelForCausalLM and AutoTokenizer from modelscope to load the pre-trained model and tokenizer.
Define Prompt and Messages: Create a prompt and define the conversation messages, including system and user roles.
Apply Chat Template: Utilize the apply_chat_template method to format the messages appropriately for the model.
Tokenize Input: Tokenize the formatted text and move it to the appropriate device (CPU/GPU).
Generate Response: Use the generate method to produce a response with a specified maximum number of new tokens.
Decode and Print: Decode the generated tokens to obtain the final text response and print it.
This model is released under the Apache License 2.0. For more details, please refer to the license information in the repository.
Feel free to reach out through our GitHub repository for any questions, issues, or contributions related to shellwork/ChatParts-qwen2.5-14b.