Downloads · 30 days
25
4% of all-time downloads
MYTH-Lab/BatGPT-15B-sirius
BatGPT-15B-sirius is a text generation model from MYTH-Lab. Use it when you need the model to write or continue text. It is set up for transformers.
Bidirectional Autoregressive Talker from Generative Pre-trained Transformer
Downloads · 30 days
25
4% of all-time downloads
All-time downloads
687
Public
Repo size
60.1 GB
Likes
5
Public
Click a slice to open those files.
.bin30.1 GB · 100%
From the Hugging Face model README
Bidirectional Autoregressive Talker from Generative Pre-trained Transformer
BatGPT-15B-sirius 是上海交通大学与武汉大学<font size=1>(或武汉大学与上海交通大学,排名不分先后)</font>联合自然语言处理团队设计、预训练、对齐的系列大型语言模型 BatGPT 中的一个开源可商用版本。 BatGPT系列模型中还包括BatGPT-30B-orion,BatGPT-70B-alhena,以及BatGPT-140B-menkalinan。
BatGPT-15B-sirius 包含 150 亿参数,在中英文 1T 语料上进行了预训练,在权威的中文和英文 benchmark 上均取得同不错的效果。BatGPT-15B-sirius 有如下几个特点:
BatGPT-15B-sirius is an open-source commercially available version of the series of large-scale language models BatGPT, designed, pretrained, and aligned by the joint natural language processing teams of Shanghai Jiao Tong University and Wuhan University <font size=1>(or Wuhan University and Shanghai Jiao Tong University, in no particular order)</font>.
The BatGPT series of models also include BatGPT-30B-orion, BatGPT-70B-alhena, and BatGPT-140B-menkalinan.
BatGPT-15B-sirius contains 15 billion parameters and has been pretrained on 1T Chinese and English corpora. It achieves excellent performance on authoritative Chinese and English benchmarks. BatGPT-15B-sirius has the following characteristics:
pip install protobuf transformers cpm_kernels torch>=2.0 streamlit sentencepiece accelerate deepspeed
如下是一个使用 BatGPT-15B-sirius 进行对话的示例:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("MLP-lab/BatGPT-15B-sirius", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("MLP-lab/BatGPT-15B-sirius", torch_dtype=torch.float16, trust_remote_code=True).cuda()
model = model.eval()
history = []
system_prompt = None # 你也可以指定系统提示
response, history = model.chat(tokenizer, "你好", history=history, system_prompt=system_prompt)
print(response)
response, history = model.chat(tokenizer, "介绍一下你自己", history=history, system_prompt=system_prompt)
print(response)
Here is an example of a conversation using BatGPT-15B-sirius:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("MLP-lab/BatGPT-15B-sirius", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("MLP-lab/BatGPT-15B-sirius", torch_dtype=torch.float16, trust_remote_code=True).cuda()
model = model.eval()
history = []
system_prompt = None # You can give a system prompt here.
response, history = model.chat(tokenizer, "Hello", history=history, system_prompt=system_prompt)
print(response)
response, history = model.chat(tokenizer, "Please introduce yourself", history=history, system_prompt=system_prompt)
print(response)
BatGPT-15B-sirius 具体参数和见下表:
| 模型名称 | 隐含层维度 | 层数 | Query头数 | Key/Value头数 | 词表大小 | 总参数量 | 训练数据(tokens) | 位置编码 | 最大长度 |
|---|---|---|---|---|---|---|---|---|---|
| BatGPT-15B-sirius | 5,632 | 48 | 44 | 2 | 65,536 | 15,030,081,024 | 1T | RoPE | 32K |
The specific parameters of BatGPT-15B-sirius are as follows:
| Model Name | Hidden Size | Num Layers | Query Heads | Key/Value Heads | Vocab Size | Total Params | Training Dats(tokens) | Position Embedding | Max Length |
|---|---|---|---|---|---|---|---|---|---|
| BatGPT-15B-sirius | 5,632 | 48 | 44 | 2 | 65,536 | 15,030,081,024 | 1T | RoPE | 32K |
BatGPT-15B-sirius 模型的使用应当遵循社会的公序良俗,不能被用于任何危害国家社会安全或违法的活动。另外,我们也要求使用者不要将 BatGPT-15B-sirius 模型用于未经适当安全审查和备案的互联网服务。我们希望所有的使用者都能遵守这个原则,确保科技的发展能在规范和合法的环境下进行。
我们已经尽我们所能,来确保模型训练过程中使用的数据的合规性。然而,尽管我们已经做出了巨大的努力,但由于模型和数据的复杂性,仍有可能存在一些无法预见的问题。如使用本项目所含模型及其修改版本提供服务产生误导性或有害性言论,造成不良影响,由服务提供方负责,与本项目无关。
The use of the BatGPT-15B-sirius model should adhere to societal norms and not be used for any activities that jeopardize national or social security or violate the law. Additionally, we also request users not to use the BatGPT-15B-sirius model for internet services that have not undergone appropriate security review and documentation. We hope that all users will abide by this principle to ensure that technological development occurs in a regulated and legal environment.
We have done our best to ensure the compliance of the data used during the model training process. However, despite our significant efforts, unforeseen issues may still arise due to the complexity of the model and data. If misleading or harmful statements are generated through the use of the models included in this project or their modified versions while providing services, the responsibility lies with the service provider and is not associated with this project.
如果你觉得我们的工作有帮助的话,请考虑引用我们的BatGPT论文:
If you find our work helpful, please consider citing our BatGPT paper:
@article{li2023batgpt,
title={BatGPT: A Bidirectional Autoregessive Talker from Generative Pre-trained Transformer},
author={Li, Zuchao and Zhang, Shitou and Zhao, Hai and Yang, Yifei and Yang, Dongjie},
journal={arXiv preprint arXiv:2307.00360},
year={2023}
}