Downloads · 30 days
4
15% of all-time downloads
zhaospei/Model_19
Model_19 is a machine learning model from zhaospei. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Đây là mô hình causal language model dành cho code với 1.3 tỷ tham số, được huấn luyện từ đầu trên 2 nghìn tỷ token (87% mã nguồn, 13% văn bản Anh–Hoa). Mô hình hỗ trợ hoàn thiện code quy mô dự án với context window l…
Downloads · 30 days
4
15% of all-time downloads
All-time downloads
27
Public
Repo size
5.4 GB
Likes
0
Public
Click a slice to open those files.
.bin2.7 GB · 100%
From the Hugging Face model README
Đây là mô hình causal language model dành cho code với 1.3 tỷ tham số, được huấn luyện từ đầu trên 2 nghìn tỷ token (87% mã nguồn, 13% văn bản Anh–Hoa). Mô hình hỗ trợ hoàn thiện code quy mô dự án với context window lên đến 16.000 token, tích hợp cả tác vụ fill‑in‑the‑blank để hỗ trợ infilling và completion trên mức file/repo.
Project-level code completion và code infilling nhờ window size lớn (16 K). Hiệu năng hàng đầu trong các benchmark mở: HumanEval, MultiPL-E, MBPP, DS-1000, APPS Hỗ trợ nhiều ngôn ngữ lập trình, khả năng tốt trong cả text kĩ thuật Anh–Hoa.
pip install torch transformers
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
tokenizer = AutoTokenizer.from_pretrained("zhaospei/Model_19", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("zhaospei/Model_19", trust_remote_code=True).cuda()
input_text = "# write a quick sort in Python\n"
inputs = tokenizer(input_text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_length=200)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
tokenizer = AutoTokenizer.from_pretrained("zhaospei/Model_19", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("zhaospei/Model_19", trust_remote_code=True).cuda()
prompt = """<|fim_begin|>def quick_sort(arr):\n if len(arr) <= 1:\n return arr\n pivot = arr[0]\n <|fim_hole|>\n"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Thông số Giá trị Tham số ~1.3 B Window Size 16 K tokens Dữ liệu huấn luyện 2T tokens (87% code, 13% văn bản) Kiến trúc Causal Transformer-model Danh mục benchmark HumanEval, MultiPL-E, MBPP, DS‑1000, APPS