Downloads · 30 days
18
20% of all-time downloads
flylcw/seq_monkey_pretrain_base_1.5B
seq_monkey_pretrain_base_1.5B is a text generation model from flylcw. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
基于 Qwen2.5-1.5B 架构、在出门问问「序列猴子」中文通用语料上 from-scratch(随机初始化)预训练 的中文 Base 模型。本模型为 Base(续写)模型,未经指令微调,适合下游继续预训练 / SFT,或直接做文本续写。
Downloads · 30 days
18
20% of all-time downloads
All-time downloads
91
Public
Parameters
1.8B
3.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.6 GB · 100%
From the Hugging Face model README
基于 Qwen2.5-1.5B 架构、在出门问问「序列猴子」中文通用语料上 from-scratch(随机初始化)预训练 的中文 Base 模型。本模型为 Base(续写)模型,未经指令微调,适合下游继续预训练 / SFT,或直接做文本续写。
| 项目 | 值 |
|---|---|
| 架构 | Qwen2ForCausalLM(Decoder-only,LLaMA 同族) |
| 参数量 | 1.5B 级(含 15 万词表 embedding,HF 统计约 2B) |
| hidden_size | 1536 |
| num_hidden_layers | 28 |
| num_attention_heads | 12 |
| num_key_value_heads | 2(GQA) |
| intermediate_size | 8960 |
| vocab_size | 151936 |
| max_position_embeddings | 131072 |
| 激活函数 | SiLU(SwiGLU) |
| 归一化 | RMSNorm(eps=1e-6) |
| 位置编码 | RoPE |
| 精度 | BF16 |
mobvoi_seq_monkey_general_open_corpus——序列猴子中文通用文本语料,约 1300 万份中文文本,来源涵盖网页、百科、书籍等,许可 Apache-2.0from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
name = "flylcw/seq_monkey_pretrain_base_1.5B"
tok = AutoTokenizer.from_pretrained(name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
name, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True)
prompt = "人工智能正在改变"
ids = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**ids, max_new_tokens=128, do_sample=True,
temperature=0.7, top_p=0.9, repetition_penalty=1.1)
print(tok.decode(out[0], skip_special_tokens=True))