Downloads · 30 days
23
10% of all-time downloads
myxy/recursive-compressor-2-7b
recursive-compressor-2-7b is a text generation model from myxy. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama2.
Downloads · 30 days
23
10% of all-time downloads
All-time downloads
229
Public
Parameters
2.7B
5.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors5.5 GB · 100%
From the Hugging Face model README
RecursiveCompressor アーキテクチャによる日英バイリンガル言語モデルの事前学習版です。 独自の再帰的圧縮機構によりトークン間の長距離依存をサブリニアな計算量で扱います。
elyza/ELYZA-japanese-Llama-2-7b-fast| パラメータ | 値 |
|---|---|
| d_model | 2048 |
| num_heads | 16 |
| d_ff | 6144 |
| chunk_size | 4 |
| compress_size | 1 |
| num_layers | 16 |
事前学習に以下の文書データセットを使用しました(対話データは含みません):
| データセット | 言語 | ライセンス |
|---|---|---|
wikimedia/wikipedia (20231101.ja) | 日本語 | CC BY-SA 4.0 / GFDL |
wikimedia/wikipedia (20231101.en) | 英語 | CC BY-SA 4.0 / GFDL |
hotchpotch/cc100-ja-documents | 日本語 | Common Crawl Terms of Use |
JeanKaddour/minipile | 英語 | MIT (The Pileの派生) |
本モデルは独自アーキテクチャ RecursiveCompressorLM を使用するため、
以下のリポジトリをクローンしてその中のクラス定義を読み込む必要があります:
リポジトリ: https://github.com/myxyy/RecursiveCompressorHF
git clone https://github.com/myxyy/RecursiveCompressorHF.git -b v1.0
cd RecursiveCompressorHF
uv sync
HuggingFaceの generate() メソッドに対応しています:
import torch
from transformers import AutoTokenizer, TextStreamer
from recursive_compressor_lm import RecursiveCompressorLM
model = RecursiveCompressorLM.from_pretrained(
"myxy/recursive-compressor-2-7b",
torch_dtype=torch.bfloat16,
).to("cuda").eval()
tokenizer = AutoTokenizer.from_pretrained("elyza/ELYZA-japanese-Llama-2-7b-fast")
prompt = "吾輩は猫である。"
input_ids = tokenizer.encode(prompt, return_tensors="pt").to("cuda")
output_ids = model.generate(
input_ids,
max_new_tokens=256,
do_sample=True,
temperature=0.8,
top_p=0.9,
streamer=TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True),
)
#print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
リポジトリ内の predict.py / predict_stream.py も同等の機能を提供します(インタラクティブREPL等)。
注: beam searchは未対応です(do_sample=Falseの貪欲生成、またはdo_sample=Trueの確率サンプリングを使用してください)。
Llama 2 Community License に従います。
トークナイザが Llama 2 派生であるため、本モデルもこのライセンスに従います。
TODO
A bilingual (Japanese / English) pretrained causal language model using the RecursiveCompressor architecture, which handles long-range token dependencies in sublinear computational complexity via recursive compression.
PreTrainedModel)elyza/ELYZA-japanese-Llama-2-7b-fast| Parameter | Value |
|---|---|
| d_model | 2048 |
| num_heads | 16 |
| d_ff | 6144 |
| chunk_size | 4 |
| compress_size | 1 |
| num_layers | 16 |
Pretrained on document datasets only (no dialogue data):
| Dataset | Language | License |
|---|---|---|
wikimedia/wikipedia (20231101.ja) | Japanese | CC BY-SA 4.0 / GFDL |
wikimedia/wikipedia (20231101.en) | English | CC BY-SA 4.0 / GFDL |
hotchpotch/cc100-ja-documents | Japanese | Common Crawl Terms of Use |
JeanKaddour/minipile | English | MIT (subset of The Pile) |
This model uses the custom RecursiveCompressorLM architecture, so you need
to clone the repository to import the class definitions:
Repository: https://github.com/myxyy/RecursiveCompressorHF
git clone https://github.com/myxyy/RecursiveCompressorHF.git -b v1.0
cd RecursiveCompressorHF
uv sync
The model supports HuggingFace's generate() method:
import torch
from transformers import AutoTokenizer, TextStreamer
from recursive_compressor_lm import RecursiveCompressorLM
model = RecursiveCompressorLM.from_pretrained(
"myxy/recursive-compressor-2-7b",
torch_dtype=torch.bfloat16,
).to("cuda").eval()
tokenizer = AutoTokenizer.from_pretrained("elyza/ELYZA-japanese-Llama-2-7b-fast")
prompt = "Once upon a time"
input_ids = tokenizer.encode(prompt, return_tensors="pt").to("cuda")
output_ids = model.generate(
input_ids,
max_new_tokens=256,
do_sample=True,
temperature=0.8,
top_p=0.9,
streamer=TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True),
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
See predict.py / predict_stream.py in the repository for an interactive REPL.
Note: beam search is not supported (use do_sample=False for greedy or do_sample=True for sampling).
Llama 2 Community License.
The tokenizer is derived from Llama 2, hence this model inherits the license.
TODO