Downloads · 30 days
23
21% of all-time downloads
Aratako/LightChatAssistant-4x7B-optimized-experimental
LightChatAssistant-4x7B-optimized-experimental is a text generation model from Aratako. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
GGUF版はこちら/Click here for the GGUF version
Downloads · 30 days
23
21% of all-time downloads
All-time downloads
108
Public
Parameters
24.2B
48.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors48.3 GB · 100%
From the Hugging Face model README
GGUF版はこちら/Click here for the GGUF version
Aratako/LightChatAssistant-2x7B-optimized-experimentalと同じようなChat Vectorの加算割合の最適化を、 Aratako/LightChatAssistant-4x7Bの設定をもとに行ったモデルです。 素晴らしい手法を提案いただいた@Sdff-Ltbaさんに感謝します。
元々のLightChatAssistant-4x7BではChat Vectorの加算割合が全レイヤで0.8と固定でした。 この加算割合の最適化をOptunaのTPESamplerを使って各レイヤごとで行い、その後MoEを行って作成したモデルです。 作成に利用したスクリプトは以下で公開しています。
https://github.com/Aratako/Task-Vector-Merge-Optimzier
最適化の大まかな流れは以下の通りです。
なお、experimentalと付いている通り、あくまで実験として作成したモデルです。主に下記の理由であまりパフォーマンスの向上が出来ていない可能性もあります。
などなど問題があるかと思いますので、あくまで実験的なものとしてご認識ください。
余談ですが、各ベースモデル単体に対して同じ流れで個別で最適化を行い、最適化されたモデルをMoEする方法だと出力がかなり悪化しました。MoE前提で最適化を行う場合は、MoEまでを全体フローに取り入れ、MoEを行ったモデルを利用した評価値で最適化したほうが良さそうです。