Downloads · 30 days
23
9% of all-time downloads
Aratako/LightChatAssistant-2x7B-optimized-experimental
LightChatAssistant-2x7B-optimized-experimental is a text generation model from Aratako. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
GGUF版はこちら/Click here for the GGUF version
Downloads · 30 days
23
9% of all-time downloads
All-time downloads
257
Public
Parameters
12.9B
25.8 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors25.8 GB · 100%
From the Hugging Face model README
GGUF版はこちら/Click here for the GGUF version
Sdff-Ltba/LightChatAssistant-2x7Bと同じ設定で、Chat Vectorの加算割合の最適化を目指したモデルです。 素晴らしい手法を提案いただいた@Sdff-Ltbaさんに感謝します。
元々のLightChatAssistant-2x7BではChat Vectorの加算割合が全レイヤで0.8と固定でした。 この加算割合の最適化をOptunaのTPESamplerを使って各レイヤごとで行い、その後MoEを行って作成したモデルです。 作成に利用したスクリプトは以下で公開しています。
https://github.com/Aratako/Task-Vector-Merge-Optimzier
最適化の大まかな流れは以下の通りです。
なお、experimentalと付いている通り、あくまで実験として作成したモデルです。主に下記の理由であまりパフォーマンスの向上が出来ていない可能性もあります。
などなど問題があるかと思いますので、あくまで実験的なものとしてご認識ください。
余談ですが、各ベースモデル単体に対して同じ流れで個別で最適化を行い、最適化されたモデルをMoEする方法だと出力がかなり悪化しました。MoE前提で最適化を行う場合は、MoEまでを全体フローに取り入れ、MoEを行ったモデルを利用した評価値で最適化したほうが良さそうです。