Downloads · 30 days
292
1% of all-time downloads
optimum/mistral-1.1b-testing
mistral-1.1b-testing is a text generation model from optimum. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
Downloads · 30 days
292
1% of all-time downloads
All-time downloads
25.6K
Public
Parameters
1.1B
4.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors4.4 GB · 100%
From the Hugging Face model README
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
it's based on the 3t base model (not chat tuned).
not extensively tested.
enjoy!
(model card is repeated due to open llm leaderboard length requirements)
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
it's based on the 3t base model (not chat tuned).
not extensively tested.
enjoy!
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
it's based on the 3t base model (not chat tuned).
not extensively tested.
enjoy!
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
it's based on the 3t base model (not chat tuned).
not extensively tested.
enjoy!
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
it's based on the 3t base model (not chat tuned).
not extensively tested.
enjoy!
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
it's based on the 3t base model (not chat tuned).
not extensively tested.
enjoy!
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
it's based on the 3t base model (not chat tuned).
not extensively tested.
enjoy!
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
it's based on the 3t base model (not chat tuned).
not extensively tested.
enjoy!
mistralized tinyllama since flash attention training on llama w/ flash-attn is buggy.
it's based on the 3t base model (not chat tuned).
not extensively tested.
enjoy!