Downloads · 30 days
19
0% of all-time downloads
cognAI/lil-c3po
lil-c3po is a text generation model from cognAI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
<div style="display: flex; justify-content: center; align-items: center;" <img src="./lil-c3po.jpg" style="width: 100%; height: auto;"/</div
Downloads · 30 days
19
0% of all-time downloads
All-time downloads
9.7K
Public
Parameters
7.2B
29 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors29 GB · 100%
From the Hugging Face model README
lil-c3po is an open-source large language model (LLM) resulting from the linear merge of two distinct fine-tuned Mistral-7B models, internally referred to as c3-1 and c3-2. These models, developed in-house, bring together unique characteristics to enhance performance and utility.
lil-c3po inherits its architecture from the combined c3-1 and c3-2 models, incorporating features such as Grouped-Query Attention, Sliding-Window Attention, and Byte-fallback BPE tokenizer. This fusion aims to capitalize on the strengths of both models for improved language understanding and generation.
lil-c3po is released under the MIT license, fostering open-source collaboration and innovation.
This merged model is suitable for a broad range of language-related tasks, inheriting the capabilities of the fine-tuned c3-1 and c3-2 models. Users interested in language tasks can leverage lil-c3po's capabilities.
While lil-c3po is versatile, it is important to note that, in most cases, fine-tuning may be necessary for specific tasks. Additionally, the model should not be used to intentionally create hostile or alienating environments for people.
Detailed results can be found here
| Metric | Value |
|---|---|
| Avg. | 68.03 |
| AI2 Reasoning Challenge (25-Shot) | 65.02 |
| HellaSwag (10-Shot) | 84.45 |
| MMLU (5-Shot) | 62.36 |
| TruthfulQA (0-shot) | 68.73 |
| Winogrande (5-shot) | 79.16 |
| GSM8k (5-shot) | 48.45 |