Downloads · 30 days
42
4% of all-time downloads
DeepAuto-AI/Explore_Llama-3.2-1B-Inst
Explore_Llama-3.2-1B-Inst is a text generation model from DeepAuto-AI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
DeepAutoAI/ExploreLlama-3.2-1B-Inst is developed by deepAuto.ai by learning the distribution of llama-3.2-1B-instruct. Our approach leverages the base model’s pretrained weights and optimizes them for the Winogrande a…
Downloads · 30 days
42
4% of all-time downloads
All-time downloads
1.2K
Public
Parameters
1.2B
2.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.5 GB · 99%
From the Hugging Face model README
DeepAutoAI/Explore_Llama-3.2-1B-Inst is developed by deepAuto.ai by learning the distribution of llama-3.2-1B-instruct. Our approach leverages the base model’s pretrained weights and optimizes them for the Winogrande and ARC-Challenge datasets by training a latent diffusion model on the pretrained weights. specifically , this model is based on learning the distrinution of transformer layers from 16 to 31.
Through this process, we learn the distribution of the base model's weight space, enabling us to explore optimal configurations. We then sample multiple sets of weights, using the model-soup averaging technique to identify the best-performing weights for both datasets. These weights are merged using linear interpolation to create the final model weights for DeepAutoAI/Explore_Llama-3.1-1B-Inst.
This approach has led to improved performance on previously unseen leaderboard tasks, all without any additional task-specific training.
The work is currently in progress
We trained a diffusion model to learn the distribution of subset of llama to enable generation weights that improve the performance. We generate task specific weights on winogrande and arc_challenge then transfer the best model for leaderboard benchmarking.
The direct use case of our work is o improve existing model performance as well as generating task specific weights with no training.
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->Performance improvement of existing large models with limited compute
No fine-tuning or architecture generalization
Using a generative model to produce weights can potentially lead to unintended or undesirable outputs. However, the generated content will still fall within the range of what the base model is inherently capable of producing.
The work is under progress
We employed a latent diffusion process on pretrained model weights, unlocking the ability to generate diverse, previously unseen neural networks. Remarkably, even within the constraints of one-shot learning, our approach consistently produces a wide range of weight variations, each offering distinct performance characteristics. These generated weights not only open opportunities for weight averaging and model merging but also have the potential to significantly enhance model performance. Moreover, they enable the creation of task-specific weights, tailored to optimize performance for specialized applications
The training data used to produced the current model is the base pretrained weights
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->We test our method on Winogrande and arc_challenge, and hellaswag
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
We used Latent diffusion for weights generation, and llama3-2-1B as target architectures.
The primary objective of this weight generation process was to demonstrate that by learning only the distribution of few layers weights (normlaization layers in this case) in an 1-billion-parameter model, it is possible to significantly enhance the model's capabilities. Notably, this is achieved using a fraction of the computational resources and without the need for fine-tuning, showcasing the efficiency and potential of this approach.
Nvidia-A100 cluster
A single Nvidia-A100
Model is tested using lm-harness tool version 0.4.3
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Detailed results can be found here
| Metric | Value |
|---|---|
| Avg. | 13.58 |
| IFEval (0-Shot) | 57.68 |
| BBH (3-Shot) | 8.31 |
| MATH Lvl 5 (4-Shot) | 4.53 |
| GPQA (0-shot) | 1.57 |
| MuSR (0-shot) | 1.09 |
| MMLU-PRO (5-shot) | 8.31 |