Downloads · 30 days
7
1% of all-time downloads
m-a-p/Amber-Reproduce-41.94B
Amber-Reproduce-41.94B is a machine learning model from m-a-p. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers.
Downloads · 30 days
7
1% of all-time downloads
All-time downloads
486
Public
Parameters
5.8B
11.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors11.6 GB · 100%
From the Hugging Face model README
Architecture & Training Configuration:
Base Model Configuration: This variant is built upon the Llama2-7B configuration, ensuring a robust foundation that aligns with the latest advancements in model architecture.
Sequence Length Adaptation: Originally processed data for a sequence length of 2048 was detokenized and re-encoded to fit a sequence length of 4096. This step follows the preprocessing strategy of Megatron-LM, enhancing our model's capacity to understand and generate more complex sequences.
Batch Size & Token Management: We adopted a batch size capable of managing 4 million tokens, tailored to accommodate the increased sequence length and ensure efficient data processing.
Integration of GQA Technologies: To boost training efficiency, our configuration includes the integration of Gradient Quantization and Aggregation technologies. With 32 attention heads and a group size of 4, this feature significantly enhances the model's learning and processing capabilities.