Downloads · 30 days
0
nozomuteruyo14/Diff_LoRA
Diff_LoRA is a text classification model from nozomuteruyo14. Use it when you need a label for a piece of text. The card lists the license as mit.
DiffLoRA is an innovative adapter architecture that extends conventional low-rank adaptation (LoRA) by fine-tuning a pre-trained large-scale model using differential low-rank matrices. Instead of updating all model pa…
Downloads · 30 days
0
Access
Public
Updated Feb 12, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
.txt71.2 KB · 75%
From the Hugging Face model README
DiffLoRA is an innovative adapter architecture that extends conventional low-rank adaptation (LoRA) by fine-tuning a pre-trained large-scale model using differential low-rank matrices. Instead of updating all model parameters, DiffLoRA updates only a small set of low-rank matrices, which allows for efficient fine-tuning with reduced trainable parameters.
DiffLoRA is an original method developed by the author and is inspired by the conceptual ideas from the Differential Transformer paper (https://arxiv.org/abs/2410.05258). It decomposes the weight update into two components—positive and negative contributions—enabling a more fine-grained adjustment than traditional LoRA. The output of a single layer is computed as:
$$ y = W x + \Delta y $$
where:
$$ x \in \mathbb{R}^{d_{in}} $$ is the input vector (or each sample in a batch).
$$ W \in \mathbb{R}^{d_{out} \times d_{in}} $$ is the fixed pre-trained weight matrix.
$$ \Delta y $$ is the differential update computed as:
$$ \Delta y = \frac{\alpha}{r} \Big( x' A_{\text{pos}} B_{\text{pos}} - \tau , x' A_{\text{neg}} B_{\text{neg}} \Big) $$
with:
$$ x' $$ being the input after dropout (or another regularization).
$$ A_{\text{pos}} \in \mathbb{R}^{d_{in} \times r} \quad \text{and} \quad B_{\text{pos}} \in \mathbb{R}^{r \times d_{out}} $$ capturing the positive contribution.
$$ A_{\text{neg}} \in \mathbb{R}^{d_{in} \times r} \quad \text{and} \quad B_{\text{neg}} \in \mathbb{R}^{r \times d_{out}} $$ capturing the negative contribution.
$$ \tau \in \mathbb{R} $$ is a learnable scalar that balances the two contributions.
$$ \alpha $$ is a scaling factor.
$$ r $$ is the chosen rank.
For computational efficiency, the two low-rank components are fused via concatenation:
$$ \text{combined_A} = \big[ A_{\text{pos}}, A_{\text{neg}} \big] \in \mathbb{R}^{d_{in} \times 2r} $$
$$ \text{combined_B} = \begin{bmatrix} B_{\text{pos}} \ -\tau , B_{\text{neg}} \end{bmatrix} \in \mathbb{R}^{2r \times d_{out}} $$
The update is then calculated as:
$$ \text{update} = x' \cdot \text{combined_A} \cdot \text{combined_B} $$
resulting in the final output:
$$ y = W x + \frac{\alpha}{r} , \text{update}. $$
DiffLoRA is intended to be integrated as an adapter module into pre-trained transformer models. It allows efficient fine-tuning by updating only a small number of low-rank parameters, making it ideal for scenarios where computational resources are limited.
DiffLoRA is not designed for training models from scratch, nor is it recommended for tasks where full parameter updates are necessary. It is optimized for transformer-based NLP tasks and may not generalize well to non-NLP domains. Also, there are only a limited number of base models that can be used.
While DiffLoRA offers a parameter-efficient fine-tuning approach, it inherits limitations from its base models (e.g., BERT, MiniLM). It may not capture all domain-specific nuances when only a limited number of parameters are updated. Users should carefully evaluate performance and consider potential biases in their applications.
Users should:
To integrate DiffLoRA into your fine-tuning workflow, check the example script in the examples/run_glue_experiment.py file.
This implementation has been demonstrated on GLUE tasks using the Hugging Face Datasets library.
DiffLoRA is applied by freezing the base model weights and updating only the low-rank adapter parameters. The procedure involves:
GLUE validation sets are used for evaluation.
Evaluations are performed across multiple GLUE tasks to ensure comprehensive performance analysis.
Evaluation metrics include accuracy, F1 score, Pearson correlation, and Spearman correlation, depending on the task.
For detailed evaluation results, please refer to the GLUE experiment script in the examples directory.
DiffLoRA achieves faster convergence and competitive performance on GLUE tasks compared to other parameter-efficient fine-tuning methods.
paper: Writing
For any questions regarding this model card, please contact: [[email protected]]