Downloads · 30 days
0
a-01a/QSolver_Decoder_V16
QSolver_Decoder_V16 is a text generation model from a-01a. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as mit.
QSolverDecoderV16 is a fine-tuned causal language model designed for context-driven scientific multiple-choice question answering. It utilizes 5-Fold Cross-Validation, Unsloth 4-bit quantization, and Low-Rank Adaptati…
Downloads · 30 days
0
Access
Public
Updated Jul 31, 2026
Repo size
4 GB
Likes
0
Public
Click a slice to open those files.
.pt2.6 GB · 49%
From the Hugging Face model README
QSolver_Decoder_V16 is a fine-tuned causal language model designed for context-driven scientific multiple-choice question answering. It utilizes 5-Fold Cross-Validation, Unsloth 4-bit quantization, and Low-Rank Adaptation (LoRA) on top of the base model Qwen/Qwen3-4B-Instruct-2507.
The model takes a context, question prompt, and five multiple-choice options (A, B, C, D, E), and ranks option token logits to produce predictions evaluated via Mean Average Precision at 3 (MAP@3).
dahaludba/QSolver_Decoder_V16Qwen/Qwen3-4B-Instruct-2507dahaludba/QSolver_TrainFastLanguageModel)The repository contains adapter checkpoints trained across 5 folds (fold_1 to fold_5). Each fold was trained using process isolation across available GPUs with dynamic memory management.
-100 so loss was calculated exclusively on completion target tokens.| Hyperparameter | Value |
|---|---|
| Base Model Quantization | 4-bit (BitsAndBytes / Unsloth) |
| LoRA Rank ($r$) | 32 |
| LoRA Alpha ($\alpha$) | 64 |
| LoRA Dropout | 0.05 |
| Target Modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Bias Term | none |
| Gradient Checkpointing | unsloth |
| Learning Rate | 1e-4 |
| Optimizer | AdamW |
| Learning Rate Schedule | Warmup Linear Decay |
| Warmup Ratio | 0.05 |
| Weight Decay | 0.01 |
| Per Device Train Batch Size | 4 |
| Per Device Eval Batch Size | 4 |
| Gradient Accumulation Steps | 4 (Effective Batch Size = 16) |
| Training Epochs | 4 per fold |
| Data Collator | DataCollatorForSeq2Seq (pad_to_multiple_of=8) |
| Evaluation Strategy | Epoch-based |
| Best Model Metric | MAP@3 (greater_is_better=True) |
| Seed | 42 |
The total training across all 5 folds completed in 18 hours and 45 minutes. Individual run logs and metrics can be reviewed via the following Weights & Biases experiment links:
Each input sample follows standard chat templates formatted as:
<|im_start|>system
You are a scientific expert. Base your answer STRICTLY on the provided Context. Output ONLY the single letter corresponding to the correct option (A, B, C, D, or E).<|im_end|>
<|im_start|>user
Context: {context}
Question: {question}
A) {option_a}
B) {option_b}
C) {option_c}
D) {option_d}
E) {option_e}<|im_end|>
<|im_start|>assistant
{answer}<|im_end|>
Performance is measured using Mean Average Precision at 3 (MAP@3):
$$\text{MAP@3} = \frac{1}{U} \sum_{i=1}^{U} \sum_{k=1}^{\min(P, 3)} P(k) \times \text{rel}(k)$$
Where:
Below is an example script to load a fold adapter and run inference:
import torch
from unsloth import FastLanguageModel
from peft import PeftModel
MODEL_REPO = "dahaludba/QSolver_Decoder_V16"
FOLD_SUBFOLDER = "fold_1"
MAX_SEQ_LENGTH = 1024
# Load Base Model & Tokenizer
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="Qwen/Qwen3-4B-Instruct-2507",
max_seq_length=MAX_SEQ_LENGTH,
dtype=None,
load_in_4bit=True,
)
# Load PEFT Fold Adapter
model = PeftModel.from_pretrained(model, MODEL_REPO, subfolder=FOLD_SUBFOLDER)
FastLanguageModel.for_inference(model)
# Construct Input Prompt
prompt = (
"<|im_start|>system\n"
"You are a scientific expert. Base your answer STRICTLY on the provided Context. "
"Output ONLY the single letter corresponding to the correct option (A, B, C, D, or E).<|im_end|>\n"
"<|im_start|>user\n"
"Context: Mitochondria generate most of the chemical energy needed to power the cell's biochemical reactions.\n"
"Question: What organelle produces most cellular energy?\n"
"A) Nucleus\nB) Mitochondria\nC) Ribosome\nD) Golgi Apparatus\nE) Endoplasmic Reticulum<|im_end|>\n"
"<|im_start|>assistant\n"
)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=2, use_cache=True)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print("Predicted Option:", response.strip())
This project and all model weight artifacts in this repository are distributed under the MIT License.