Downloads · 30 days
10
48% of all-time downloads
yyymk/WLLM-7B-SPS
WLLM-7B-SPS is a machine learning model from yyymk. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
basemodel: deepseek-ai/DeepSeek-R1-Distill-Qwen-7B libraryname: unsloth
Downloads · 30 days
10
48% of all-time downloads
All-time downloads
21
Public
Repo size
14.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors14.3 GB · 100%
From the Hugging Face model README
base_model: deepseek-ai/DeepSeek-R1-Distill-Qwen-7B library_name: unsloth
Model Card for DeepSeek-R1-7B-COT-SPS-EN
This model, fine-tuned on the DeepSeek-R1-7B base model, specializes in predicting TBs in SPS scenarios. The model has been trained to handle complex reasoning tasks, including sequence prediction and resource allocation in communication networks. It uses a combination of complex reasoning and contextual understanding to provide accurate predictions for V2X applications.
This model can be used directly to predict TBs in V2X communication scenarios. It can analyze input data (e.g. network conditions) and output predictions related to resource allocation and usage in simulated environments.
The following code allows you to easily integrate the model into your environment and ask it questions.
from transformers import AutoModelForCausalLM, AutoTokenizer
from unsloth import FastLanguageModel
model_path = "/model_path/"
tokenizer = AutoTokenizer.from_pretrained(model_path)
device = "cuda:0"
model = AutoModelForCausalLM.from_pretrained(
model_path,
device_map="auto",
torch_dtype="auto"
).to(device)
model
tokenizer
FastLanguageModel.for_inference(model)
prompt_style_chat = """Please provide a rigorous answer with mathematical reasoning logic to complete the current conversation task. Before answering, please think carefully about the question and proceed with reasoning step by step to ensure the answer is logical and accurate.
### Instruction:
You are an expert with reasoning ability. Please answer the following question according to the 3GPP specification.
### Question:
{}
### Response:
<think>{}"""
questions = [
"TaskX: XXXX."
]
import time
inputs_list = [tokenizer([prompt_style_chat.format(q, "")], return_tensors="pt") for q in questions]
inputs = [input_data.to(device) for input_data in inputs_list]
outputs = []
for input_data in inputs:
output = model.generate(
input_ids=input_data.input_ids,
max_new_tokens=2048,
attention_mask=input_data.attention_mask,
use_cache=True,
)
outputs.append(output)
responses = [tokenizer.batch_decode(output, skip_special_tokens=True) for output in outputs]
for response in responses:
print(response)
The model was trained on 576 labeled examples of V2X communication scenarios with different values of c and ρ.
The raw dataset comes from the open-access repository: QiangFuNWPU/Rawdataset_WLLM.
From this dataset, we processed the power patterns measured at TBs during Alice’s transmission. Based on these patterns, we defined three rules to construct the training set:
C, ρ). To improve accuracy and stability, we distilled the CoT reasoning ability of a base LLM and created a tailored dataset.
Training sequences are constrained to 2048 tokens to ensure efficient and coherent reasoning.Using the dataset from Training Data, the base model's CoT is fine-tuned by Walter through the Low-Rank Adaptation (LoRA) method. LoRA is an efficient fine-tuning technique, with its primary advantage being the lightweight adaptation of the model via low-rank matrix factorization. Specifically, LoRA inserts two low-rank matrices, W1 and W2, in parallel with the original weight matrix, W. W1 reduces the input dimension, while W2 expands it back to the original size. During training, the original parameters of the LLM are frozen, and only W1 and W2 are updated. As a result, the training parameters account for just 0.1% to 1% of the total parameters. This significantly reduces the training cost and enables execution on consumer-grade hardware. After the fine-tuning process, the base model is transformed into WLLM, which can now perform CoT reasoning based on the three rules from Step I. LoRA is applied to all linear layers in the base model, resulting in a substantial performance boost for WLLM.
Training regime: Mixed precision (fp16) The model was trained using mixed precision with fp16 (half-precision floating-point) arithmetic. This allows for faster training by reducing memory usage and improving computational efficiency, while still maintaining model accuracy. Mixed precision training is particularly beneficial when working with large models, as it reduces the memory footprint and accelerates training on GPUs, especially on consumer-grade hardware.
Optimizer: AdamW The AdamW optimizer was used for fine-tuning, which is a variant of the Adam optimizer with weight decay regularization. It is known for its stability and effectiveness in training large language models like the one used in this experiment.
Optimization method: AdamW_8bit For further optimization, an 8-bit version of AdamW was used. This reduces memory usage during training, providing further efficiency without compromising the performance of the optimization process.
Learning rate: 2e-4 A learning rate of 2e-4 was used for training. The learning rate was selected to ensure smooth convergence without causing instability during the optimization process.
Batch size: 2 The batch size was set to 2 per device, with gradient accumulation steps of 4. This strategy helps to simulate larger batch sizes on GPUs with limited memory.
Gradient accumulation steps: 4 The gradient accumulation technique was used to effectively increase the batch size without requiring additional memory. This allows for more stable gradient updates, especially when working with large models.
Epochs: 2 The model was fine-tuned over 2 epochs. Given the size of the model and the dataset, this number of epochs was sufficient for the model to learn the necessary task-specific patterns without overfitting.
Warmup steps: 5 A warmup period of 5 steps was used to gradually increase the learning rate from 0 to the initial learning rate (2e-4). This helps to avoid the risk of large, unstable updates early in the training process.
Weight decay: 0.01 Weight decay of 0.01 was applied to prevent overfitting by regularizing the weights during training. This helps to improve generalization performance by penalizing overly large weights.
FP16 mixed precision: Enabled Since the model was trained on consumer-grade hardware, mixed-precision training with fp16 was enabled to save memory and accelerate the training process. This significantly reduces the memory consumption, enabling the training of larger models.
Learning rate scheduler: Linear decay A linear learning rate scheduler was used, which gradually decreases the learning rate from the initial value to 0 over the course of training. This helps to stabilize the training process as the model approaches convergence.
Random seed: 3407 The random seed 3407 was used to ensure reproducibility of the results. Using a fixed random seed allows the training process to be repeated with the same results, which is essential for model evaluation and comparison.