Downloads · 30 days
9
24% of all-time downloads
Eugenios/qwen2.5-coder-1.5b-secure-codegen
qwen2.5-coder-1.5b-secure-codegen is a text generation model from Eugenios. Use it when you need the model to write or continue text. It is set up for peft.
Адаптер DoRA/LoRA для Qwen/Qwen2.5-Coder-1.5B-Instruct, дообученный под генерацию более безопасного кода и исправление уязвимостей. Обучение велось на парах «уязвимый код → безопасный код» на основе датасета Big‑Vul.
Downloads · 30 days
9
24% of all-time downloads
All-time downloads
37
Public
Repo size
29.3 MB
Likes
0
Public
Click a slice to open those files.
.safetensors17.9 MB · 61%
From the Hugging Face model README
Адаптер DoRA/LoRA для Qwen/Qwen2.5-Coder-1.5B-Instruct, дообученный под генерацию более безопасного кода и исправление уязвимостей.
Обучение велось на парах «уязвимый код → безопасный код» на основе датасета Big‑Vul.
Qwen/Qwen2.5-Coder-1.5B-Instructapply_chat_templateАдаптер предназначен для:
Примеры задач:
user_id безопасным образом».cursor.execute(f"...") — объясни, почему он небезопасен, и перепиши безопасно».from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
base_model_id = "Qwen/Qwen2.5-Coder-1.5B-Instruct"
adapter_id = "<your-username>/qwen2.5-coder-1.5b-secure-codegen" # замените на ваш репозиторий
tokenizer = AutoTokenizer.from_pretrained(base_model_id, trust_remote_code=True)
base = AutoModelForCausalLM.from_pretrained(
base_model_id,
torch_dtype=torch.bfloat16 if torch.backends.mps.is_available() else torch.float32,
device_map="auto" if torch.backends.mps.is_available() else None,
trust_remote_code=True,
)
model = PeftModel.from_pretrained(base, adapter_id)
model.eval()
device = next(model.parameters()).device
prompt = "Напиши на Python функцию, которая выполняет SQL-запрос по user_id безопасным образом."
messages = [{"role": "user", "content": prompt}]
chat_text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
enc = tokenizer(chat_text, return_tensors="pt").to(device)
with torch.no_grad():
out = model.generate(
**enc,
max_new_tokens=256,
do_sample=False,
pad_token_id=tokenizer.eos_token_id,
)
generated = out[0, enc["input_ids"].shape[-1]:]
print(tokenizer.decode(generated, skip_special_tokens=True))
Источник: bstee615/bigvul (Hugging Face)
Подготовка:
берутся пары func_src_before / func_src_after (уязвимый/исправленный код);
длина ограничена (≤4000 символов);
формируется instruction в стиле:
{
"messages": [
{
"role": "user",
"content": "Исправь этот код, устранив уязвимость. Верни только исправленный код.\n\n<уязвимый_код>"
},
{
"role": "assistant",
"content": "<безопасный_код>"
}
]
}
Размер: 2000 пар (train+val), train/test‑split 90/10 на уровне готовых примеров.
Qwen/Qwen2.5-Coder-1.5B-Instructpeft.LoraConfig(use_dora=True) (если поддерживается версией PEFT)transformers.Trainer)bf16 на Apple MPS (Mac), float32 на CPUgradient_accumulation_steps=4 (эффективный batch 8 примеров)warmup_ratio=0.05, далее линейный decayТренинг запускался через scripts/train_qlora_secure.py из этого репозитория.
Неформальное сравнение с базовой моделью Qwen/Qwen2.5-Coder-1.5B-Instruct на
ручных промптах показывает:
SQL по user_id
Базовая модель уже иногда использует параметризованные запросы, но ответы
бывают размыты и не всегда подчёркивают безопасный шаблон.
Дообученная модель стабильно генерирует код вида
cursor.execute("SELECT * FROM users WHERE id = ?", (user_id,)) и даёт
более фокусированное объяснение, почему это защищает от SQL‑инъекции.
Определение небезопасного кода
На промптах вида
cursor.execute(f"SELECT * FROM users WHERE id = {user_id}")
адаптер точнее указывает на SQL‑инъекцию и сразу предлагает безопасный
вариант с параметрами.
Общие свойства
Адаптер не меняет архитектуру базовой модели и добавляет всего ~0.3%
обучаемых параметров, поэтому:
Важно: адаптер не гарантирует полную безопасность генерируемого кода. Его следует использовать как помощника, а не как единственный источник истины, вместе с ручным ревью, статическим/динамическим анализом и существующими инструментами безопасности.
Рекомендуется:
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct library_name: peft pipeline_tag: text-generation tags:
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
Use the code below to get started with the model.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
BibTeX:
[More Information Needed]
APA:
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]