Downloads · 30 days
9
11% of all-time downloads
ABrain/Delta-NAS-DeepSeek-Coder-7B
Delta-NAS-DeepSeek-Coder-7B is a text generation model from ABrain. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as mit.
This is a fully merged model (LoRA weights merged into base) for DeepSeek-Coder-7B-Instruct-v1.5, fine-tuned for delta-based Neural Architecture Search (NAS) — generating novel PyTorch image-classification architectur…
Downloads · 30 days
9
11% of all-time downloads
All-time downloads
84
Public
Parameters
6.9B
15.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors15.8 GB · 100%
From the Hugging Face model README
This is a fully merged model (LoRA weights merged into base) for DeepSeek-Coder-7B-Instruct-v1.5, fine-tuned for delta-based Neural Architecture Search (NAS) — generating novel PyTorch image-classification architectures via unified code diffs.
This adapter is the result of 22 iterative fine-tuning cycles on the delta-NAS pipeline described in "Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs". The model generates unified diffs that modify a baseline neural network architecture to produce new, functional PyTorch models.
deepseek-ai/deepseek-coder-7b-instruct-v1.5Models were evaluated on 6 LEMUR image-classification benchmarks:
| Metric | Value |
|---|---|
| Trained candidates | 828 |
| Valid rate (compiles + trains) | 49.5% |
| Mean 1-epoch accuracy | 33.9% (±7.9% SD across cycles) |
| ≥40% accuracy rate | 16.6% |
| Novel architectures admitted to LEMUR | 83 |
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
# Load base model
base_model = AutoModelForCausalLM.from_pretrained(
"deepseek-ai/deepseek-coder-7b-instruct-v1.5",
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("deepseek-ai/deepseek-coder-7b-instruct-v1.5")
# Load LoRA adapter
model = PeftModel.from_pretrained(base_model, "ABrain/Delta-NAS-DeepSeek-Coder-7B")
# Generate a diff to modify a baseline architecture
prompt = """Given the following PyTorch neural network baseline:
[baseline code here]
Generate a unified diff that creates a novel architecture variant."""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
@article{deltanas2026,
title={Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs},
author={Adhikari, Santosh and Ignatov, Dmitry},
year={2026}
}
MIT License (same as the base model and LEMUR dataset)