Downloads · 30 days
47
100% of all-time downloads
doryno/NeuroObfuscator-ai
NeuroObfuscator-ai is a text generation model from doryno. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
A fine-tuned Qwen2.5-Coder-7B-Instruct that turns a JavaScript function plus its AST features into an obfuscation plan (JSON). It never writes code: a deterministic Babel engine applies the plan, and a differential te…
Downloads · 30 days
47
100% of all-time downloads
All-time downloads
47
Public
Repo size
4.7 GB
Likes
1
Public
Click a slice to open those files.
.gguf4.7 GB · 100%
From the Hugging Face model README
A fine-tuned Qwen2.5-Coder-7B-Instruct that turns a JavaScript function plus its AST features into an obfuscation plan (JSON). It never writes code: a deterministic Babel engine applies the plan, and a differential test proves the obfuscated function behaves identically.
Qwen/Qwen2.5-Coder-7B-Instruct (Apache-2.0)r=32, alpha=64, dropout=0, 3 epochs, lr 2e-4, cosine, effective batch 16q8_0, q4_k_m)Neural planning of JavaScript obfuscation for standalone top-level named functions, with explicit user control over aggressiveness via target intensity:
| Target intensity | Plan shape the model must produce |
|---|---|
light | rename + dead_code only |
medium | 2–4 transforms, string_encode/operator_sub when applicable, no opaque_predicates |
heavy | all relevant transforms, always includes opaque_predicates |
Out of scope: async functions, generators, JSX/TypeScript, DOM-dependent code, and code with external dependencies — the dataset generation pipeline rejects them.
The model was trained on a raw [INST] template, not ChatML. Use exactly this layout:
[INST] <<SYS>>
{SYSTEM_PROMPT}
<</SYS>>
=== CODE ===
{javascript_source}
=== END CODE ===
=== AST FEATURES ===
{json_of_18_ast_features}
=== END AST FEATURES ===
complexity_class=medium (cyclomatic_complexity=4)
Target intensity: heavy
seed=3451783347
Generate the obfuscation plan JSON: [/INST]
SYSTEM_PROMPT used during training:
You are NeuroObfuscator. Given JavaScript code and its AST features, generate an optimal obfuscation plan as a JSON object.
Available transformations (apply in this order when enabled):
1. rename - Rename local identifiers to hex-like names. Almost always recommended.
2. string_encode - Encode string literals. Methods: charcode_array, charcode_concat, hex_escape, unicode_escape. Only enable if string_count > 0.
3. operator_sub - Substitute arithmetic/comparison operators (a+b -> a-(-b), a===b -> !(a!==b)). Use when operator_count > 2.
4. dead_code - Insert unreachable code blocks. count: 1-5. More complex code tolerates more.
5. opaque_predicates - Insert always-true/always-false conditions. count: 1-3. Primarily for medium/heavy intensity; may also be used sparingly on light functions when extra diversity is needed.
Intensity guide:
- light: cyclomatic_complexity <= 2. Prefer rename + dead_code only.
- medium: complexity 3-5. Add string_encode and operator_sub if applicable.
- heavy: complexity > 5. Use all relevant transforms aggressively.
Rules:
- You MUST honor the requested "Target intensity" when it is provided, even if
it differs from what the complexity alone would suggest. Intensity determines
the plan shape:
light -> minimal plan: rename + dead_code ONLY (no string_encode,
no operator_sub, no opaque_predicates),
medium -> moderate plan: rename + dead_code + string_encode/operator_sub
when applicable, NO opaque_predicates,
heavy -> aggressive plan: all relevant transforms INCLUDING opaque_predicates.
- Only include enabled transforms in "order" array.
- Order MUST follow: rename, string_encode, operator_sub, dead_code, opaque_predicates.
- Do NOT include a "seed" field in your JSON; the runtime injects the provided seed automatically.
- Avoid over-bloating small functions.
Output ONLY valid JSON. No explanations, no markdown.
The AST features block is produced by the project's Babel engine
(scripts/inference.py::NeuroObfuscatorInference.get_features or
node engine/index.js --json with {"operation":"extract_features",...}).
The model emits only the plan body. The seed is not predicted — the runtime injects the seed from the prompt before handing the plan to the engine.
{
"intensity": "heavy",
"transforms": {
"rename": {"enabled": true, "keep": []},
"string_encode": {"enabled": true, "method": "charcode_array", "min_length": 2},
"operator_sub": {"enabled": true, "rate": 0.95},
"dead_code": {"enabled": true, "count": 2},
"opaque_predicates": {"enabled": true, "count": 3}
},
"order": ["rename", "string_encode", "operator_sub", "dead_code", "opaque_predicates"]
}
from llama_cpp import Llama
llm = Llama(model_path="neuroobfuscator-v7.1-q8_0.gguf", n_gpu_layers=-1, n_ctx=4096)
out = llm(prompt, max_tokens=256, temperature=0.0, stop=["<|im_end|>", "</s>"], echo=False)
raw = out["choices"][0]["text"]
prompt is the [INST] block above. Stop tokens are optional — the model terminates with EOS.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct", device_map="auto")
model = PeftModel.from_pretrained(base, "neuroobfuscator-v7.1-adapter")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-Coder-7B-Instruct")
ids = tok(prompt, return_tensors="pt").to(model.device)
print(tok.decode(model.generate(**ids, max_new_tokens=256, do_sample=False)[0][ids["input_ids"].shape[1]:],
skip_special_tokens=True))
The plan is only useful with the deterministic engine + differential validation from the GitHub repository:
node engine/index.js --input input.js --plan plan.json --output obfuscated.js
Measured on 750 held-out test functions with the exported q8_0 GGUF through llama.cpp, applying every plan with the real engine and comparing original vs obfuscated behaviour on 50 argument sets:
| Metric | Result |
|---|---|
| JSON parse rate | 100.0% |
| Schema valid rate | 100.0% |
| Intensity obedience (field) | 100.0% |
| Intensity obedience (plan shape) | 100.0% |
| Light purity (light ⇒ rename + dead_code only) | 100.0% |
| Semantic pass rate | 100.0% |
| Semantic pass by intensity | light 112/112, medium 355/355, heavy 283/283 |
| Unique transform orders / top non-light share | 10 / 19.0% |
Dataset-side quality gates (scripts/09_audit_dataset.py --enforce): top non-light order share
≤15%, ≥20 unique orders, per-transform coverage floors, intensity 20/45/35 ±5 pp, zero
cross-split leakage, zero prompt/plan contradictions (rules R1–R6).
| Setting | Value |
|---|---|
| Method | QLoRA (4-bit NF4), all attention + MLP projections |
| LoRA | r=32, alpha=64, dropout=0 |
| Epochs / LR / schedule | 3 / 2e-4 / cosine, warmup 3% |
| Effective batch | 16 |
| Max seq length | 2048 |
| Loss | completion-only (prompt tokens masked) |
| Hardware | NVIDIA A100 40 GB (also runs on L4/T4 with a smaller batch) |
Data: 7,500 records (6,000 train / 750 val / 750 test) built from 103,580 differentially validated
candidate plans. Each record pairs the code + AST features + Target intensity with the
best-scoring plan for that (function, intensity) cell. 1,152 functions appear with several
intensity variants, and 100% of those variants have different transform orders — that contrast
is what teaches the model to obey the intensity control in the plan content.
vm with a timeout — not a security
boundary; run validation in an isolated container for untrusted input.Apache-2.0 (inherited from Qwen2.5-Coder-7B-Instruct). Dataset provenance records repository and license for real-code sources; no raw third-party source is redistributed.