Downloads · 30 days
156
16% of all-time downloads
AuricErgeson/shellwhiz-7b
shellwhiz-7b is a text generation model from AuricErgeson. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A fine-tune of Qwen2.5-7B-Instruct that turns a plain-English request into a shell command. Type what you want to do, get back find, grep, docker, git, or whatever fits.
Downloads · 30 days
156
16% of all-time downloads
All-time downloads
961
Public
Parameters
7.6B
19.9 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors15.2 GB · 76%
From the Hugging Face model README
A fine-tune of Qwen2.5-7B-Instruct that turns a plain-English request into a shell command. Type what you want to do, get back find, grep, docker, git, or whatever fits.
I wanted something I could actually type "show me the 5 biggest files in this folder" into and get a working du/sort/head pipeline back, instead of half-remembering the flags myself. There are commercial tools that do this (Warp, some IDE plugins), but I couldn't find a small open model that just did the one thing well, so I built one.
Trained on 697 natural-language-to-shell-command pairs covering:
find, cp, mv, rm, chmod, du)grep, sed, awk, sort, cut)ps, kill, systemctl)curl, ssh, scp, ping)tar, zip, apt, pip, npm)These are from the actual post-training sanity check, not cherry-picked from the training set:
> find all python files modified in the last 24 hours
find <directory> -name '*.py' -mtime -1
> show me the 5 largest files in this directory
find . -type f -exec du -h {} + | sort -rh | head -5
> list all running docker containers
docker ps
The <directory> placeholder is intentional. The model was trained to use placeholders where a real path would depend on context, rather than guessing one.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_name = "AuricErgeson/shellwhiz-7b"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16, device_map="auto")
system_msg = "You are a helpful assistant that converts natural language requests into precise shell commands. Respond with ONLY the shell command, no explanation."
messages = [
{"role": "system", "content": system_msg},
{"role": "user", "content": "find all files larger than 100MB"},
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=100, temperature=0.1)
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
A GGUF (q4_k_m) build is also available in this repo if you want to run it locally through Ollama or llama.cpp.
Loss dropped from 2.96 at step 10 to 0.18 by the end of training and flattened out around step 190, with no spikes or divergence:
| Step | Loss |
|---|---|
| 10 | 2.960 |
| 50 | 0.391 |
| 100 | 0.268 |
| 150 | 0.258 |
| 200 | 0.188 |
| 260 | 0.188 |
I ran the model against 105 held-out prompts it never saw during training, phrased differently from the training set on purpose to test generalization rather than recall. Each output was judged by Claude against a known-correct reference command, allowing for different-but-equivalent approaches (there's rarely only one right way to write a shell command).
| Verdict | Count | Percent |
|---|---|---|
| Correct | 58 | 55.2% |
| Partial (right idea, has a bug) | 26 | 24.8% |
| Wrong | 21 | 20.0% |
Syntax validity (does bash -n parse it without error) came out at 100/105, or 95.2 percent.
The five syntax failures were almost all cases where the model left a bracketed placeholder like
<filename> or <output_file> in a spot where bash needs an actual token, which reads as a
formatting habit rather than the model not understanding the command it's building.
The wrong and partial cases cluster into a few recognizable patterns, worth knowing before you rely on this for anything important:
docker images --sort, pkill --exclude). These would fail immediately with an error, so
at least they're not silently wrong.<image_name>:<tag>
style placeholders even when the prompt spelled out a concrete value like nginx:latest.None of this is surprising for 700 training examples on a 7B model, but it's worth knowing which categories to double check rather than trusting blindly.
The training data is published separately at AuricErgeson/text-to-shell-dataset, generated synthetically and deduplicated on instruction text.