Downloads · 30 days
120
18% of all-time downloads
dreeseaw/cleo
cleo is a text generation model from dreeseaw. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Cleo is a small SQL analyst hardel: a Qwen3.5-2B fine-tune paired with a read-only SQL harness/runtime for live database connections. The model is trained to inspect schemas, gather real values from the database, repa…
Downloads · 30 days
120
18% of all-time downloads
All-time downloads
662
Public
Parameters
1.9B
9.6 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors7.5 GB · 65%
From the Hugging Face model README
Cleo is a small SQL analyst hardel: a Qwen3.5-2B fine-tune paired with a read-only SQL harness/runtime for live database connections. The model is trained to inspect schemas, gather real values from the database, repair SQL from execution feedback, and return analyst-ready read-only queries.
The recommended entry point is the Python package and MCP server:
github.com/Dreeseaw/cleo.
pip install "cleo-sql[hf] @ git+https://github.com/Dreeseaw/cleo.git@master"
from cleo import Cleo
cleo = Cleo.from_hf("dreeseaw/cleo")
ans = cleo("How many employees are currently in each department?", conn)
print(ans.sql)
print(ans.rows)
print(ans.clarification)
cleo(...) and cleo.ask(...) use the hardel runtime by default: a greedy candidate plus sampled candidates are executed through the same read-only harness, then selected with product-visible execution evidence. Use cleo.ask_once(...) when you explicitly want a single-candidate, lower-latency path.
| file | purpose |
|---|---|
| root model files | Current hardel Transformers weights in bf16-compatible safetensors format. |
v1.4-hardel-v3/ | Archived copy of the current hardel checkpoint. |
cleo-Q8_0.gguf | Legacy llama-cpp-python GGUF alias from a prior release. |
cleo_v1_2_bird-no_mtp-Q8_0.gguf | Prior versioned Q8_0 GGUF artifact. |
v1.0/ | Earlier archived tool-use checkpoint. |
For the current release, use the Hugging Face Transformers backend through Cleo.from_hf("dreeseaw/cleo"). The GGUF files are retained for compatibility with earlier runtime paths.
All rows below use denotation scoring on BIRD minidev SQLite: predicted and gold SQL are executed, normalized row sets are compared, and the formula_1 database is excluded because of training overlap. BIRD-434 is reported as a broad analytical SQL benchmark, not as the only measure of the product runtime.
| model / runtime | BIRD-434 execution accuracy |
|---|---|
| Cleo v1.4 hardel K=4 | 143/434 = 33.0% |
| Gemini 2.5 Flash | 55.5% |
| DeepSeek Chat | 50.5% |
Cleo was evaluated through the public package runtime with k=4, temperature=0.7, max_gather=3, and max_repair=2.
Cleo starts from Qwen/Qwen3.5-2B-Base, then adds SQL analyst behavior in stages:
The root model files are current bf16 Transformers weights. They are not stored as int8 weights. For CUDA machines that need lower VRAM, install the optional extra and load with runtime quantization:
pip install "cleo-sql[hf,int8] @ git+https://github.com/Dreeseaw/cleo.git@master"
cleo = Cleo.from_hf("dreeseaw/cleo", quantization="int8")
By default, Cleo.from_hf() asks PyTorch what is available and chooses CUDA, then XPU, then MPS, then CPU. CUDA is the tested fast path for this release; CPU is supported by compatible Transformers installs but is slower.
tables= or a provided schema= string so the runtime sees the relevant part of the database.Dreeseaw/cleocleo-sqlQwen/Qwen3.5-2B-Basedreeseaw/cleo-value-discoverydreeseaw/cleo-process-analytics-v1