Downloads · 30 days
18
4% of all-time downloads
sfsx/Z-Image-Engineer-V6
Z-Image-Engineer-V6 is a text generation model from sfsx. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
18
4% of all-time downloads
All-time downloads
418
Public
Parameters
4B
33.4 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
Follow me on X @BennyDaBall_OG !
| Key | Value |
|---|---|
| License | Apache-2.0 |
| Language | English (en) |
| Base Model | Tongyi-MAI/Z-Image-Turbo |
| Library | transformers |
| Pipeline Tag | text-generation |
| Format | HF Safetensors |
The Z-Engineer returns, fully rebuilt around the SMART DoRA training system for Z-Image Turbo.
Yes, we jump from V4 to V6. Unlike the usual guy math, this one actually brought the extra two inches.
Z-Image-Engineer V6 is a fine-tuned 4B Qwen text encoder (Tongyi-MAI/Z-Image-Turbo) optimized for dual-role performance: a local prompt-enhancement model and a merged HF text encoder for Z-Image workflows. The ComfyUI-Z-Engineer node runs both roles fully inside ComfyUI from this release.

V6 transforms minimal seed prompts into rich, highly structured visual narratives. It adds explicit scene composition, lighting direction, material texture, and depth separation while stripping out empty prompt sludge like "8k, masterpiece, trending on ArtStation."
It can also be used directly as a Z-Image text encoder. This repo contains the merged HF safetensors. The GGUF quantized release lives in the companion repo: Z-Image-Engineer-V6-GGUF.
llama.cpp. No API logs, no external telemetry.V4 pioneered SMART training. V6 adapts that system into a Weight-Decomposed Low-Rank Adaptation (DoRA) framework.
DoRA provides surgical adapter updates by decoupling directional and magnitude adjustments. SMART adds auxiliary pressure so the model does not collapse into repetitive prompt loops or superficial sentence patterns.
| Regularizer | What it Does | Why it Matters |
|---|---|---|
| Entropic | Broadens output probability diversity. | Reduces repetitive loops and generic vocabulary. |
| Holographic | Enforces structured, depth-wise feature logic. | Improves foreground/background hierarchy. |
| Topological | Stabilizes coherent latent trajectories. | Keeps prompts flowing naturally instead of stalling out. |
| Manifold | Regulates overall weight distributions. | Keeps model behavior stable under high-pressure refinement. |
V6 was not a simple one-and-done training run. The final architecture is a blended composite:
Use this merged HF release directly where supported, or download a GGUF quant from Z-Image-Engineer-V6-GGUF for LM Studio. No complex system prompt is required.
Enhance this image prompt for Z-Image Turbo: a unicorn
The comparison examples were generated from direct LM Studio user requests like this, with no separate system prompt. V6_SYSTEM_PROMPT.md is included only as an optional preset for people who want a stricter prompt-only chat setup.
Use the ComfyUI-Z-Engineer custom node (v2.0+). It loads this repo's sharded safetensors release directly and runs V6 as both the Z-Image text encoder and an in-ComfyUI prompt enhancer - no LM Studio or external server required.
ComfyUI/models/text_encoders/Z-Image-Engineer-V6/ (the three model-0000X-of-00003.safetensors shards plus model.safetensors.index.json).Z-Image-Engineer-V6/ from the dropdown.clip into your Z-Image CLIP Text Encode - V6 replaces the stock Qwen text encoder.clip to rewrite seed prompts in-process; the enhanced prompt is previewed right on the node.A ready-made workflow ships with the node repo: example_workflows/z_image_turbo_z_engineer.json.
Prefer smaller files? Use a quant from Z-Image-Engineer-V6-GGUF with the node's Z-Engineer CLIP Loader (GGUF) instead.
UNET: z_image_turbo_bf16.safetensors
VAE: ae.safetensors
Text Encoder: Z-Image-Engineer-V6 (this repo's sharded safetensors, or a GGUF quant)
Resolution: 1024x1024
Steps: 8
CFG: 1.0
Sampler: res_multistep
Scheduler: simple
Shift: 3.0
| Parameter | Specification |
|---|---|
| Base Text Encoder | Tongyi-MAI/Z-Image-Turbo/text_encoder |
| Tokenizer | Tongyi-MAI/Z-Image-Turbo/tokenizer |
| Method | SMART DoRA / PEFT Adapter Training |
| Rank / Alpha / Dropout | 64 / 64 / 0.03 |
| Target Modules | q_proj, k_proj, v_proj, o_proj, gate_proj, down_proj, up_proj |
| Refinement Stack | Supervised Style SFT + Binary Anti-Repeat |
| Final Packaging | Merged HF safetensors |
The quantized release is separate on purpose:
BennyDaBall/Z-Image-Engineer-V6-GGUF
That repo contains the full GGUF ladder: F16, Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q3_K_M, and MXFP4.
The bundled comparison image is:
evidence/gallery_z_image_engineer_v6_simple_ab_with_rewrites_CONTACT.png
It compares foundational prompts across four isolated control paths:
This model is a prompt engineer and text encoder. Diffusion is still diffusion; structural expansion improves compositional adherence, but it does not mathematically guarantee a perfect seed every single time. Use creative judgment locally.
Built & trained locally with care by BennyDaBall.
Follow me on X @BennyDaBall_OG !