Downloads · 30 days
372
65% of all-time downloads
iBotIA/Qwen3.8-4B-Empero-AI-FullStack-GGUF
Qwen3.8-4B-Empero-AI-FullStack-GGUF is a text generation model from iBotIA. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This repository contains the optimized GGUF quantization of Qwen3.8-4B-Empero-AI-Distill-FullStack, fine-tuned using the Unsloth framework for advanced full-stack web and mobile software development pipelines.
Downloads · 30 days
372
65% of all-time downloads
All-time downloads
574
Public
Repo size
9.9 GB
Likes
0
Public
Click a slice to open those files.
.gguf6.3 GB · 100%
From the Hugging Face model README
This repository contains the optimized GGUF quantization of Qwen3.8-4B-Empero-AI-Distill-FullStack, fine-tuned using the Unsloth framework for advanced full-stack web and mobile software development pipelines.
The base architecture features a full-parameter distillation of reasoning traces (Chain-of-Thought via <think>...</think> tags) from the frontier-scale Qwen3.8 2.4T A95B teacher model developed by Empero-AI. This configuration offers advanced local planning, logic, and code compilation compliance within a highly efficient 4-billion parameter footprint.
The model underwent continuous pre-training on 279,049 curated data segments across 8 strictly isolated developer knowledge directories:
The fine-tuning process completed 250 hardware-optimized steps on a T4 GPU. The learning curve showed a definitive late convergence ("Eureka" moment) near step 140, where the weights successfully aligned cross-stack framework logic.
3.6434802.8297892.7884932.508277Qwen3.8-4B-Empero-AI-Distill-FullStack-Q6_K.ggufWhen deploying this GGUF file inside Unsloth Desktop, LM Studio, Jan, or Ollama, configure these 3 runtime settings to prevent system stuttering:
temperature=0.6, top_p=0.95, and top_k=20. Avoid a raw greedy search (temperature=0) to prevent the reasoning tokens from falling into endless structural loops.16384 or 32768. This expanded context window allows autonomous agents to evaluate several source files at the same time.To launch this model as an active developer backend connected to your terminal agent, run the OpenAI-compatible local engine server:
unsloth start opencode --context-length 32000
If this model helped you, consider supporting the project:
18cBC5sFjtctw121ULTkxTbTZPurginJBsltc1q3jrcwrx66xpz4k92p08u8c5v8zwywk3dqpzdkvTGKVpbbznmvEusKbuZZj4WSK6XxtHcG6FE (TRX chain)0x1059cb5a1F8467e5b56a9bdf082cE86FFB002D15 (POL chain)0x18b2AA731daeFD47DFFa278f3F856eAF80376fd6 (ETH chain)0x3bEcddC7c49bDba5503eB1677628b4519439884c (BNB chain)Quantizations are built upon empero-ai/Qwen3.8-4B-Distill. Weights inherit the permissive Apache-2.0 license from the base Qwen repository and are shared as-is.