Downloads · 30 days
720
51% of all-time downloads
FredQuant/corriente-qwen38-flash-next
corriente-qwen38-flash-next is a machine learning model from FredQuant. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
The full 82GB model, distilled into a single 16GB file that runs anywhere. We took Qwen3.8-Flash-Next — Alibaba's 125B-parameter Mixture-of-Experts model — and extracted its core intelligence into a single, deployment…
Downloads · 30 days
720
51% of all-time downloads
All-time downloads
1.4K
Public
Repo size
17.1 GB
Likes
8
Public
Click a slice to open those files.
.gguf17.1 GB · 100%
From the Hugging Face model README
The full 82GB model, distilled into a single 16GB file that runs anywhere. We took Qwen3.8-Flash-Next — Alibaba's 125B-parameter Mixture-of-Experts model — and extracted its core intelligence into a single, deployment-ready GGUF. Same brain. 80% smaller. Runs on a laptop.
The original model ships as 82GB across multiple shard files. That's fine for data centers, but most teams can't deploy that. We extracted the backbone — the 6B active parameters that actually think — and requantized it to Q4_K_M. The result: a single 16GB file that fits in 20GB of RAM. What we kept: Full reasoning capability, 262K context, all 512 experts. What we removed: The 24GB N-gram embedding lookup table (training-only, not needed for inference).
| Property | Value |
|---|---|
| Total parameters | 125B |
| Active parameters | 6B per token |
| Experts | 512 (10 routed + 1 shared) |
| Quantization | Q4_K_M |
| File size | 16GB |
| Context length | 262,144 tokens |
| Architecture | Gated DeltaNet + Qwen Sparse Attention |
| Original size | 82GB |
| Size reduction | 80% |
| License | Apache 2.0 |
# Create the model
cat > Modelfile << 'EOF'
FROM ./qwen3.8-flash-next-Q4.gguf
PARAMETER num_ctx 16384
EOF
ollama create qwen38-flash -f Modelfile
ollama run qwen38-flash
./llama-server \
-m qwen3.8-flash-next-Q4.gguf \
-c 16384 \
--host 0.0.0.0 \
--port 8080
from llama_cpp import Llama
llm = Llama(model_path="./qwen3.8-flash-next-Q4.gguf", n_ctx=16384)
output = llm("Explain quantum computing in simple terms", max_tokens=200)
print(output["choices"][0]["text"])
This model is a reasoning model: before answering, it runs an internal thinking pass (same style as DeepSeek-style chain-of-thought). How you see it depends on the runtime:
content contains only the final
answer, and the thinking pass comes back in its own thinking field. Ollama cannot bake the
toggle into the model (PARAMETER think is unsupported), so pass it per request:
curl http://localhost:11434/api/chat -d '{
"model": "qwen38-flash",
"think": false,
"messages": [{"role": "user", "content": "Hello"}]
}'
Set "think": false for tool-calling / agentic loops and chat UIs where you want the
final answer only. Models that don't reason simply ignore the flag.think on, the answer stream stays clean and the
reasoning is reported separately; with think off, replies are fully direct. Native tool
calls work correctly either way.We don't just share models. We make them work for you. This GGUF is one example of what we do at Corriente. We take powerful open-source models and make them deployable, customizable, and aligned to your specific work.
Model Distillation & Quantization We extract the intelligence from large models and compress them for your hardware. From 82GB to 16GB. From cloud-only to run-anywhere. We've done this with Qwen, Llama, DeepSeek, and more. Custom Model Training We fine-tune models on your data, your domain, your terminology. Legal firms, hospitals, engineering teams — we build models that speak your language. Not generic AI. Your AI. Fleet Deployment We deploy models across multi-node clusters for parallel inference. Our infrastructure runs on NVIDIA DGX Sparks — enterprise-grade hardware at a fraction of the cost. We handle the orchestration so you don't have to. Data Cleaning & Preparation Garbage in, garbage out. We clean, structure, and prepare your data for training. Our pipeline handles deduplication, quality filtering, domain tagging, and format normalization. We've processed millions of records across dozens of domains. Integration & Support We don't drop a model on your doorstep and walk away. We integrate it into your workflow, train your team, and provide ongoing support. We're builders, not vendors.
If you need a model that thinks about your work, not just general knowledge — we can build it. Email: [email protected] Website: corriente.ai We leave no one behind.
This model is free because we believe powerful AI should be in everyone's hands — but a small star goes a long way.
If this model helped you, we'd truly appreciate a ⭐ on this repo. It takes one click, it's free, and it tells us (and the world) this work matters. It also helps more people find a model they can actually run.
And if you want to go further — tell us what you built with it, or share it with a team that needs local AI. Word of mouth from people who actually use the work is the best fuel we know.
Built by Corriente — the future is frequency. ⭐ if you agree.