Downloads ยท 30 days
43
22% of all-time downloads
Lamapi/next-32b-GGUF
next-32b-GGUF is a text generation model from Lamapi. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Downloads ยท 30 days
43
22% of all-time downloads
All-time downloads
198
Public
Repo size
34.8 GB
Likes
1
Public
Click a slice to open those files.
.gguf34.8 GB ยท 100%
From the Hugging Face model README

Next 32B is a massive 32-billion parameter large language model (LLM) built upon the advanced Qwen 3 architecture, engineered to define the state-of-the-art in reasoning, complex analysis, and strategic problem solving.
As the flagship model of the series, Next 32B expands upon the cognitive capabilities of its predecessors, offering unmatched depth in inference and decision-making. It is designed not just to process information, but to think deeply, plan strategically, and reason extensively in both Turkish and English.
Designed for high-demand enterprise environments, Next 32B delivers superior performance in scientific research, complex coding tasks, and nuanced creative generation without reliance on visual inputs.
Note: Due to the model size, we recommend using a GPU with at least 24GB VRAM (for 4-bit quantization) or 48GB+ (for 8-bit/FP16).
!pip install unsloth
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained("Lamapi/next-32b")
messages = [
{"role": "system", "content": "You are Next-X1, an AI assistant created by Lamapi. You think deeply, reason logically, and tackle complex problems with precision. You are an helpful, smart, kind, concise AI assistant."},
{"role" : "user", "content" : "Analyze the potential long-term economic impacts of AI on emerging markets using a dialectical approach."}
]
text = tokenizer.apply_chat_template(
messages,
tokenize = False,
add_generation_prompt = True,
enable_thinking = True, # Enable thinking
)
from transformers import TextStreamer
_ = model.generate(
**tokenizer(text, return_tensors = "pt").to("cuda"),
max_new_tokens = 1024, # Increase for longer outputs!
temperature = 0.7, top_p = 0.95, top_k = 400,
streamer = TextStreamer(tokenizer, skip_prompt = True),
)
| Feature | Description |
|---|---|
| ๐ง Deep Cognitive Architecture | Capable of handling massive context windows and multi-step logical chains. |
| ๐น๐ท Cultural Mastery | Native-level nuance in Turkish idioms, history, and law, alongside global fluency. |
| โ๏ธ High-Performance Scaling | Optimized for multi-GPU inference and heavy workload batching. |
| ๐งฎ Scientific & Coding Excellence | Solves graduate-level physics, math, and complex software architecture problems. |
| ๐งฉ Pure Reasoning Focus | Specialized textual intelligence without the overhead of vision encoders. |
| ๐ข Enterprise Reliability | Deterministic outputs suitable for legal, medical, and financial analysis. |
| Specification | Details |
|---|---|
| Base Model | Qwen 3 |
| Parameters | 32 Billion |
| Architecture | Transformer (Causal LLM) |
| Modalities | Text-only |
| Fine-Tuning | Advanced SFT & RLHF on Cognitive Kernel & KAG-Thinker datasets |
| Optimizations | GQA, Flash Attention 3, Quantization-ready |
| Primary Focus | Deep Reasoning, Complex System Analysis, Strategic Planning |
Licensed under the MIT License โ free for commercial and non-commercial use. Attribution is appreciated.
Next 32B โ Tรผrkiyeโs flagship reasoning model. Built for those who demand depth, precision, and massive intelligence.