Downloads · 30 days
109
56% of all-time downloads
kiel2/Kiel-2-Sol
Kiel-2-Sol is a text generation model from kiel2. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as apache-2.0.
Kiel-2-Sol is an enterprise-grade, high-performance conversational language model fine-tuned specifically to anchor the premium tier of the KielTech AI production ecosystem. Built upon Meta's meta-llama/Llama-3.2-3B-I…
Downloads · 30 days
109
56% of all-time downloads
All-time downloads
194
Public
Parameters
3.2B
11.9 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors6.4 GB · 54%
From the Hugging Face model README
Kiel-2-Sol is an enterprise-grade, high-performance conversational language model fine-tuned specifically to anchor the premium tier of the KielTech AI production ecosystem. Built upon Meta's meta-llama/Llama-3.2-3B-Instruct architecture, the model undergoes an intensive parameter-efficient adaptation process targeting both attention mechanism layers and core multi-layer perceptron (MLP) blocks.
This model balances dense reasoning capacities and highly structured instruction-following capabilities, making it ideal for robust enterprise automation, complex multi-turn API workflows, and rapid serverless deployment via engines like vLLM.
Kiel-2-Sol is engineered for deployment within high-volume production setups. It directly services complex systemic tasks including:
This model is not intended for unmonitored critical safety systems, malicious text generation, or downstream applications that lack safety guardrails or validation layers.
The intelligence profile of Kiel-2-Sol is derived from a highly curated 10,000-sample strategic mixture ingested via real-time cloud streaming (streaming=True):
mlabonne/FineTome-100k: Optimized to maximize natural conversational pacing, verbal crispness, and conversational alignment.Arcee-AI/Llama-3.1-SuperNova-Lite: A heavily distilled dataset used to inject advanced multi-step reasoning patterns and complex instruction-following capabilities.Training was completed within a highly optimized 4-bit NormalFloat (nf4) workspace, applying the official Llama 3 structural chat template tokens during streaming ingestion to guarantee exact template cohesion.
q_proj, v_proj, k_proj, o_proj, gate_proj, up_proj, down_projpaged_adamw_8bitbfloat16 compute precision to maximize gradient calculation throughput within a standard 16GB VRAM constraint.transformers, peft, trl (Supervised Fine-Tuning Trainer), and bitsandbytes.For business backend pipelines, loading Kiel-2-Sol into a vLLM offline engine or server instance provides optimal throughput:
from vllm import LLM, SamplingParams
# Load the premium enterprise model directly from the Hub
llm = LLM(
model="kiel2/Kiel-2-Sol",
quantization="bitsandbytes",
load_format="bitsandbytes",
max_model_len=2048
)
sampling_params = SamplingParams(temperature=0.7, top_p=0.9, max_tokens=256)
prompts = ["<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\nGenerate an enterprise system-architecture report summary for KielTech AI.<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n"]
outputs = llm.generate(prompts, sampling_params)
for output in outputs:
print(output.outputs[0].text)