Downloads · 30 days
624
50% of all-time downloads
KordAI/KeawGPT
KeawGPT is a text generation model from KordAI. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
KeawGPT is a 4B-parameter language model by KordAI, fine-tuned from KordAI/KeawGPT-Base (Qwen3 architecture) to specialize in Science, Technology, Engineering, Mathematics (STEM), and Philosophy. It is designed to giv…
Downloads · 30 days
624
50% of all-time downloads
All-time downloads
1.2K
Public
Parameters
4B
8.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
KeawGPT is a 4B-parameter language model by KordAI, fine-tuned from KordAI/KeawGPT-Base (Qwen3 architecture) to specialize in Science, Technology, Engineering, Mathematics (STEM), and Philosophy. It is designed to give clear, reasoned answers across technical and conceptual domains, and supports structured tool calling for tasks like web search and code execution.
KeawGPT is intended for:
It is not intended as a substitute for professional advice (medical, legal, financial) or as an authoritative source without verification — like any language model, it can produce confident-sounding but incorrect output, especially on niche technical claims or citations.
KeawGPT uses a custom plain-text turn format rather than a standard chat markup:
# SYSTEM:
{system message}
# USER:
{user message}
# ASSISTANT:
{assistant message}
# TOOL:
{tool result}
Each assistant turn ends with the model's EOS token. Use tokenizer.apply_chat_template() to render this — a custom chat_template.jinja ships with the tokenizer config and handles turn formatting, the default system message, and tool-schema injection automatically. Building the prompt by hand is unnecessary and error-prone (whitespace/newline placement, tool-schema formatting, etc. are already handled by the template).
KeawGPT can emit structured tool calls instead of a direct answer. When it decides to call a tool, it replies only in this format, with no other content before or after:
<tools_call>
<function=tool_name>
<parameter=param_name>
value
</parameter>
</function>
</tools_call>
The calling application is expected to parse this block, execute the corresponding tool, and feed the result back as a # TOOL: turn before continuing generation. The model may optionally reason in natural language before a tool call, but should not add anything after the closing </tools_call> tag.
Inference with 🤗 Transformers, loaded in 4-bit via bitsandbytes:
import re
import json
import torch
from transformers import (
AutoModelForCausalLM,
AutoTokenizer,
BitsAndBytesConfig,
TextStreamer,
StoppingCriteria,
StoppingCriteriaList,
)
MODEL_NAME = "KordAI/KeawGPT"
# --------------------------------------------------------------------------
# 1. Load model in 4-bit
# --------------------------------------------------------------------------
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
)
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModelForCausalLM.from_pretrained(
MODEL_NAME,
quantization_config=bnb_config,
device_map="auto",
)
model.eval()
SYSTEM_PROMPT = None # unused — chat_template.jinja builds the default system
# message + tools instruction automatically from `tools`
# --------------------------------------------------------------------------
# 2. Prompting via tokenizer.apply_chat_template
# --------------------------------------------------------------------------
# KeawGPT ships with a custom chat_template.jinja (bundled in this repo's
# tokenizer_config.json) that renders conversations in the model's native
# # SYSTEM: / # USER: / # ASSISTANT: / # TOOL: turn format and injects the
# tool schema into the default system message. Just pass `tools=tools` and
# let apply_chat_template build the prompt — no manual string building needed.
def build_prompt(history):
return tokenizer.apply_chat_template(
history,
tools=tools,
tokenize=False,
add_generation_prompt=True,
)
STOP_STRINGS = ["</tool_call>", "</tools_call>", "\n# USER:", "\n# TOOL:"]
class StopOnSubstrings(StoppingCriteria):
def __init__(self, tokenizer, stop_strings, prompt_len, check_every=4):
self.tokenizer = tokenizer
self.stop_strings = stop_strings
self.prompt_len = prompt_len
self.check_every = check_every
self._count = 0
def __call__(self, input_ids, scores, **kwargs):
self._count += 1
if self._count % self.check_every != 0:
return False
generated = self.tokenizer.decode(input_ids[0][self.prompt_len:], skip_special_tokens=True)
return any(s in generated for s in self.stop_strings)
def generate(prompt, max_new_tokens=1024):
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
prompt_len = inputs["input_ids"].shape[1]
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
stopping_criteria = StoppingCriteriaList([StopOnSubstrings(tokenizer, STOP_STRINGS, prompt_len)])
with torch.no_grad():
output_ids = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
do_sample=True,
temperature=0.7,
top_p=0.9,
pad_token_id=tokenizer.eos_token_id,
streamer=streamer,
stopping_criteria=stopping_criteria,
)
return tokenizer.decode(output_ids[0][prompt_len:], skip_special_tokens=True)
# --------------------------------------------------------------------------
# 3. Example single-turn call
# --------------------------------------------------------------------------
history = [{"role": "user", "content": "What's the difference between a P and an NP problem?"}]
prompt = build_prompt(history)
output = generate(prompt)
print(output)
Wrap this in a loop that appends {"role": "assistant", ...} and, when parse_tool_call() returns non-None, dispatches the call, appends the result as a {"role": "tool", ...} turn, and re-generates — this is how tool-calling conversations are extended turn by turn.
<tools_call> blocks and returning results as # TOOL: turns; without that harness, the model may describe a tool call in text without it actually being executed.