Downloads · 30 days
65
4% of all-time downloads
applexml/kimi-k2-poc2
kimi-k2-poc2 is a text generation model from applexml. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
NanoAgent is a compact 135M parameter, 8k context-length language model trained to perform tool calls and generate responses based on tool outputs. Despite its small size (~135 MB in 8-bit precision), it’s optimized f…
Downloads · 30 days
65
4% of all-time downloads
All-time downloads
1.6K
Public
Parameters
135M
269 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors269 MB · 96%
From the Hugging Face model README
NanoAgent is a compact 135M parameter, 8k context-length language model trained to perform tool calls and generate responses based on tool outputs.
Despite its small size (~135 MB in 8-bit precision), it’s optimized for agentic use cases and runs easily on personal devices.
Github: NanoAgent
Inference resource: link
Base model: SmolLM2-135M-Instruct
Fine-tuning method: Dynamic Fine-Tuning (DFT)
Hardware: Apple Mac M1 (16 GB Unified Memory) using MLX.
microsoft/orca-agentinstruct-1M-v1 — agentic tasks, RAG answers, classificationmicrosoft/orca-math-word-problems-200k — lightweight reasoningallenai/tulu-3-sft-personas-instruction-following — instruction followingxingyaoww/code-act — ReAct style reasoning and actionm-a-p/Code-Feedback — alignment via feedbackHuggingFaceTB/smoltalk + /apigen — tool calling stabilizationweijie210/gsm8k_decomposed — question decompositionLocutusque/function-calling-chatml — tool call response structureThis is a beta model.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "quwsarohi/NanoAgent-135M"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name)
def inference(messages, max_new_tokens=256, temperature=0.3, min_p=0.15, **kwargs):
input_text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer.encode(input_text, return_tensors="pt")
outputs = model.generate(
inputs,
max_new_tokens=max_new_tokens,
do_sample=True,
min_p=0.15,
temperature=temperature,
**kwargs
)
return tokenizer.decode(outputs[0][inputs.shape[1] :], skip_special_tokens=True)
messages = [{"role": "user", "content": "Hi! Do you have a name?"}]
print(inference(messages))
Use the following template for tool calling:
TOOL_TEMPLATE = """You are a helpful AI assistant. You have a set of possible functions/tools inside <tools></tools> tags.
Based on question, you may need to make one or more function/tool calls to answer user.
You have access to the following tools/functions:
<tools>{tools}</tools>
For each function call, return a JSON list object with function name and arguments within <tool_call></tool_call> tags."""
Sample tool call definition:
{
"name": "web_search",
"description": "Performs a web search for a query and returns a string of the top search results formatted as markdown with titles, links, and descriptions.",
"parameters": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "The search query to perform.",
}
},
"required": ["query"],
},
}