Downloads · 30 days
16
14% of all-time downloads
crystalai35/tinyclaude-1b
tinyclaude-1b is a machine learning model from crystalai35. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A lightweight, locally-runnable language model based on TinyLlama 1.1B, enhanced with a sophisticated system prompt inspired by Claude's behavioral guidelines.
Downloads · 30 days
16
14% of all-time downloads
All-time downloads
118
Public
Repo size
638 MB
Likes
0
Public
Click a slice to open those files.
.gguf638 MB · 100%
From the Hugging Face model README
A lightweight, locally-runnable language model based on TinyLlama 1.1B, enhanced with a sophisticated system prompt inspired by Claude's behavioral guidelines.
TinyClaude-1B brings thoughtful AI assistant behavior to edge devices and resource-constrained environments. Built on the efficient TinyLlama architecture, this model incorporates carefully crafted system instructions emphasizing helpfulness, safety, and nuanced conversation.
# Pull the model
ollama pull thatdamai/tinyclaude-1b
# Run interactively
ollama run thatdamai/tinyclaude-1b
| Component | Minimum | Recommended |
|---|---|---|
| RAM | 4GB | 8GB |
| VRAM | 2GB | 4GB |
| Storage | 1GB | 2GB |
ollama run thatdamai/tinyclaude-1b
curl http://localhost:11434/api/generate -d '{
"model": "thatdamai/tinyclaude-1b",
"prompt": "Explain quantum computing simply.",
"stream": false
}'
import requests
response = requests.post('http://localhost:11434/api/generate', json={
'model': 'thatdamai/tinyclaude-1b',
'prompt': 'What is machine learning?',
'stream': False
})
print(response.json()['response'])
Simply select thatdamai/tinyclaude-1b from the model dropdown after pulling.
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load model and tokenizer
model_name = "TinyLlama/TinyLlama-1.1B-Chat-v1.0"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")
# Define the TinyClaude system prompt
system_prompt = """You are a helpful, harmless, and honest AI assistant..."""
# Format with chat template
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": "Explain quantum computing simply."}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
# Generate response
outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
from llama_cpp import Llama
# Download GGUF from Hugging Face Hub
llm = Llama.from_pretrained(
repo_id="TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF",
filename="tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf",
n_ctx=2048,
n_gpu_layers=-1 # Use all GPU layers
)
system_prompt = """You are a helpful, harmless, and honest AI assistant..."""
output = llm.create_chat_completion(
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": "What is machine learning?"}
],
temperature=0.7,
max_tokens=512
)
print(output['choices'][0]['message']['content'])
# Install huggingface_hub
pip install huggingface_hub
# Download model files
huggingface-cli download TinyLlama/TinyLlama-1.1B-Chat-v1.0 --local-dir ./tinyllama
# Download GGUF quantized version
huggingface-cli download TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf --local-dir ./tinyllama-gguf
# Run with Docker
docker run --gpus all --shm-size 1g -p 8080:80 \
ghcr.io/huggingface/text-generation-inference:latest \
--model-id TinyLlama/TinyLlama-1.1B-Chat-v1.0 \
--max-input-length 1024 \
--max-total-tokens 2048
# Query the endpoint
curl http://localhost:8080/generate \
-X POST \
-H 'Content-Type: application/json' \
-d '{"inputs": "<|system|>\nYou are a helpful assistant.</s>\n<|user|>\nHello!</s>\n<|assistant|>\n", "parameters": {"max_new_tokens": 256}}'
| Property | Value |
|---|---|
| Base Model | TinyLlama 1.1B |
| Parameters | 1.1 Billion |
| Context Window | 2048 tokens |
| License | Apache 2.0 |
| Quantization | Q4_0 (default) |
TinyClaude-1B is well-suited for:
As a 1.1B parameter model, TinyClaude-1B has inherent limitations:
For demanding tasks, consider larger models like Llama 3.1 8B, Mistral 7B, or Qwen 14B.
Create your own variant:
# Create a Modelfile
cat << 'EOF' > Modelfile
FROM tinyllama
SYSTEM """
Your custom system prompt here.
"""
PARAMETER temperature 0.7
PARAMETER num_ctx 2048
EOF
# Build the model
ollama create my-tinyclaude -f Modelfile
# Test it
ollama run my-tinyclaude
# Install required tools
pip install huggingface_hub
# Login to Hugging Face
huggingface-cli login
# Create a new model repository
huggingface-cli repo create tinyclaude-1b --type model
# Upload model files
huggingface-cli upload thatdamai/tinyclaude-1b ./model-files --repo-type model
# Find your Ollama model location
ollama show thatdamai/tinyclaude-1b --modelfile
# Models are stored in ~/.ollama/models or /usr/share/ollama/.ollama/models
# Copy the blob files and upload to HF
# Alternative: Use ollama's model export (if available)
cp /usr/share/ollama/.ollama/models/blobs/<sha256-hash> ./tinyclaude.gguf
Create a README.md in your HF repo with YAML frontmatter:
---
license: apache-2.0
base_model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
tags:
- tinyllama
- gguf
- ollama
- assistant
- conversational
model_type: llama
pipeline_tag: text-generation
inference: false
---
# Method 1: Create Modelfile pointing to HF GGUF
cat << 'EOF' > Modelfile
FROM hf.co/thatdamai/tinyclaude-1b-gguf
EOF
ollama create tinyclaude-local -f Modelfile
# Method 2: Download GGUF first, then import
huggingface-cli download thatdamai/tinyclaude-1b-gguf tinyclaude-1b.Q4_K_M.gguf --local-dir ./
cat << EOF > Modelfile
FROM ./tinyclaude-1b.Q4_K_M.gguf
EOF
ollama create tinyclaude-local -f Modelfile
Suggestions and improvements are welcome. Feel free to:
This model inherits the Apache 2.0 license from TinyLlama. The system prompt and configuration are provided as-is for educational and personal use.
Author: thatdamai
Model: thatdamai/tinyclaude-1b
Platform: Ollama