Downloads · 30 days
18
25% of all-time downloads
theaicompany02/Shivik-1B
Shivik-1B is a text generation model from theaicompany02. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
A 1B parameter model optimized for mathematical reasoning and chain-of-thought problem solving.
Downloads · 30 days
18
25% of all-time downloads
All-time downloads
71
Public
Parameters
1.2B
5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.9 GB · 100%
From the Hugging Face model README
A 1B parameter model optimized for mathematical reasoning and chain-of-thought problem solving.
| Benchmark | Score |
|---|---|
| GSM8K (5-shot) | 29.72% |
| Component | Value |
|---|---|
| Parameters | ~1.07B |
| Hidden Size | 2048 |
| Layers | 16 |
| Attention Heads | 32 (8 KV heads - GQA) |
| Context Length | 131,072 tokens |
| Vocabulary | 128,262 tokens |
| Precision | bfloat16 |
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "theaicompany02/Shivik-1B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True
)
# Math problem
prompt = """Question: A store sells apples for $2 each. If John buys 5 apples and pays with a $20 bill, how much change does he get?
Answer: Let me solve this step by step.
"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=256,
temperature=0.7,
do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Apache 2.0
Built as part of the SHIVIK project - creating competitive small language models through intelligent training.