Downloads ยท 30 days
239
38% of all-time downloads
WhirlwindAI/MetaNova-1-60M
MetaNova-1-60M is a text generation model from WhirlwindAI. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Downloads ยท 30 days
239
38% of all-time downloads
All-time downloads
634
Public
Parameters
62.7M
125 MB on disk
Likes
8
Trending 1
Click a slice to open those files.
.safetensors125 MB ยท 98%
From the Hugging Face model README
<div align="center">MetaNova-1 leads the series on ArithMark-2 โ outperforming GPT-2 (124M) with less than half the parameters.
| Benchmark | MetaNova-Test | MetaNova-0.1 | MetaNova-1 | Supra-50M-Base | OpenAI/GPT-2 |
|---|---|---|---|---|---|
| Params | 70.55M | 70.55M | 62.70M | 51.79M | 124M |
| HellaSwag | 25.37% | 27.40% | 27.08% | 31.65% | 31.26% |
| ARC-Easy | 26.89% | 31.73% | 30.56% | 45.58% | 39.35% |
| ARC-Challenge | 25.09% | 25.00% | 24.15% | 24.66% | 22.35% |
| PIQA | 52.29% | 58.54% | 56.04% | 61.53% | 62.08% |
| ArithMark-2 | 24.12% | 27.12% | 34.12% | 26.08% | 26.48% |
| ARC Avg | 25.99% | 28.37% | 27.35% | 35.12% | 30.85% |
| Final Avg | 31.94% | 35.36% | 36.15% | 38.60% | 37.67% |
MetaNova-1 uses the DeepSeek / Qwen im_start format with optional thinking mode.
<|im_start|>user
{query} /think<|im_end|>
<|im_start|>assistant
<think>
{thinking_content}
</think>
{response}<|im_end|>
</details>
<details>
<summary><b>โก Non-Thinking Mode</b> โ append <code>/no_think</code></summary>
<|im_start|>user
{query} /no_think<|im_end|>
<|im_start|>assistant
<think>
</think>
{response}<|im_end|>
</details>
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "WhirlwindAI/MetaNova-1-60M"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
prompt = "<|im_start|>user\nWhat is 14 ร 27? /think<|im_end|>\n<|im_start|>assistant\n"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0], skip_special_tokens=False))
| Architecture | Causal Language Model |
| Parameters | 62.70M |
| Organization | WhirlwindAI |
| License | Apache 2.0 |
| Language | English |
| Chat Format | DeepSeek / Qwen (im_start) |
| Thinking Mode | Yes (/think / /no_think) |
<sub>Built by <a href="https://huggingface.co/WhirlwindAI">WhirlwindAI</a> ยท Apache 2.0</sub>
</div>