Downloads · 30 days
408
53% of all-time downloads
SurjoLabs/Spark-2A
Spark-2A is a text generation model from SurjoLabs. Use it when you need the model to write or continue text. The card lists the license as mit.
Spark-2 is a high-efficiency, sub-10M parameter language model developed by SurjoLabs. With an Intelligence Index of 9.14, it proves that combining 3:1 Grouped-Query Attention (GQA), value-subtraction projection (XSA)…
Downloads · 30 days
408
53% of all-time downloads
All-time downloads
771
Public
Parameters
9.1M
3.1 GB on disk
Likes
4
Public
Click a slice to open those files.
.bin1.7 GB · 53%
From the Hugging Face model README
Spark-2 is a high-efficiency, sub-10M parameter language model developed by SurjoLabs. With an Intelligence Index of 9.14, it proves that combining 3:1 Grouped-Query Attention (GQA), value-subtraction projection (XSA), and deep recurrent weight-sharing yields exceptional reasoning capabilities at micro scale.
Spark-2 demonstrates that specialized recurrent architectures can match or exceed standard transformers multiple times their size when trained on dense, high-quality data.
Evaluated at 0-shot using normalized accuracy (acc_norm):
| Benchmark | Spark | Spark-2 |
|---|---|---|
| HellaSwag | 28.17% | 28.13% |
| ARC-Easy | 35.02% | 34.26% |
| ARC-Challenge | 20.99% | 22.18% |
| PIQA | 55.55% | 57.24% |
| ArithMark-3.0 | 36.10% | 37.00% |
| Intelligence Index | 7.93 | 9.14 |
The model was trained for 20,000 steps with a WSD scheduler (decay phase starting at step 16,000). Extensive checkpoint evaluation revealed that checkpoint-17500 achieved peak consolidation across math, commonsense, and logic benchmarks (9.14 INT INDEX). The final step 20,000 was discarded, and step 17,500 is the official version hosted in the root of this repository.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "SurjoLabs/Spark-2"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16
).cuda()
prompt = "The capital of France is"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=20)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
We would like to thank AxiomicLabs for pioneer work proving the token-efficiency of the XSA attention mechanism.
Spark-2 is an experimental sub-10M parameter model developed under the broader Surjo Project. While it demonstrates strong reasoning for its size, language generation length and coherence remain constrained by absolute capacity. The modeling files are intended for research and benchmark evaluation.