Downloads · 30 days
30
2% of all-time downloads
oscowlai/Wiola13M
Wiola13M is a text generation model from oscowlai. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
30
2% of all-time downloads
All-time downloads
1.8K
Public
Parameters
12.9M
465 MB on disk
Likes
5
Public
Click a slice to open those files.
.pt310 MB · 60%
From the Hugging Face model README

Wiola is a novel decoder-only Small Language Model (SLM) developed by OSCOWL-AI. It introduces several architectural improvements over conventional Transformer decoder models, focusing on improving contextual reasoning, computational efficiency, and parameter utilization while remaining compatible with the Hugging Face ecosystem.
Note
Wiola uses a custom architecture. Before loading the model, install the Wiola package:
pip install git+https://github.com/Wiola-OSCOWL-ai/Wiola13M.gitor (once available)
pip install wiola
Wiola is designed as a research-focused decoder-only language model that explores new methods for improving attention, positional encoding, and feed-forward computation.
The architecture introduces five core innovations:
These components are integrated into a standard autoregressive language modeling framework and are compatible with Hugging Face Transformers.
Install directly from GitHub:
pip install git+https://github.com/Wiola-OSCOWL-ai/Wiola13M.git
or after the PyPI release:
pip install wiola
Create a new Python file (for example, test.py) and paste the following code into it:
from wiola13m import WiolaForCausalLM
from transformers import AutoTokenizer
model_path = "oscowlai/Wiola13M"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = WiolaForCausalLM.from_pretrained(model_path)
prompt = "Once upon a time"
inputs = tokenizer(
prompt,
return_tensors="pt",
return_token_type_ids=False,
)
outputs = model.generate(
**inputs,
max_new_tokens=100,
do_sample=True,
temperature=0.8,
top_p=0.95
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Run the script from your terminal:
python test.py
The first time you run the script, the model and tokenizer will be downloaded from Hugging Face and stored in the local cache. This may take a few moments depending on your internet connection. Subsequent runs will use the cached files and start much faster.
To generate text for a different prompt, replace:
prompt = "Once upon a time"
with any text you would like the model to continue.
For example:
prompt = "Artificial Intelligence will"
prompt = "Write a short story about a robot."
prompt = "Explain gravity in simple terms."
Running the script prints the generated text to your terminal.
Example:
Once upon a time, a simple idea turned into a powerful innovation. It all started with coding ...
Since text generation uses sampling (do_sample=True), the generated output will be different each time you run the script.
The model accepts tokenized text.
Input shape:
(batch_size, sequence_length)
Example:
inputs = tokenizer(
"Hello, how are you?",
return_tensors="pt"
)
The model returns:
For text generation, use:
outputs = model.generate(...)
The generated token IDs can be converted back into readable text using:
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Wiola is a decoder-only Transformer architecture composed of:
Each decoder layer consists of:
The released checkpoints are trained from scratch using the Wiola architecture.
The architecture supports training on standard causal language modeling corpora such as:
The released checkpoints may use different datasets depending on the model version.
Training data is cleaned using standard preprocessing:
The model is designed to be evaluated using common language modeling benchmarks including:
Benchmark results will be released alongside future checkpoints.
Training:
Inference:
CPU:
GPU:
Wiola is a research model and has several limitations:
Suitable for:
Not recommended for:
Like other large language models, Wiola may inherit biases from training data.
Developers should:
Input Tokens
│
▼
Token Embeddings
│
▼
┌─────────────────────────────────────┐
│ Spiral Rotary Encoding │
│ Gated Cross-Layer Attention │
│ Adaptive Token Merging │
│ Dual-Stream FFN │
│ WiolaRMSNorm │
└─────────────────────────────────────┘
│
▼
Repeated N Layers
│
▼
Final RMSNorm
│
▼
LM Head
│
▼
Generated Tokens
If you use Wiola in your research, please cite:
@software{wiola2026,
title={Wiola: A Novel Decoder-Only Small Language Model},
author={OSCOWL-AI},
year={2026},
url={https://github.com/Wiola-OSCOWL-ai/wiola}
}
Apache License 2.0
Wiola is an open research project developed by OSCOWL-AI to explore efficient and scalable Small Language Model architectures. Contributions, feedback, and research collaborations are welcome.