Downloads · 30 days
17
3% of all-time downloads
North-ML1/willow-alpha-base
willow-alpha-base is a text generation model from North-ML1. Use it when you need the model to write or continue text. The card lists the license as mit.
<h1 align="center" style="font-size: 54px;" Willow Alpha </h1
Downloads · 30 days
17
3% of all-time downloads
All-time downloads
573
Public
Parameters
287M
1.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 100%
From the Hugging Face model README
Willow Alpha is an early-stage base model checkpoint in the Forge-1V model line.
This model is currently experimental and should be treated as a research checkpoint rather than a polished assistant model. It is useful for testing architecture, pretraining quality, tokenizer behavior, evaluation pipelines, and future SFT/RLHF improvements.
| Field | Value |
|---|---|
| Model name | Willow Alpha |
| Project | Forge-1V |
| Organization | North ML |
| Model type | Causal Language Model |
| Language | English |
| License | MIT |
| Status | Early-stage / Alpha |
All benchmarks below were run in 0-shot mode.
| Benchmark | Metric | Score | Runtime |
|---|---|---|---|
| HellaSwag | acc_norm | 26.71% | 318.67s |
| PIQA | acc_norm | 53.86% | 38.85s |
| WinoGrande | acc | 50.67% | 23.73s |
| BoolQ | acc | 40.21% | 144.80s |
| ARC-Easy | acc_norm | 34.68% | 51.41s |
| ARC-Challenge | acc_norm | 25.60% | 37.69s |
| OpenBookQA | acc_norm | 25.00% | 21.14s |
| CommonsenseQA | acc | 20.31% | 27.66s |
| LAMBADA | acc | 0.23% | 96.28s |
| BLiMP | acc | 59.23% | 354.79s |
| MMLU | acc | 23.89% | 388.62s |
| WikiText-2 | word_perplexity | 12524.42 | 182.89s |
| WikiText-2 | byte_perplexity | 5.84 | 181.42s |
| SciQ | acc_norm | 35.60% | 87.15s |
| COPA | acc | 64.00% | 17.21s |
| RACE | acc | 23.16% | 334.70s |
| SWAG | acc_norm | 29.13% | 252.00s |
| TruthfulQA MC2 | acc | 48.74% | 126.29s |
| Category | Result |
|---|---|
| Number of completed benchmark runs | 18 |
| Successful runs | 18 |
| Failed runs | 0 |
| Best accuracy-style score | COPA — 64.00% |
| Best language-structure score | BLiMP — 59.23% |
| MMLU score | 23.89% |
| WikiText-2 byte perplexity | 5.84 |
| WikiText-2 word perplexity | 12524.42 |
Willow Alpha is still in a very early stage. Some results are near-random or unstable, especially on knowledge-heavy and long-context tasks.
The strongest early signals are:
The weakest areas are:
These results suggest the model has some early reasoning and grammar signal, but still needs substantially more pretraining, higher-quality data, and post-training before being useful as a general assistant.
Willow Alpha is intended for:
It is not yet recommended for production use.
This model may:
@misc{willow-alpha,
title = {Willow Alpha},
author = {North ML},
year = {2026},
note = {Early-stage Forge-1V checkpoint}
}