Downloads · 30 days
0
Orvyth/engrym-seed-flash-4b
engrym-seed-flash-4b is a text generation model from Orvyth. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
Part of Engrym Seed — Orvyth's seed-tier local model family for tool-using agents. Qwen3.5 hybrid linear-attention architecture, 262,144-token native context, first-class tool calling.
Downloads · 30 days
0
Access
Public
Updated Aug 15, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md2.9 KB · 65%
From the Hugging Face model README
Part of Engrym Seed — Orvyth's seed-tier local model family for tool-using agents. Qwen3.5 hybrid linear-attention architecture, 262,144-token native context, first-class tool calling.
The best agent build — fastest to act, and the value pick of the ladder.
Weights are distributed via the Ollama registry.
ollama run Orvyth/engrym-seed:flash
| Size | 4.6 GB |
| 77-task score | 124/143 (86.7%) |
| Tool calling | 12/12 verified |
| Context | 262,144 native (default num_ctx 32,768) |
| Quantization | Q8_0 |
| Tag | Class | Size | 77-task |
|---|---|---|---|
:nano | Nano 2B | 2.1 GB | 90.8/143 |
:flash | Flash 4B | 4.6 GB | 124/143 |
:base | Base 9B | 9.5 GB | 131/143 |
:pro-27b-q4 | Pro 27B v2 Q4 | 16.5 GB | 134/143 |
:pro | Pro 27B v2 Q8 | 28.6 GB | 134/143 |
:pro-e | Pro-E 27B (experimental) | 28.6 GB | 137/143 |
77 tasks · 143 points · temperature=0 · max_tokens=16384 · seed=42 · one attempt ·
deterministic validators · no LLM judge. Scores are bound to the exact published blobs.
First-party numbers. Repeated runs on an uncontended GPU are deterministic (zero spread across n=2 for every model measured). Receipts: https://huggingface.co/datasets/Orvyth/engrym-seed-receipts
Asking the model to work deliberately recovers points on tasks it otherwise fails, and the gain is largest for the smallest models (Nano +11.8, Flash +6, Base +3, Pro-E +0). On the small end that is worth more than a model upgrade.
temperature 0.2 · top_p 0.9 · top_k 20 · num_ctx 32768 · num_predict 8192
If a prompt exceeds num_ctx, Ollama returns HTTP 400 — it does not silently truncate.
Qwen/Qwen3.5 base → Ornith-1.0-9B x Qwythos-9B TIES merge (9B line) → Orvyth identity and
chip-calling LoRA merged into the weights → converted and quantized in-house.
ORVYTH - Intelligence. Governed. Ground truth over hype. Prove before you claim.