Downloads · 30 days
411
100% of all-time downloads
ElytronAI/DeepSeek-v4-Flash
DeepSeek-v4-Flash is a image-text-to-text model from ElytronAI. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for pulsar. The card lists the license as mit.
Pulsar's build of DeepSeek-v4-Flash, derived from deepseek-ai/DeepSeek-V4-Flash-Vision-Exp at revision 6821d6ad3681a4b137b066b76094fa82ebd0a380.
Downloads · 30 days
411
100% of all-time downloads
All-time downloads
411
Public
Parameters
166B
168 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors168 GB · 100%
How the weights are stored.
U8164B · 99%
From the Hugging Face model README
Pulsar's build of DeepSeek-v4-Flash, derived from
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp at revision
6821d6ad3681a4b137b066b76094fa82ebd0a380.
This is the unqualified member of the family: the source-precision build.
Every routed expert is stored in cutlass_mxfp4, repacked from the upstream
checkpoint's own FP4 bytes — nothing is quantized down, and there is no
iq2_xxs_mmq_k tensor anywhere. Variants that deviate from source precision
carry a suffix (-IQ2, -Mix-…); the bare name always means this one.
At ~168 GB it does not fit a single GB10. It is the artifact for a two-Spark
rig; the 2-bit build (DeepSeek-v4-Flash-IQ2, ~92 GB) is the single-Spark one.
| Layout | Used by | Bytes |
|---|---|---|
cutlass_mxfp4 | main routed experts (43 layers × 256) and drafter experts (3 × 256) | per-expert E2M1 data then a swizzled E8M0 group-32 SF tile |
mxfp8_lt | attention / dense / shared-expert projections | de-interleaved E4M3 then a swizzled E8M0 group-32 scale |
fp8_e4m3_soa_k | one drafter head projection | E8M0 scale plane [rows][cols/32] then an E4M3 payload plane |
native | norms, router, embeddings, hyper-connection params | plain BF16 / F32 / I32 as declared in the header |
Engine-layout tensors are stored as flat U8 blobs. Each shard's
__metadata__ is self-describing:
pulsar.tensors — per tensor: layout and dims_ne (the logical shape in
ggml ne[] order, which is the HF shape reversed; recorded for every tensor
so a reader never has to reason about orientation).pulsar.experts — per routed-expert projection: n_experts, expert_bytes,
layout, contiguous, and the original stacked tensor name. The expert
tensors of a projection are written back-to-back with no inter-expert
padding, so a kernel can address them as base + xid * expert_bytes from a
single pointer.pulsar.kv_arch — the 53 architecture keys the engine reads. pulsar.kv
(primary shard only) additionally carries the tokenizer block.format — pt, and pulsar.alignment — 32. Every tensor offset is
aligned to 32 bytes, which the swizzled scale tiles depend on.| component | upstream | this artifact | delta |
|---|---|---|---|
| routed-expert payloads | 157.437 GB | 157.437 GB | 0 |
| everything else | 10.374 GB | 10.529 GB | +155.2 MB |
| total | 167.819 GB | 167.979 GB | +0.09% |
The expert payloads are byte-count identical — 157,437,394,944 B in both — and
a pure permutation: the upstream E2M1 nibble plane is copied verbatim and the
upstream E8M0 group-32 plane is scattered into the CUTLASS SFB swizzle. The
+155 MB is the dense scale plane: mxfp8_lt carries one E8M0 scale per 32
columns, where the upstream FP8 weight block is 128×128. That re-encode is
lossless to within E4M3 rounding on 0.0013% of dense elements (max |Δ| 1.5e-5).
This runs on pulsar only. Stock transformers / vLLM cannot load it — they
have no decoder for the fused layouts, and U8 blobs have no meaning without
the metadata above. config.json's quantization_config.quant_method is
pulsar-safetensors precisely so a standard loader fails loudly instead of
misreading the bytes.
cutlass_mxfp4 and contiguous.iq2.See BUILD-RECORD.md for the full provenance and gate list.