Downloads · 30 days
10
42% of all-time downloads
code-and-canvas/Walkyrie-1.3B-v2.0-CoreML-float16
Walkyrie-1.3B-v2.0-CoreML-float16 is a text-to-image model from code-and-canvas. Use it when you need an image from a text prompt. It is set up for diffusers. The card lists the license as apache-2.0.
This repository contains the first native Apple Silicon Core ML conversion of the Walkyrie-1.3B-v2.0 core transformer brain, an image model built on top of the Wan 2.1 Diffusion Transformer (DiT) framework.
Downloads · 30 days
10
42% of all-time downloads
All-time downloads
24
Public
Repo size
5.8 GB
Likes
0
Public
Click a slice to open those files.
.safetensors3 GB · 51%
From the Hugging Face model README
This repository contains the first native Apple Silicon Core ML conversion of the Walkyrie-1.3B-v2.0 core transformer brain, an image model built on top of the Wan 2.1 Diffusion Transformer (DiT) framework.
Walkyrie_1.3B_v2.0_float16.mlpackage: The complete 30-block core DiT transformer layer, fully optimized to execute on the Apple Neural Engine (ANE) and Apple Graphics Processor (GPU).This asset contains only the core transformer block. To build a complete text-to-image pipeline inside a native Swift application, you will need to pair this core package with a text tokenizer and a VAE decoder:
swift-tokenizers or mlx-swift.If you want to re-compile or modify this setup from scratch using the silicon-alloy converter or direct coremltools tracing, you must bypass several legacy architectural structural mismatches hardcoded into older diffusion conversion scripts.
The original codebase must be patched with the following workflow modifications:
The newer Wan 2.1 architecture uses updated property names. Legacy scripts searching for sub-modules will throw immediate AttributeErrors unless mapped to the following properties:
.transformer_blocks references to .blocks.patch_embed references to .patch_embeddingOlder models process prompt token arrays and timesteps via isolated .text_embed() and .time_embed() functions. Wan 2.1 consolidates these into a single unified block.
temb, timestep_proj, encoder_hidden_states, _ = self.model.condition_embedder(timestep, encoder_hidden_states, None)timestep_proj = timestep_proj.unflatten(1, (6, -1))The patch embedding layer outputs a 5D spatial video matrix structured as [Batch, Hidden_Dim, Frames, Height, Width]. The transformer blocks, however, expect a flattened 3D sequence token vector [Batch, Sequence_Length, Hidden_Dim]. Crucially, the Rotary Position Embedding (.rope) module still requires the 5D spatial layout to calculate coordinates.
.rope() module first to extract your rotary embedding parameters: image_rotary_emb = self.model.rope(hidden_states_5d)hidden_states = hidden_states_5d.flatten(2).transpose(1, 2)silicon-alloy framework.