Downloads · 30 days
36
56% of all-time downloads
sookie90/indic-transcribe-flex-coreml
indic-transcribe-flex-coreml is a automatic speech recognition model from sookie90. Use it when you need speech turned into text. The card lists the license as other.
A CoreML port of bodhan-ai/indic-transcribe-flex (1.22B params, FastConformer encoder + Transformer decoder, fine-tuned from nvidia/canary-1b-v2), converted for on-device inference on Apple Silicon (macOS 15+). Produc…
Downloads · 30 days
36
56% of all-time downloads
All-time downloads
64
Public
Repo size
3.6 GB
Likes
0
Public
Click a slice to open those files.
.bin3.6 GB · 100%
From the Hugging Face model README
A CoreML port of bodhan-ai/indic-transcribe-flex
(1.22B params, FastConformer encoder + Transformer decoder, fine-tuned from
nvidia/canary-1b-v2), converted for on-device inference on Apple Silicon
(macOS 15+). Produced for Veere as Task 84 of an
internal engineering plan — a spike to see whether this model could run as
Veere's background re-transcription pass without leaving the device.
Built with Indic-Transcribe-Flex from Bodhan AI / AI4Bharat.
Three CoreML models (.mlpackage, fp16 compute), converted directly from
the original fp32 checkpoint with a from-scratch conversion script — no
existing Canary CoreML port covers this architecture's decoder (see
"Provenance" below):
| File | Role | I/O |
|---|---|---|
indic_flex_preprocessor.mlpackage | mel front-end (STFT + filterbank + per-utterance normalization) — FLOAT32 compute precision, matching the original's own "runs in fp32 regardless of model dtype" | audio [1,480000] f32 (30s @ 16kHz, pad short audio with trailing silence) + sample_lens [1] i32 → features [1,128,3001] f32, feat_lens [1] i32 |
indic_flex_encoder.mlpackage | FastConformer encoder, 32 layers | input_features [1,128,3001] f32 + attention_mask [1,3001] i32 → encoder_states [1,376,1024] f32, encoder_lengths [1] i32 |
indic_flex_decoder_stateful.mlpackage | Transformer decoder, 24 layers, cross-attention to the encoder output, autoregressive with a CoreML State self-attention KV cache (macOS 15+) | input_ids [1,Q] i32, position_ids [1,Q] i32, encoder_states [1,376,1024] f32, cross_mask [1,1,1,376] f32, self_attn_mask [1,1,Q,end_step] f32 (Q, end_step: RangeDim) → logits [1,Q,7152] f32; 48 State tensors (self_k_cache_0..23, self_v_cache_0..23, each [1,8,320,128] fp16) |
vocab.json | flat, id-indexed list of 7,152 SentencePiece piece strings | detokenize with ''.join(pieces).replace('▁', ' ').strip() — no SentencePiece binding needed |
tokenizer_config.json | copied verbatim from the source checkpoint: special token ids + the fixed 10-token canary2 prompt per language | e.g. hi native = [7, 4, 18, 89, 89, 5, 9, 11, 13, 15]; romanized swaps index 7 from 11 to 10 |
Indic_Open_Model_License.md | the license text, verbatim | required by the license's Distribution clause (§5.3) |
Fixed 30-second input shape. Every chunk must be exactly 480,000 samples
(30s @ 16kHz); pad shorter audio with trailing silence and pass the true
sample count in sample_lens. This matches the checkpoint's own training
max_duration and the chunking already used to measure its real accuracy
(see "Measured numbers" below) — the original PyTorch wrapper explicitly
does not chunk internally and degrades sharply past ~45s, so a caller must
chunk before calling this model, not after.
Cross-attention is recomputed every decode step, not cached as a second
State. The original model caches cross-attention K/V once per utterance;
since it is a pure, deterministic function of encoder_states (fixed for
the whole utterance), recomputing it every step is numerically identical
and removes the one piece of this conversion that would have needed
step-conditional state writes — torch.jit.trace cannot express those at
all. See the conversion script's module docstring for the full reasoning.
No existing Canary or FastConformer-AED-decoder CoreML port covers this
architecture — checked against vendor/FluidAudio (no Canary* manager at
all: Parakeet/TDT, Qwen3, Cohere, Paraformer, SenseVoice only) and against
FluidInference/canary-1b-v2-coreml
(the base model's own community port — a 4-component split: Preprocessor,
EncoderInt4, DecoderInt4, Projection; useful as a component-split
precedent, not reused directly here, and not weight-compatible with this
fine-tune regardless). The conversion technique — register_buffer KV-cache
tensors turned into CoreML States via ct.convert(..., states=[ct.StateType(...)]),
torch.jit.trace (no data-dependent branching survives State conversion),
explicit matmul+softmax attention, .cpuAndGPU compute units — is the same
one used earlier in this project to port Qwen3-ASR to CoreML
(FluidInference's own mobius converter, adapted for a decoder-only model);
this port adapts it to an encoder-decoder architecture with real
cross-attention, which that precedent did not need to handle.
Conversion script: tools/coreml/convert_indic_transcribe.py. Full technical
writeup, including every numeric parity check run before this was trusted:
see this project's reviews/2026-09-09-indic-transcribe-coreml-port.md.
.mlpackage; no int4/int8 pass
attempted).Full numbers, methodology, and every blocker hit: reviews/2026-09-09-indic-transcribe-coreml-port.md
in this project's repo.
Indic Open Model License v1.0 (full text: Indic_Open_Model_License.md
in this repo, copied verbatim from
Bodhan-AI/bodhan-model-info).
This is a derivative of bodhan-ai/indic-transcribe-flex (format
conversion + quantization are explicitly derivative-creating acts under the
license's own definitions) and is distributed under the same license,
as the license's §5.1 share-alike clause requires. Key points for anyone
building on this repo — not a substitute for reading the license itself:
nvidia/canary-1b-v2 carries its own separate license —
ensure your use complies with both.int8/, 2026-09-10)Weight-only linear per-channel int8 of the same fp16 encoder + stateful decoder (preprocessor stays fp16),
via coremltools.optimize.coreml. Same 200-clip CoSHE dev set, 30 s chunking, romanised mode:
| variant | bundle | script-blind WER [95 % CI] | strict | RTF | compute units |
|---|---|---|---|---|---|
| fp16 (root) | 2.29 GB | 13.40 % [11.9–14.8] | 43.6 % | 0.39 | CPU_AND_GPU |
| int8 | 1.1 GB | 14.04 % [12.2–15.9] | 43.8 % | 0.24 | ALL (encoder 98.8 % on the Neural Engine) |
| int4 (not published) | 632 MB | 14.42 % [12.9–16.0] | 44.5 % | 0.37 | ALL — no gain: the stateful decoder cannot run on the Neural Engine |
Load every component with compute_units=ALL: the FastConformer encoder runs almost entirely on the Neural
Engine; the decoder's CoreML State-API ops (the KV cache) have no Neural Engine support and run on the GPU.
Licence unchanged (Indic Open Model License v1.0, see the licence file at the root). Recipe and measurements:
tools/coreml/convert_indic_transcribe.py --quantize int8, tools/coreml/ane_check.py in the Veere repo.