Downloads · 30 days
61
80% of all-time downloads
mlboydaisuke/clip-vit-base-patch32-CoreAI-official
clip-vit-base-patch32-CoreAI-official is a machine learning model from mlboydaisuke. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for coreai. The card lists the license as mit.
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or t…
Downloads · 30 days
61
80% of all-time downloads
All-time downloads
76
Public
Repo size
605 MB
Likes
0
Public
Click a slice to open those files.
.mlirb303 MB · 99%
From the Hugging Face model README
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta 26A5353q, 2026-06-11).
This model has no row on DeviceMark, the on-device LLM leaderboard.
<!-- gen-cards:devicemark end -->fp16 static export of openai/clip-vit-base-patch32
via apple/coreai-models' official recipe (models/clip/export.py), with one change: text
inputs are padded to the full 77-token context (padding="max_length") so free-text
queries work, instead of the recipe's 7-token example trace.
Runs out of the box with CoreAIKit's
ImageTextEncoder:
let encoder = try await ImageTextEncoder() // downloads this repo
let imageVec = try await encoder.encode(image: cgImage)
let textVec = try await encoder.encode(text: "red bike at the beach")
let score = ImageTextEncoder.cosineSimilarity(imageVec, textVec)
model/
├── clip-vit-base-patch32_float16_static.aimodel
└── tokenizer.json
| name | shape | dtype | |
|---|---|---|---|
| input | pixel_values | [1, 3, 224, 224] | fp16 |
| input | input_ids | [3, 77] | int32 |
| input | attention_mask | [3, 77] | int32 |
| output | image_embeds | [1, 512] | fp16, L2-normalized |
| output | text_embeds | [3, 512] | fp16, L2-normalized |
| output | logits_per_image / logits_per_text | [1, 3] / [3, 1] | fp16 |
Preprocessing: 224×224 resize + CLIP mean/std normalization (handled by
ImageTextEncoder).
M4 Max: ~3.7 ms per image on the Neural Engine (fp16). Requires macOS 27 / iOS 27 (device — the CoreAI framework is not in the iOS Simulator SDK).
Model weights: MIT (OpenAI CLIP); see the upstream repo. Export recipe: BSD-3-Clause (apple/coreai-models).
<!-- funnel:v1 -->More models in this format: Core AI Model Zoo — 75 models, each with the recipe that produced it.
Want a different model on-device? Open a request — free, open weights only; the export and its measured numbers get published publicly.
<!-- /funnel:v1 -->