Downloads · 30 days
157
7% of all-time downloads
Xenova/sam-vit-large
sam-vit-large is a mask generation model from Xenova. Use it for the mask generation task on the model card, and read the license before you ship it in a product. It is set up for transformers.js. The card lists the license as apache-2.0.
https://huggingface.co/facebook/sam-vit-large with ONNX weights to be compatible with Transformers.js.
Downloads · 30 days
157
7% of all-time downloads
All-time downloads
2.4K
Public
Repo size
4.4 GB
Likes
1
Trending 1
Click a slice to open those files.
.onnx4.4 GB · 100%
From the Hugging Face model README
https://huggingface.co/facebook/sam-vit-large with ONNX weights to be compatible with Transformers.js.
If you haven't already, you can install the Transformers.js JavaScript library from NPM using:
npm i @huggingface/transformers
Example: Perform mask generation with Xenova/sam-vit-large.
import { SamModel, AutoProcessor, RawImage } from "@huggingface/transformers";
// Load model and processor
const model = await SamModel.from_pretrained("Xenova/sam-vit-large");
const processor = await AutoProcessor.from_pretrained("Xenova/sam-vit-large");
// Prepare image and input points
const img_url = "https://huggingface.co/datasets/Xenova/transformers.js-docs/resolve/main/corgi.jpg";
const raw_image = await RawImage.read(img_url);
const input_points = [[[340, 250]]];
// Process inputs and perform mask generation
const inputs = await processor(raw_image, { input_points });
const outputs = await model(inputs);
// Post-process masks
const masks = await processor.post_process_masks(outputs.pred_masks, inputs.original_sizes, inputs.reshaped_input_sizes);
console.log(masks);
// [
// Tensor {
// dims: [ 1, 3, 410, 614 ],
// type: 'bool',
// data: Uint8Array(755220) [ ... ],
// size: 755220
// }
// ]
const scores = outputs.iou_scores;
console.log(scores);
// Tensor {
// dims: [ 1, 1, 3 ],
// type: 'float32',
// data: Float32Array(3) [
// 1.0122944116592407,
// 0.9184409976005554,
// 0.9796935319900513
// ],
// size: 3
// }
You can then visualize the generated mask with:
const image = RawImage.fromTensor(masks[0][0].mul(255));
image.save('mask.png');

Next, select the channel with the highest IoU score, which in this case is the first (red) channel. Intersecting this with the original image gives us an isolated version of the subject:

We've also got an online demo, which you can try out here.
<video controls autoplay src="https://cdn-uploads.huggingface.co/production/uploads/61b253b7ac5ecaae3d1efe0c/Y0wAOw6hz9rWpwiuMoz2A.mp4"></video>
Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using 🤗 Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).