Downloads Β· 30 days
14
10% of all-time downloads
gj5520/KoalaSeg
KoalaSeg is a image segmentation model from gj5520. Use it for the image segmentation task on the model card, and read the license before you ship it in a product. It is set up for transformers.
[](https://colab.research.google.com/drive/1LXWqtv-7lba128iEzhgSwXtEzRpRF7I0?usp=sharing)
Downloads Β· 30 days
14
10% of all-time downloads
All-time downloads
146
Public
Parameters
219M
882 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors879 MB Β· 99%
How the weights are stored.
F32219M Β· 100%
From the Hugging Face model README
KOrean lAyered assistive Segmentation

νκ΅ λλ‘·보ν νκ²½ μ μ© Universal Segmentation λͺ¨λΈμ
λλ€.
shi-labs/oneformer_cityscapes_swin_large κΈ°λ° OneFormer κ΅μ¬ λͺ¨λΈμ
shi-labs/oneformer_cityscapes_swin_largeAIHUB μΈλ·보ννκ²½ λ°μ΄ν° (https://aihub.or.kr/aihubdata/data/view.do?dataSetSn=189):
μ΄ 18,369μ₯ (AIHUB 5.5k + μκ° μ΄¬μ 9k + Street View 3.7k) λ μ΄μ΄ μμλΈ β
Morph Open/Close + MedianBlur(17px) ν GT μμ±.
| Device | Baseline Cityscapes | Ensemble (3-layer) | Custom (K-Road) | koalaseg |
|---|---|---|---|---|
| A100 | 3.58 s β 0.28 FPS | 3.74 s β 0.27 FPS | 0.15 s β 6.67 FPS | 0.14 s β 7.25 FPS |
| T4 | 5.61 s β 0.18 FPS | 6.01 s β 0.17 FPS | 0.39 s β 2.60 FPS | 0.31 s β 3.27 FPS |
| CPU (i9-12900K) | 124 s β 0.008 FPS | 150 s β 0.007 FPS | 26.6 s β 0.038 FPS | 18.4 s β 0.054 FPS |
from transformers import AutoProcessor, AutoModelForUniversalSegmentation
import torch, requests, matplotlib.pyplot as plt, numpy as np
from PIL import Image
from io import BytesIO
# 0. Load model & processor -----------------------------------
model_id = "gj5520/KoalaSeg"
proc = AutoProcessor.from_pretrained(model_id)
model = AutoModelForUniversalSegmentation.from_pretrained(model_id).to("cuda").eval()
# 1. Download image -------------------------------------------
url = "https://pds.joongang.co.kr/news/component/htmlphoto_mmdata/202205/21/1200738c-61c0-4a51-83c4-331f53d4dcdc.jpg"
resp = requests.get(url, stream=True)
img = Image.open(BytesIO(resp.content)).convert("RGB")
# 2. Pre-process & inference ----------------------------------
inputs = proc(images=img, task_inputs=["semantic"], return_tensors="pt").to("cuda")
with torch.no_grad():
out = model(**inputs)
# 3-A. Get class-id map ---------------------------------------
idmap = proc.post_process_semantic_segmentation(
out, target_sizes=[img.size[::-1]]
)[0].cpu().numpy()
# 3-B. Convert idmap β RGB mask + overlay ---------------------
cmap = plt.get_cmap("tab20", max(20, len(np.unique(idmap))))
mask_rgb = np.zeros((*idmap.shape, 3), dtype=np.uint8)
for idx, cid in enumerate(np.unique(idmap)):
if cid == 0: # keep background black
continue
mask_rgb[idmap == cid] = (np.array(cmap(idx)[:3]) * 255).astype(np.uint8)
mask_img = Image.fromarray(mask_rgb)
overlay = Image.blend(img, mask_img, alpha=0.6) # 0.6 β mask κ°μ‘°
# 4. Show overlay ---------------------------------------------
plt.figure(figsize=(8, 8))
plt.imshow(overlay)
plt.axis("off")
plt.show()
@misc{KoalaSeg2025, title = {KoalaSeg: Layered Distillation for Korean Road Universal Segmentation}, author = {RoadSight Team}, year = {2025}, url = {https://huggingface.co/gj5520/KoalaSeg} }