Downloads · 30 days
0
visurg/SurgLIME
SurgLIME is a machine learning model from visurg. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
[](https://arxiv.org/abs/2604.18134) [](https://huggingface.co/datasets/visurg/LIME) [](https://huggingface.co/visurg/SurgLIME/blob/main/surglime.pth)
Downloads · 30 days
0
Access
Public
Updated Aug 4, 2026
Repo size
794 MB
Likes
0
Public
Click a slice to open those files.
.pth794 MB · 100%
From the Hugging Face model README
This is the official repository for the CVPRW 2026 Oral paper: Can LLM-Generated Text Empower Surgical Vision-Language Pre-training?
Star ⭐ the repository if you find this work useful.
<p align="center"> <img src="https://github.com/user-attachments/assets/4d397e7e-3262-43ba-b970-c6213d5171c4" alt="SurgLIME overview" /> </p>Recent advancements in self-supervised learning have led to powerful surgical vision encoders capable of spatiotemporal understanding. However, extending these visual foundations to multimodal reasoning tasks is severely bottlenecked by the prohibitive cost of expert textual annotations.
To overcome this scalability limitation, we introduce LIME, a large-scale multimodal dataset derived from open-access surgical videos using human-free, Large Language Model (LLM)-generated narratives. While LIME offers immense scalability, unverified generated texts may contain errors, including hallucinations, that could potentially lead to catastrophically degraded pretrained medical priors in standard contrastive pipelines.
To mitigate this issue, we propose SurgLIME, a parameter-efficient Vision-Language Pre-training (VLP) framework designed to learn reliable cross-modal alignments from noisy narratives. SurgLIME preserves foundational medical priors using a LoRA-adapted dual-encoder architecture and introduces an automated confidence estimation mechanism that dynamically down-weights uncertain text during contrastive alignment.
Evaluations on the AutoLaparo and Cholec80 benchmarks show that SurgLIME achieves competitive zero-shot cross-modal alignment while preserving the robust linear-probing performance of the visual foundation model.
Clone the repository and install the required dependencies:
git clone [email protected]:visurg-ai/SurgLIME.git
cd SurgLIME
pip install -r requirements.txt