← back to paper
arxiv: 2608.09302 · 2 revisions
Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation