Pith. sign in

REVIEW 3 cited by

Surgical-VQLA: Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.11692 v1 pith:GI6YAQTW submitted 2023-05-19 cs.CV cs.AIcs.CLcs.LGcs.RO

classification cs.CVcs.AIcs.CLcs.LGcs.RO
keywords surgicalanswerembeddingquestionquestion-answeringsurgical-vqlavideosvisual
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite the availability of computer-aided simulators and recorded videos of surgical procedures, junior residents still heavily rely on experts to answer their queries. However, expert surgeons are often overloaded with clinical and academic workloads and limit their time in answering. For this purpose, we develop a surgical question-answering system to facilitate robot-assisted surgical scene and activity understanding from recorded videos. Most of the existing VQA methods require an object detector and regions based feature extractor to extract visual features and fuse them with the embedded text of the question for answer generation. However, (1) surgical object detection model is scarce due to smaller datasets and lack of bounding box annotation; (2) current fusion strategy of heterogeneous modalities like text and image is naive; (3) the localized answering is missing, which is crucial in complex surgical scenarios. In this paper, we propose Visual Question Localized-Answering in Robotic Surgery (Surgical-VQLA) to localize the specific surgical area during the answer prediction. To deal with the fusion of the heterogeneous modalities, we design gated vision-language embedding (GVLE) to build input patches for the Language Vision Transformer (LViT) to predict the answer. To get localization, we add the detection head in parallel with the prediction head of the LViT. We also integrate GIoU loss to boost localization performance by preserving the accuracy of the question-answering model. We annotate two datasets of VQLA by utilizing publicly available surgical videos from MICCAI challenges EndoVis-17 and 18. Our validation results suggest that Surgical-VQLA can better understand the surgical scene and localize the specific area related to the question-answering. GVLE presents an efficient language-vision embedding technique by showing superior performance over the existing benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

    cs.CV 2025-01 reject novelty 5.0 of 10

    EndoChat is a grounded multimodal LLM for endoscopic surgery, trained on the new Surg-396K dataset and reported to outperform prior MLLMs, though its evaluation is confounded by training-data overlap.

  2. Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs

    cs.CV 2026-07 conditional novelty 4.5 of 10

    Multi-task LoRA fine-tuning with Grad-CAM grounding and terminology-free descriptions raises small-VLM GI VQA accuracy and implicit answer-to-region alignment on in- and out-of-distribution data.

  3. EndoARSS: Adapting Spatially-Aware Foundation Model for Efficient Activity Recognition and Semantic Segmentation in Endoscopic Surgery

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A DINOv2-based multi-task framework with task-specific low-rank adapters and a spatial attention module reports state-of-the-art joint activity recognition and semantic segmentation on three endoscopic surgery datasets.

Pith tools