REVIEW 4 cited by
PitVis-2023 Challenge: Workflow Recognition in videos of Endoscopic Pituitary Surgery
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery: including which surgical steps are performed; and which surgical instruments are used. This information can later be used to assist clinicians when learning the surgery; during live surgery; and when writing operation notes. The Pituitary Vision (PitVis) 2023 Challenge tasks the community to step and instrument recognition in videos of endoscopic pituitary surgery. This is a unique task when compared to other minimally invasive surgeries due to the smaller working space, which limits and distorts vision; and higher frequency of instrument and step switching, which requires more precise model predictions. Participants were provided with 25-videos, with results presented at the MICCAI-2023 conference as part of the Endoscopic Vision 2023 Challenge in Vancouver, Canada, on 08-Oct-2023. There were 18-submissions from 9-teams across 6-countries, using a variety of deep learning models. A commonality between the top performing models was incorporating spatio-temporal and multi-task methods, with greater than 50% and 10% macro-F1-score improvement over purely spacial single-task models in step and instrument recognition respectively. The PitVis-2023 Challenge therefore demonstrates state-of-the-art computer vision models in minimally invasive surgery are transferable to a new dataset, with surgery specific techniques used to enhance performance, progressing the field further. Benchmark results are provided in the paper, and the dataset is publicly available at: https://doi.org/10.5522/04/26531686.
Forward citations
Cited by 4 Pith papers
-
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
MedVideoCap-55K, a 55,803-clip caption-rich medical video dataset, enables MedGen, a LoRA fine-tune of HunyuanVideo that reports top open-source scores and near-commercial quality on medical video benchmarks.
-
Anatomy Might Be All You Need: Forecasting What to Do During Surgery
A transformer that reads anatomy and instrument detections from endoscopic video can forecast the instrument's next movement direction, with anatomy improving 8-frame direction accuracy from 50.7% to 60.6%.
-
Large-scale Self-supervised Video Foundation Model for Intelligent Surgery
SurgVISTA is a masked-reconstruction surgical video foundation model whose joint spatiotemporal pretraining plus expert distillation outperforms image-level and natural-video pretrained models on 13 surgical benchmarks.
-
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
A fine-tuned 8B medical vision-language model that claims explainable grounding, uncertainty estimates and cancer prognosis, but its core uncertainty formula is mathematically inconsistent and key comparisons use the ...
Discussion (0). Continue with ORCID to comment.