REVIEW 3 major objections 5 minor 80 references
Fine-tuning a vision encoder with latent-action and forward-prediction losses, plus a patch-level distribution regularizer, improves surgical interaction recognition and grounds change features in instrument-tissue regions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 11:24 UTC pith:ZHMMNLWF
load-bearing objection A well-engineered fine-tuning recipe with consistent but modest gains; the latent-action mechanism story is weaker than the abstract implies, and the grounding gains are carried by the patch-SIG regularizer, not the IDM/FWM. the 3 major comments →
LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central claim is that adding latent-action supervision to vision-language fine-tuning reshapes a pretrained vision encoder toward interaction-relevant spatial and transition features, and that this yields better performance than standard fine-tuning of the same encoder. The authors' mechanism: an inverse dynamics model that operates on the difference between consecutive pooled states sends that change through a low-dimensional bottleneck to produce a compact latent action; the forward world model is conditioned on this action via adaptive layer normalization and must reconstruct the next state from a stop-gradient target, so the action must carry exactly the tra
What carries the argument
The load-bearing object is a cheap latent-action bottleneck: a difference-based inverse dynamics model maps the difference between consecutive frame states to a 128-dimensional action token, and a DiT-style forward world model uses that token as the only transition path to predict the next state from the current one. Because the action dimension is far smaller than the state dimension, the forward model cannot copy the next state through the action, so the encoder must compress frame-to-frame change into few informative directions. The second mechanistic piece is the patch-level SIG regularizer, a sliced Cramer-Wold Epps-Pulley test that drives randomly sampled per-patch token distributions
Load-bearing premise
The paper's central grounding claim rests on the untested premise that supervising the forward model with a single average feature per frame, together with a Gaussianity regularizer on random patches, is enough to push patch-level features toward the instrument-tissue interaction region.
What would settle it
Train the same framework but supervise the forward model with patch-level rather than pooled-state targets, or recompute the grounding IoU using manually annotated interaction masks instead of propagated ones; if the reported IoU gains vanish or reverse under either change, the paper's attribution of grounding to the state-level IDM/FWM loss and its evaluation both come into question.
If this is right
- Across a vanilla transformer encoder, a self-distillation image encoder, and a surgically pretrained variant of the same architecture, latent-action guidance improves triplet mean average precision over each encoder's standard fine-tuned baseline on both benchmarks, with the largest gains on the weakest encoder.
- An effective-dimension measure of the transition covariance drops substantially under IDM/FWM training (for example, from about 157 to 34 for the vanilla encoder), showing that inter-state change is compressed into fewer dominant directions.
- Without the patch-level regularizer, compression collapses further (effective dimensions around 13-18 and the leading-component count capturing 90 percent of variance falls as low as 19), while grounding IoU stagnates or drops; with it, grounding IoU rises to 0.074, 0.137, and 0.181 across the three encoders.
- Dropping action tokens at inference lowers triplet mAP (for example, from 17.22 to 15.79 on one benchmark and 23.58 to 21.84 on the other), so the latent action carries information beyond the training-time objective.
- Both too small and too large action dimensions weaken subspace-grounding IoU and cluster quality, suggesting an intermediate bottleneck size is needed to keep action-relevant transitions concentrated.
Where Pith is reading between the lines
- If the proposed mechanism transfers, the same inverse-dynamics plus forward-model plus patch-regularizer recipe could improve other fine-grained video-recognition tasks that need spatially grounded change features, such as hand-object interaction or robot manipulation, without region annotations.
- The paper supervises the forward model with mean-pooled frame states; a natural stress-test is to replace the pooled target with patch- or block-level targets. If grounding improves further, the current attribution understates the role of target granularity; if it does not, the pooled-signal premise is corroborated.
- Because the reported grounding numbers inherit the accuracy of the propagated segmentation masks used as ground truth, an evaluation on a small set of manually annotated frames would establish how much of the IoU gain survives mask-estimation error.
- The results show CH separability and grounding IoU can move in opposite directions (the patch regularizer sometimes lowers CH while raising mAP), so representation studies of fine-tuning should report cluster-quality and grounding metrics jointly rather than in isolation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes LAViFiT, a fine-tuning framework for surgical action-triplet recognition that couples a pretrained vision encoder with an inverse dynamics model (IDM), a forward world model (FWM), and a patch-level SIG regularizer. The IDM maps pooled-frame-state differences into low-dimensional latent actions; the FWM predicts the next pooled state from the current state and action; and SIG regularizers are applied to states, actions, and sampled patch tokens. The method is evaluated on CholecT50 and ProstaTD across ViT-L, DINOv2-L, and SurgeNet-L backbones, reporting gains in triplet mAP over standard fine-tuning and over a CLIP4Clip-style states-only baseline. The paper also presents mechanism analyses: PCA visualizations, Calinski-Harabasz separability, transition-covariance effective rank, and IoU of dominant change directions with an instrument-tissue interaction mask.
Significance. If the recognition claims hold, LAViFiT offers a practical fine-tuning recipe for surgical vision encoders that needs no region-level supervision and is specified with enough detail (App. A.1) to reproduce. The paper ships a public code link, uses external benchmarks and standard metrics, and its recognition ablations are internally consistent. The mechanism story, however, goes beyond what Table 2 supports: the IDM/FWM module alone does not improve the reported grounding IoU, and the patch-level regularizer is spatially agnostic, so the causal chain from latent-action compression to interaction-region grounding is not established. The paper's central contribution is defensible, but the overclaimed mechanism and the missing control condition require a major revision.
major comments (3)
- [§6, §5.4, Table 2] The conclusion states that 'compact IDM-FWM training ... improves their spatial alignment with action-relevant instrument-tissue regions.' Table 2 contradicts this: comparing the No-IDM&FWM row with the No-Patch-SIG row (which retains IDM/FWM), IoU_norm changes only 0.057→0.060 for ViT-L, 0.081→0.080 for DINOv2, and 0.111→0.096 for SurgeNet. The IDM/FWM module compresses transitions (PR/d90 drop) but does not by itself improve grounding; the large IoU gains appear only when patch-SIG is added. The supported claims are (a) IDM/FWM compresses transition covariance and (b) patch-SIG improves IoU; the claim that IDM/FWM improves spatial alignment must be re-stated or supported with additional evidence.
- [§3.2, Eqs. (7)-(8), Table 2] No ablation isolates the patch-SIG regularizer without IDM/FWM, so the paper cannot attribute the final grounding improvement to the combination. Ordinary patch-SIG is spatially agnostic: it samples 32 random patches per frame and matches them to an isotropic Gaussian (Eq. 8), with no exposure to the interaction-region mask M_act (Eq. 11) or any spatial conditioning. A control with 'No IDM/FWM + patch-SIG' would determine whether patch-SIG alone yields the IoU increases reported in Table 2. Without such a control, the causal narrative in §5.4 ('IDM/FWM compresses -> patch-SIG restores diversity -> grounding improves') is underdetermined and should be labeled a hypothesis.
- [§3.1, Eq. (1); §3.2, Eq. (5)] The IDM/FWM supervision is routed through frame-level pooled states s_t = pool(p_t), and the forward-prediction loss (Eq. 5) supervises only these pooled states. The latent action therefore carries no direct patch-level spatial information. Yet the paper claims that this pooled-state bottleneck shapes patch-level feature geometry and grounds transitions in the instrument-tissue region. This is a load-bearing gap: the mechanism by which a spatially impoverished state signal transfers to patch-level grounding is neither derived nor tested. The CH and IoU analyses are post-hoc evaluations; no training signal connects them to the interaction mask. Please either provide evidence for this transfer (e.g., show action-token dependence on per-patch motion) or explicitly acknowledge it as an untested premise.
minor comments (5)
- [Table 1, §4.1] The main table compares single-frame image-encoder baselines with an 8-frame LAViFiT pipeline. Although a states-only CLIP4Clip-style baseline is provided in the ablations, the main-table caption should clearly state that the standard fine-tuning baselines for image encoders use a single frame while the method uses an 8-frame clip, to avoid misleading readers about the source of temporal gain.
- [§1, §4, Table 1] Text in §4.1 refers to 'fine-tuned CLIP ViT-L/16 baseline' while the table and method use ViT-L/14. Please correct the mismatched notation and specify the exact patch size used for each encoder.
- [Eq. (10) vs. Eq. (2)] The symbol v is used both for the aggregated clip embedding (Eq. 2) and for the verb component (Eq. 10, y_t = (i_t, v_t, r_t)). This is confusing in a paper where verb AP is a reported metric; please disambiguate.
- [§5, Table 1] CholecT50 results are from a single run, and the SurgeNet-L improvement over standard fine-tuning is small (18.04→18.50). ProstaTD results include means and std across folds, but no paired significance or confidence intervals are reported. Please add per-fold numbers or a statistical test to support the claim of consistent gains, particularly for the smallest effect.
- [Table 1, No action tokens at inference] Removing action tokens only at inference is an input-corruption test, not a training-ablation: the model was optimized with action tokens and then evaluated without them. The interpretation should be framed accordingly, not as evidence that action tokens are the sole driver of training.
Circularity Check
No significant circularity: recognition and grounding claims rest on external benchmarks, evaluation masks, and ablations; self-citations are contextual only.
full rationale
LAViFiT's central recognition result is an empirical comparison against external benchmarks (CholecT50 RDV split, ProstaTD 5-fold) using the standard ivtmetrics mAP; no fitted parameter is defined in terms of the reported mAP. The grounding analysis uses SASVI-propagated CholecSeg8k masks and ground-truth triplet instrument/target annotations only as an evaluation region (Eq. 11), not as supervision in the training objective (Eq. 9); the patch-level SIG regularizer (Eq. 8) samples random patches and matches them to an isotropic Gaussian with no access to M_act, so the reported IoU_norm improvements are not forced by construction. The same-author citations [3] and [49] appear in related-work or motivation contexts (explainability concerns; prior trajectory-supervised methods) and are not load-bearing for any derivation. The paper's own Table 2 does raise a correctness concern about attributing grounding gains to the IDM/FWM module, but that is an underdetermination/mechanism-interpretation issue, not a circular reduction: no equation makes the grounding metric equal to an optimized loss or to a self-citation. Thus the derivation chain is self-contained with respect to circularity; independent support comes from external benchmarks, fixed evaluation masks, and standard held-out test protocols.
Axiom & Free-Parameter Ledger
free parameters (5)
- latent action dimension d_a =
128
- loss weights λ_pred, λ_s, λ_p =
1.0, 0.1, 0.3
- SIGReg projection counts M =
1024 (state/action), 256 (patch)
- per-component learning rates =
3e-5 to 1e-4 depending on encoder (Table 3)
- patch subset size |K_t| =
32 patches per frame
axioms (5)
- domain assumption A compact latent-action bottleneck forces the encoder to represent action-induced transition change and suppresses task-irrelevant nuisance variation.
- domain assumption Mean-pooled patch states retain enough information for the IDM/FWM objective to steer patch-level features toward the interaction region.
- domain assumption SIGReg Gaussianization prevents representational collapse without destroying task-relevant feature structure.
- domain assumption SASVI-propagated CholecSeg8k masks are a valid ground truth for the instrument-tissue interaction region in evaluation.
- domain assumption Eight-frame clips at 1 fps capture the action-relevant transitions between consecutive states.
invented entities (1)
-
latent action a_t (d_a=128)
no independent evidence
read the original abstract
Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language models (VLMs) and vision encoders offer an alternative to conventional interaction classifiers by transferring broad visual and semantic knowledge. However, adapting them to fine-grained surgical interactions remains challenging: (1) freezing the vision encoder depends entirely on pretrained representations that may retain noise and provide weak spatial localization, while (2) full fine-tuning can improve global semantic alignment without ensuring that the encoder learns meaningful features in the correct action region. We address these limitations by introducing LAViFiT, an end-to-end latent-action-guided framework for vision-language fine-tuning. An inverse dynamics model captures the visual changes induced by each action, while a forward world model drives the encoder to represent action-relevant regions. A patch-level SIG Regularizer further prevents local feature collapse without additional supervision, such as bounding boxes or pseudo-labels. Experiments across multiple encoders and datasets improve recognition and image-text alignment, while representation analyses show stronger grounding over the complete instrument-tissue interaction region and more spatially coherent features.
Figures
Reference graph
Works this paper leans on
-
[1]
& Speidel, S
Kati ´c, D., Wekerle, A., Gärtner, F., Kenngott, H., Müller-Stich, B., Dillmann, R. & Speidel, S. A System for Context-Aware Intra-Operative Augmented Reality and Surgical Action Triplet Modeling.Information Processing In Computer-Assisted Interventions (IPCAI). (2014)
2014
-
[2]
& Pinto, L
Cui, Z., Pan, H., Iyer, A., Haldar, S. & Pinto, L. Dynamo: In-domain dynamics pretraining for visuo-motor control.Advances In Neural Information Processing Systems.37pp. 33933-33961 (2024)
2024
-
[3]
& Lin, S
Cheng, J., Zhao, X., Liu, S., Yu, X., Prakash, R., Codd, P., Katz, J. & Lin, S. Surgxbench: Explainable vision- language model benchmark for surgery.Proceedings Of The IEEE/CVF Winter Conference On Applications Of Computer Vision. pp. 8188-8198 (2026)
2026
-
[4]
& Padoy, N
Nwoye, C., Gonzalez, C., Yu, T., Mascagni, P., Mutter, D., Marescaux, J. & Padoy, N. Recognition of Instrument- Tissue Interactions in Endoscopic Videos via Action Triplets.Medical Image Computing And Computer Assisted Intervention (MICCAI). pp. 364-374 (2020)
2020
-
[5]
& Padoy, N
Nwoye, C., Yu, T., Gonzalez, C., Seeliger, B., Mascagni, P., Mutter, D., Marescaux, J. & Padoy, N. Rendezvous: Attention Mechanisms for the Recognition of Surgical Action Triplets in Endoscopic Videos.Medical Image Analysis.78pp. 102433 (2022)
2022
-
[6]
Luo, H., Ji, L., Zhong, M., Chen, Y ., Lei, W., Duan, N. & Li, T. Clip4clip: An empirical study of clip for end to end video clip retrieval.ArXiv Preprint ArXiv:2104.08860. (2021)
Pith/arXiv arXiv 2021
-
[7]
Wu, J., Holm, F., Chen, C., Wang, A., Hu, Y ., Ye, X., Zang, Z., Xu, M., Zhou, L., Liao, H. & Others UniSurg: A Video-Native Foundation Model for Universal Understanding of Surgical Videos.ArXiv Preprint ArXiv:2602.05638. (2026)
Pith/arXiv arXiv 2026
-
[8]
Sharma, S., Mutter, D. & Padoy, N. fine-clip: Enhancing zero-shot fine-grained surgical action recognition with vision-language models.ArXiv Preprint ArXiv:2503.19670. (2025)
Pith/arXiv arXiv 2025
-
[9]
& Torresani, L
Bertasius, G., Wang, H. & Torresani, L. Is space-time attention all you need for video understanding?.Icml.2pp. 4 (2021)
2021
-
[10]
& Qin, J
Chen, Y ., Li, Z., Xu, C., Liu, A., Cui, R., Xu, X., Teoh, J., He, S. & Qin, J. ProstaTD: Bridging Surgical Triplet from Classification to Fully Supervised Detection.The Fourteenth International Conference On Learning Representations. (2026), https://openreview.net/forum?id=0NkXZ98BjJ
2026
-
[11]
Vera, H., Dua, S., Zhang, B., Salz, D., Mullins, R., Panyam, S., Smoot, S., Naim, I., Zou, J., Chen, F. & Others Embeddinggemma: Powerful and lightweight text representations.ArXiv Preprint ArXiv:2509.20354. (2025)
Pith/arXiv arXiv 2025
-
[12]
& Sommen, F
Jaspers, T., Jong, R., Li, Y ., Kusters, C., Bakker, F., Jaarsveld, R., Kuiper, G., Hillegersberg, R., Ruurda, J., Brinkman, W., Pluim, J., With, P., Breeuwer, M., Khalil, Y . & Sommen, F. Scaling up self-supervised learning for improved surgical foundation models.Medical Image Analysis.108pp. 103873 (2026), https://www.sciencedirect.com/science/article/p...
2026
-
[13]
& Others Learning transferable visual models from natural language supervision.International Conference On Machine Learning
Radford, A., Kim, J., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J. & Others Learning transferable visual models from natural language supervision.International Conference On Machine Learning. pp. 8748-8763 (2021)
2021
-
[14]
& Others Dinov2: Learning robust visual features without supervision.ArXiv Preprint ArXiv:2304.07193
Oquab, M., Darcet, T., Moutakanni, T., V o, H., Szafraniec, M., Khalidov, V ., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A. & Others Dinov2: Learning robust visual features without supervision.ArXiv Preprint ArXiv:2304.07193. (2023)
Pith/arXiv arXiv 2023
-
[15]
Mur-Labadia, L., Muckley, M., Bar, A., Assran, M., Sinha, K., Rabbat, M., LeCun, Y ., Ballas, N. & Bardes, A. V-jepa 2.1: Unlocking dense features in video self-supervised learning.ArXiv Preprint ArXiv:2603.14482. (2026)
Pith/arXiv arXiv 2026
-
[17]
& Liang, P
Kumar, A., Raghunathan, A., Jones, R., Ma, T. & Liang, P. Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution.International Conference On Learning Representations. (2022)
2022
-
[18]
& Bach, S
Menghini, C., Delworth, A. & Bach, S. Enhancing clip with clip: Exploring pseudolabeling for limited-label prompt tuning.Advances In Neural Information Processing Systems.36pp. 60984-61007 (2023)
2023
-
[19]
& Bojanowski, P
Darcet, T., Oquab, M., Mairal, J. & Bojanowski, P. Vision transformers need registers.International Conference On Learning Representations.2024pp. 2632-2652 (2024)
2024
-
[20]
& Yuille, A
Wang, F., Mei, J. & Yuille, A. Sclip: Rethinking self-attention for dense vision-language inference.European Conference On Computer Vision. pp. 315-332 (2024)
2024
-
[21]
Bu, Q., Yang, Y ., Cai, J., Gao, S., Ren, G., Yao, M., Luo, P. & Li, H. Univla: Learning to act anywhere with task-centric latent actions.ArXiv Preprint ArXiv:2505.06111. (2025)
Pith/arXiv arXiv 2025
- [22]
-
[23]
& Zhang, W
Lan, M., Chen, C., Ke, Y ., Wang, X., Feng, L. & Zhang, W. Clearclip: Decomposing clip representations for dense vision-language inference.European Conference On Computer Vision. pp. 143-160 (2024)
2024
-
[24]
& Jampani, V
El Banani, M., Raj, A., Maninis, K., Kar, A., Li, Y ., Rubinstein, M., Sun, D., Guibas, L., Johnson, J. & Jampani, V . Probing the 3d awareness of visual foundation models.Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 21795-21806 (2024)
2024
-
[25]
& Ghanem, B
Zohra, F., Zhao, C., Itani, H. & Ghanem, B. b-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment.Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 680-689 (2026)
2026
-
[26]
Cali ´nski, T. & JA, H. A Dendrite Method for Cluster Analysis.Communications In Statistics - Theory And Methods.3pp. 1-27 (1974,1)
1974
-
[27]
& Eom, C
Choi, H., Jang, Y . & Eom, C. Goal: Global-local object alignment learning.Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 4070-4079 (2025)
2025
-
[28]
& Xiong, H
Ding, S., Li, M., Yang, T., Qian, R., Xu, H., Chen, Q., Wang, J. & Xiong, H. Motion-aware contrastive video representation learning via foreground-background merging.Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 9716-9726 (2022)
2022
-
[29]
Hong, W., Kao, C., Kuo, Y ., Wang, J., Chang, W. & Shih, C. Cholecseg8k: a semantic segmentation dataset for laparoscopic cholecystectomy based on cholec80.ArXiv Preprint ArXiv:2012.12453. (2020)
Pith/arXiv arXiv 2012
-
[30]
& Mukhopadhyay, A
Sivakumar, S., Frisch, Y ., Ranem, A. & Mukhopadhyay, A. SASVi: segment any surgical video.International Journal Of Computer Assisted Radiology And Surgery.20, 1409-1419 (2025)
2025
-
[31]
& Khan, F
Khattak, M., Rasheed, H., Maaz, M., Khan, S. & Khan, F. Maple: Multi-modal prompt learning.Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 19113-19122 (2023)
2023
-
[32]
& Yuan, J
Xi, N., Meng, J. & Yuan, J. Chain-of-Look Prompting for Verb-Centric Surgical Triplet Recognition in Endoscopic Videos.Proceedings Of The 31st ACM International Conference On Multimedia. pp. 5007-5016 (2023)
2023
-
[33]
Balestriero, R. & LeCun, Y . Lejepa: Provable and scalable self-supervised learning without the heuristics.ArXiv Preprint ArXiv:2511.08544. (2025)
Pith/arXiv arXiv 2025
-
[34]
& Tang, J
Li, P., Shu, X., Feng, Y ., Feng, Y ., Zuo, W. & Tang, J. Surgical Video Workflow Analysis via Visual-Language Learning.Npj Health Systems.2, 5 (2025)
2025
-
[35]
Jamal, M. & Mohareri, O. SurgMAE: Masked autoencoders for long surgical video analysis.ArXiv Preprint ArXiv:2305.11451. (2023) 14 LA VIFiT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction RecognitionA PREPRINT
Pith/arXiv arXiv 2023
-
[36]
Bati ´c, D., Holm, F., Özsoy, E., Czempiel, T. & Navab, N. Whether and When does Endoscopy Domain Pretraining Make Sense?.ArXiv Preprint ArXiv:2303.17636. (2023)
Pith/arXiv arXiv 2023
-
[38]
& Garcia-Peraza-Herrera, L
Che, C., Wang, C., Vercauteren, T., Tsoka, S. & Garcia-Peraza-Herrera, L. Lemon: A large endoscopic monocular dataset and foundation model for perception in surgical settings.Proceedings Of The IEEE/CVF Conference On Computer Vision And Pattern Recognition. pp. 42659-42669 (2026)
2026
-
[39]
& Padoy, N
Yuan, K., Srivastav, V ., Yu, T., Lavanchy, J., Marescaux, J., Mascagni, P., Navab, N. & Padoy, N. Learning multi- modal representations by watching hundreds of surgical video lectures.Medical Image Analysis. pp. 103644 (2025)
2025
-
[40]
& Padoy, N
Yuan, K., Srivastav, V ., Navab, N. & Padoy, N. HecVL: hierarchical video-language pretraining for zero-shot surgical phase recognition.International Conference On Medical Image Computing And Computer-Assisted Intervention. pp. 306-316 (2024)
2024
-
[41]
& Padoy, N
Yuan, K., Srivastav, V ., Navab, N. & Padoy, N. Procedure-aware surgical video-language pretraining with hierarchi- cal knowledge augmentation.Advances In Neural Information Processing Systems.37pp. 122952-122983 (2024)
2024
-
[42]
The robot will see you now: Foundation models are the path forward for autonomous robotic surgery
Yip, M. The robot will see you now: Foundation models are the path forward for autonomous robotic surgery. Science Robotics.10, eadt0684 (2025)
2025
-
[44]
& Padoy, N
Nwoye, C., Yu, T., Gonzalez, C., Seeliger, B., Mascagni, P., Mutter, D., Marescaux, J. & Padoy, N. Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos.Medical Image Analysis. 78pp. 102433 (2022)
2022
-
[45]
& Wang, Z
Gui, S. & Wang, Z. Tail-enhanced representation learning for surgical triplet recognition.International Conference On Medical Image Computing And Computer-Assisted Intervention. pp. 689-699 (2024)
2024
-
[46]
Liu, D., Hu, A., Shah, M. & Xu, C. Surgical triplet recognition via diffusion model.ArXiv Preprint ArXiv:2406.13210. (2024)
Pith/arXiv arXiv 2024
-
[47]
What do latent action models actually learn?
C. Zhang, T. Pearce, P. Zhang, K. Wang, X. Chen, W. Shen, L. Zhao, and J. Bian, “What do latent action models actually learn?”Advances in Neural Information Processing Systems, vol. 38, pp. 146676-146697, 2026
2026
-
[48]
Nwoye, C. & Padoy, N. Data Splits and Metrics for Method Benchmarking on Surgical Action Triplet Datasets. ArXiv Preprint ArXiv:2204.05235. (2022)
Pith/arXiv arXiv 2022
- [49]
-
[50]
& Xie, S
Peebles, W. & Xie, S. Scalable diffusion models with transformers.Proceedings Of The IEEE/CVF International Conference On Computer Vision. pp. 4195-4205 (2023)
2023
-
[51]
& Heng, P
Pei, J., Zhang, J., Qin, G., Wang, K., Jin, Y . & Heng, P. Instrument-Tissue-Guided Surgical Action Triplet Detection via Textual-Temporal Trail Exploration.IEEE Transactions On Medical Imaging.44, 5278-5289 (2025)
2025
-
[52]
& Others CholecTriplet2021: A Benchmark Challenge for Surgical Action Triplet Recognition
Nwoye, C. & Others CholecTriplet2021: A Benchmark Challenge for Surgical Action Triplet Recognition. Medical Image Analysis.86pp. 102803 (2023)
2023
-
[53]
& Padoy, N
Sharma, S., Nwoye, C., Mutter, D. & Padoy, N. Surgical action triplet detection by mixed supervised learning of instrument-tissue interactions.International Conference On Medical Image Computing And Computer-Assisted Intervention. pp. 505-514 (2023)
2023
-
[54]
& Others CholecTriplet2022: Show Me a Tool and Tell Me the Triplet — An Endoscopic Vision Challenge for Surgical Action Triplet Detection.Medical Image Analysis
Nwoye, C. & Others CholecTriplet2022: Show Me a Tool and Tell Me the Triplet — An Endoscopic Vision Challenge for Surgical Action Triplet Detection.Medical Image Analysis. (2023)
2023
-
[55]
Delta-JEPA: Learn- ing Action-Sensitive World Models via Latent Difference Decoding,
Z. Zhang, Y . Wang, Z. Guan, Y . Yang, B. Shi, T. Zong, H. Yi, G. Chao, X. Chen, T. Yang,et al., “Delta-JEPA: Learn- ing Action-Sensitive World Models via Latent Difference Decoding,”arXiv preprint arXiv:2606.31232, 2026
Pith/arXiv arXiv 2026
-
[56]
& Padoy, N
Sharma, S., Nwoye, C., Mutter, D. & Padoy, N. Rendezvous in Time: An Attention-Based Temporal Fusion Approach for Surgical Triplet Recognition.International Journal Of Computer Assisted Radiology And Surgery. (2023)
2023
-
[57]
& Jia, F
Li, Y ., Xia, T., Luo, H., He, B. & Jia, F. MT-FiST: A Multi-Task Fine-Grained Spatial-Temporal Framework for Surgical Action Triplet Recognition.IEEE Journal Of Biomedical And Health Informatics.27, 4983-4994 (2023) 15 LA VIFiT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction RecognitionA PREPRINT
2023
-
[58]
& Cao, Y
Gui, S., Wang, Z., Chen, J., Zhou, X., Zhang, C. & Cao, Y . MT4MTL-KD: A Multi-Teacher Knowledge Distillation Framework for Triplet Recognition.IEEE Transactions On Medical Imaging.43, 1628-1639 (2024)
2024
-
[59]
& Others Capturing Action Triplet Correlations for Accurate Surgical Activity Recognition.Computerized Medical Imaging And Graphics
Li, Y . & Others Capturing Action Triplet Correlations for Accurate Surgical Activity Recognition.Computerized Medical Imaging And Graphics. (2025)
2025
-
[60]
& Jia, F
Li, Y ., Bai, B. & Jia, F. Parameter-Efficient Framework for Surgical Action Triplet Recognition.International Journal Of Computer Assisted Radiology And Surgery.19, 1291-1299 (2024)
2024
-
[61]
& Others ITG-Trip: Integrating Temporal and Visual-Linguistic Cues for Surgical Action Triplet Recognition.ArXiv Preprint
Pei, J. & Others ITG-Trip: Integrating Temporal and Visual-Linguistic Cues for Surgical Action Triplet Recognition.ArXiv Preprint. (2025), verify authors/venue before camera-ready
2025
-
[62]
& Xiong, J
Li, J., Quaranto, B., Xu, C., Mishra, I., Qin, R., Liu, D., Kim, P. & Xiong, J. Recognize any surgical object: unleashing the power of weakly-supervised data.International Conference On Learning Representations.2025 pp. 33811-33823 (2025)
2025
-
[63]
Others SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence.ArXiv Preprint ArXiv:2506.02555. (2025)
Pith/arXiv arXiv 2025
-
[64]
& Padoy, N
Yuan, K., Kattel, M., Lavanchy, J., Navab, N., Srivastav, V . & Padoy, N. Advancing Surgical VQA with Scene Graph Knowledge.International Journal Of Computer Assisted Radiology And Surgery (IPCAI). (2024)
2024
-
[65]
& Others Egocentric Video-Language Pretraining.Advances In Neural Information Processing Systems (NeurIPS)
Lin, K., Wang, A., Soldan, M. & Others Egocentric Video-Language Pretraining.Advances In Neural Information Processing Systems (NeurIPS). (2022)
2022
-
[66]
& Others EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone.International Conference On Computer Vision (ICCV)
Pramanick, S., Song, Y ., Nag, S. & Others EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone.International Conference On Computer Vision (ICCV). (2023)
2023
-
[67]
& Others InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.ArXiv Preprint
Wang, Y . & Others InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.ArXiv Preprint. (2024)
2024
-
[68]
& Others Ego4D: Around the World in 3,000 Hours of Egocentric Video.IEEE/CVF Conference On Computer Vision And Pattern Recognition (CVPR)
Grauman, K. & Others Ego4D: Around the World in 3,000 Hours of Egocentric Video.IEEE/CVF Conference On Computer Vision And Pattern Recognition (CVPR). (2022)
2022
-
[69]
Others EgoHOIBench: Benchmarking Egocentric Hand-Object Interaction Understanding.ArXiv Preprint ArXiv:2405.17719. (2024)
Pith/arXiv arXiv 2024
-
[70]
& Isbell, C
Edwards, A., Sahni, H., Schroecker, Y . & Isbell, C. Imitating Latent Policies from Observation.International Conference On Machine Learning (ICML). (2019)
2019
-
[71]
& Jiang, M
Schmidt, D. & Jiang, M. Learning to Act without Actions.International Conference On Learning Representations (ICLR). (2024)
2024
-
[72]
& Others Genie: Generative Interactive Environments.International Conference On Machine Learning (ICML)
Bruce, J., Dennis, M., Edwards, A. & Others Genie: Generative Interactive Environments.International Conference On Machine Learning (ICML). (2024)
2024
-
[73]
Nikulin, A. & Others Latent Action Learning Requires Supervision in the Presence of Distractors.ArXiv Preprint ArXiv:2502.00379. (2025)
Pith/arXiv arXiv 2025
-
[74]
Liang, A., Czempin, P., Hong, M., Zhou, Y ., Biyik, E. & Tu, S. Clam: Continuous latent action models for robot learning from unlabeled demonstrations.ArXiv Preprint ArXiv:2505.04999. (2025)
Pith/arXiv arXiv 2025
-
[75]
& Others Latent Action Pretraining from Videos.International Conference On Learning Representations (ICLR)
Ye, S. & Others Latent Action Pretraining from Videos.International Conference On Learning Representations (ICLR). (2024)
2024
-
[76]
A Path Towards Autonomous Machine Intelligence
LeCun, Y . A Path Towards Autonomous Machine Intelligence. (2022), OpenReview
2022
-
[77]
& Others Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
Assran, M. & Others Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. IEEE/CVF Conference On Computer Vision And Pattern Recognition (CVPR). (2023)
2023
-
[78]
& Others V-JEPA: Latent Video Prediction for Visual Representation Learning
Bardes, A., Garrido, Q., Ponce, J. & Others V-JEPA: Latent Video Prediction for Visual Representation Learning. International Conference On Learning Representations (ICLR). (2024)
2024
-
[79]
& Others V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Assran, M. & Others V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. ArXiv Preprint ArXiv:2506.09985. (2025)
Pith/arXiv arXiv 2025
-
[80]
Maes, L., Le Lidec, Q., Scieur, D., LeCun, Y . & Balestriero, R. LeWorldModel: Stable End-to-End Joint- Embedding Predictive Architecture from Pixels.ArXiv Preprint ArXiv:2603.19312. (2026)
Pith/arXiv arXiv 2026
-
[81]
& Others Deep Delta Learning.ArXiv Preprint ArXiv:2601.00417
Zhang, Y . & Others Deep Delta Learning.ArXiv Preprint ArXiv:2601.00417. (2026)
Pith/arXiv arXiv 2026
-
[82]
Contrastive localized language-image pre-training,
H.-Y . Chen, Z. Lai, H. Zhang, X. Wang, M. Eichner, K. You, M. Cao, B. Zhang, Y . Yang, and Z. Gan, “Contrastive localized language-image pre-training,”arXiv preprint arXiv:2410.02746, 2024
Pith/arXiv arXiv 2024
-
[83]
Multi-grained vision-language pre-training: Aligning texts with visual concepts,
Y . Zeng, X. Zhang, and H. Li, “Multi-grained vision-language pre-training: Aligning texts with visual concepts,” arXiv preprint arXiv:2111.08276, 2021. 16 LA VIFiT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction RecognitionA PREPRINT A Appendix A.1 Implementation details. All three encoders share the same latent-action components and di...
Pith/arXiv arXiv 2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.