Pith. sign in

REVIEW 5 major objections 5 minor 112 references

Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance

T0 review · 5 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A self-supervised compression objective trained on 0.78 million unlabeled surgical frames produces an encoder that fine-tunes to improved results across surgical phase recognition, action triplet detection, segmentation, and polyp…

desk verdict Potentially useful large-scale surgical pretraining study undermined by a disconnect between its theoretical framing and its implemented MAE-style loss. read the letter →

arxiv 2506.01980 v1 pith:LRVYWDCU submitted 2025-05-16 eess.IV cs.AIcs.CV

classification eess.IVcs.AIcs.CV
keywords surgicalfoundationmodelself-supervisedlearningKolmogorovcomplexityentropymaximizationmaskedautoencoderphaserecognitionactiontripletsfew-shot
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes Compress-to-Explore (C2E), a self-supervised pretraining method that trains a vision transformer encoder on 0.78 million unlabeled images drawn from 2122 minimally invasive surgeries. The idea is to make the encoder learn representations by compressing images through stages that reduce Kolmogorov complexity, while a decoder built on entropy-maximizing energy functions reconstructs clinically relevant detail. The paper claims that the resulting encoder, fine-tuned on downstream tasks, outperforms or matches existing surgical vision baselines in phase recognition, instrument-tissue action triplets, segmentation, and polyp diagnosis, and that it generalizes with very few labeled videos on new procedure types. A sympathetic reader would take the central claim to be that an unlabeled-data compression objective, rather than task-specific labels, can carry a general surgical visual foundation model.

What carries the argument

The load-bearing object is the compression encoder: each layer removes a constant dimension and applies a shrinkage step derived from the objective $\max_\theta\,-H(Z)+H(I)$, which the paper connects to Kolmogorov complexity through the asymptotic identity $E[\frac{1}{n}K(Z^n)]\to H(Z)$. The gradient update uses an SVD-based form $Z_{i+\frac12}=Z_{i+1}-\beta(S_{i+1}V_{i+1}^{-1}D_{i+1})$ with a residual-like projection bypass, so each layer is meant to solve a concave entropy-maximization subproblem. The decoder inverts this with a temperature-dependent Langevin sampler, where the temperature is estimated from a conditional slice of the hidden state. In the actual pretraining, however, the stated loss is mean absolute error between reconstructed and original patches, run as a standard masked autoencoder pipeline.

What would settle it

Train the same architecture with the same data, masking, and training schedule but replace the decoder and losses with a plain MAE decoder and L2 or MAE reconstruction; if phase, action-triplet, segmentation, and polyp metrics are statistically indistinguishable from C2E's, the compression and entropy machinery is not carrying the reported gains. Conversely, the paper's claim would be supported by showing that removing the entropy-maximizing Langevin decoder or the dimension-reduction projections measurably degrades downstream accuracy.

Watch

Extended reading notes

Core claim

On the paper's own terms, C2E establishes that maximizing the entropy difference between input and a dimension-reduced hidden state, equivalently minimizing Kolmogorov complexity under the identification of complexity with entropy, yields an encoder whose latent representations are compact and disentangled by surgery type, phase, and instrument. The reconstruction side uses a decoder derived from the closed-form entropy-maximizing distribution $p(z)\propto e^{-E(z)/E[E(z)]}$ sampled via Langevin dynamics, intended to recover perceptually fine details such as tool textures and organ boundaries. The paper reports, with the encoder used as a backbone, 92.5% accuracy on Cholec80 phase recognition, 41.8 mAP on action triplet recognition, 0.74 IoU on CholecSeg8k segmentation, and 94.0% accuracy on polyp diagnosis, and substantially higher few-shot phase accuracy than a vision transformer pretrained on endoscopic images.

Load-bearing premise

The pretraining that produced the reported results is the same network described by the compression and entropy-maximization objectives, rather than a standard masked-autoencoder reconstruction loss with the theoretical machinery playing no functional role.

Editorial extensions

If this is right

  • If C2E's compression objective is what drives its representations, unlabeled video libraries become the main resource for building surgical foundation models rather than labeled benchmarks.
  • Fine-tuning C2E on as few as two or eight videos yields phase recognition gains of roughly 8 to 26 points over a ViT baseline on colorectal procedure types, suggesting new procedures can be modeled with minimal annotation.
  • Because one pretrained encoder helps classification, triplet detection, segmentation, and diagnosis, downstream systems can share a single backbone instead of task-specific pretraining.
  • The claimed disentanglement of surgery type, phase, and instrument in the latent space could enable interpretable monitoring, with attention heads tracking local details and their aggregation composing global context.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reported comparisons mostly keep the downstream decoder fixed while swapping the backbone, so the incremental gains may come from the larger and more diverse pretraining corpus rather than from the compression theory; a controlled experiment pretraining plain MAE on the identical corpus would separate the two.
  • If the effective objective is standard MAE, C2E can be read as evidence that scaling up diverse surgical videos with a masked-autoencoder objective is sufficient, with the Kolmogorov-complexity framing serving as an interpretation rather than a mechanism.
  • A testable extension is to evaluate C2E on other fine-grained medical video domains, such as endoscopy or retinal surgery, where local texture is decisive; the paper's own logic predicts the largest gains where global and local features must be balanced.
  • The temperature-conditional decoder suggests a route to per-image uncertainty estimates in downstream predictions, though the paper does not attempt this.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes Compress-to-Explore (C2E), a self-supervised pretraining framework for surgical video understanding. The method claims a compression encoder that maximizes a Kolmogorov-complexity difference objective and an exploration decoder based on entropy-maximizing Langevin sampling, pretrained on 0.78M unlabeled surgical images from 2122 procedures. The pretrained encoder is then fine-tuned on phase recognition, action-triplet classification, segmentation, polyp diagnosis, and few-shot phase recognition. The central claims are that C2E produces compact, disentangled representations and improves downstream accuracy relative to ViT/MAE-style baselines, thereby serving as a surgical visual foundation model.

Significance. If the empirical results were reproducible and the mechanism were actually implemented, the scale of pretraining (0.78M frames, 2122 surgeries) and the breadth of downstream tasks would be a useful contribution to surgical computer vision. The paper also includes few-shot evaluations and representation visualizations, which are appropriate for a foundation-model claim. However, the theoretical core supporting the C2E mechanism is not sound, the implemented loss appears to be standard MAE reconstruction, and the empirical comparisons do not consistently show the claimed improvements. The paper does not release code or the private MGH dataset, so the reported gains cannot currently be attributed to the proposed mechanism rather than to data scale or an architectural variant of MAE. The significance as stated, namely a new compression- and entropy-based self-supervised principle, is therefore not established.

major comments (5)
  1. [Section II-D] The paper explicitly states that 'the loss function minimizes the mean average error of the predicted images and the original images, making the overall training run as a standard MAE [17] pipeline.' None of the terms in Eq. 3, such as -1/2 ln|Σ_Z| and the dimension-change term, appears in this loss, and Eq. 10's energy, temperature kT(h), and Langevin noise are likewise absent from any weight update. Consequently, the trained encoder is not the one whose compression and entropy-maximization properties are proved in Section V. This is a load-bearing inconsistency: the central claim that C2E 'leverages Kolmogorov complexity' and 'uses entropy-maximizing decoders' is disconnected from the implemented training objective.
  2. [Section V, Theorems 1 and 3] Theorem 1 states that the entropy H(X) is the Kolmogorov complexity K(X), but the proof establishes only the asymptotic relation H(X) ≤ E[(1/n)K(X)] ≤ H(X) + ... for finite alphabets, i.e., that the expected normalized Kolmogorov complexity converges to H(X). This does not imply equality of H(X) and K(X) for individual embeddings. In addition, Theorem 3 refers to H(Z) as a 'higher bound' of K(X), while Eq. 13 places H(X) as a lower bound of the expected complexity; the inequality is used in the wrong direction. These errors undermine the theoretical foundation of the compression encoder.
  3. [Section II-B, Eq. (4), and Section V, Theorems 2 and 5] The algebraic derivation of Eq. 4 is not valid as written. From Z = SVD, the simplification Z(Z^T Z)^{-1} = S V^{-1} D requires orthogonality and invertibility assumptions that are not stated and are not satisfied by the rectangular matrix Z ∈ R^{N,C}; the computation also drops D^{-1} and misplaces the transpose structure. Separately, Theorems 2 and 5 invoke the concavity of ln|A| on positive definite Hermitian matrices and assert that 'Z is a positive definite Hermitian matrix', but the paper's own notation defines Z as an N×C matrix of hidden states. The matrix-analysis theorems therefore do not apply as stated, leaving Eqs. 3-5 without a valid justification.
  4. [Section II-C, Eqs. (6)-(8)] The maximum-entropy solution is misidentified in Eqs. 6-8. Under a constraint on E[E(z)], the Lagrange multiplier λ is the inverse temperature, and the solution is p(z) ∝ exp(-λ E(z)); it is not p(z) ∝ exp(-E(z)/E[E(z)]). Writing E[E(z)] = kT(h) in Eq. 8 conflates the expectation of the energy with the temperature parameter, yet this identification is used directly in the Langevin sampling of Eq. 10. The exploration decoder's theoretical grounding is therefore also unsupported.
  5. [Section III, Tables II-VI] The reported results do not consistently support the claim of improved performance. In Table II, C2E (92.5±6.9) is statistically indistinguishable from SurgFormer (92.4±6.4), and its precision is lower than several baselines. In Table III, C2E's mAPit (42.3±0.9) is below MT4MTL's (43.1±2.0), even though the text claims better average precision across the action-triplet metrics. More importantly, no ablation controls for the 531,565 private MGH images or for the specific architecture changes against a standard MAE trained on the same data; the only ViT comparison in Table VI uses a previously published model rather than a same-data, same-compute MAE baseline. Without such controls, the reported gains cannot be attributed to the C2E mechanism.
minor comments (5)
  1. [Throughout] There are numerous typos, including 'Massachusset' in the author affiliation, 'adoped' in Section III-D, 'few-short learning' in the contributions list and Section III-F, and 'Precession' instead of 'Precision' in the Table II header.
  2. [Table VII] The entry '58.8 ±11.' is incomplete; the standard deviation is missing its decimal digits.
  3. [Section III-F and Figure 3] The naming is inconsistent: 'VIT' is used in Figure 3 and Table VI while 'ViT' is used elsewhere; please unify.
  4. [Section II, notation] The symbol E is used both for the energy function and for the expectation operator, e.g., E[E(z)] in Eqs. 6-8 and Eq. 10; this is confusing and should be disambiguated with different symbols.
  5. [Section II-A] The claim that Z0 is a 'disentangled and sparse normal distribution' is not supported by any quantitative evaluation in the paper; the t-SNE visualizations in Section III-F are not a substitute for a measured sparsity or disentanglement metric.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: downstream results are external-benchmark evaluations with no labels in pretraining; the theory–implementation gap is a validity issue, not a circular reduction.

full rationale

The derivation chain is not circular. Pretraining in Section II-D is described as a standard MAE reconstruction pipeline over 0.78M unlabeled frames, and all reported downstream numbers (Cholec80 phase recognition, CholecT45 action triplets, CholecSeg8k segmentation, PolypDiag diagnosis, and few-shot HeiCo evaluations) are measured against external public benchmarks and external baseline methods, with no downstream label used during pretraining and with claimed removal of overlap images to avoid leakage. The compression and entropy-maximization equations (Eq. 3 and Eq. 10) are presented as motivation for the encoder/decoder architecture, and even if the paper's own statement that training 'runs as a standard MAE pipeline' means those information-theoretic objectives are not separately optimized, that is an implementation/validity gap rather than a circular reduction of the predicted result to a fitted input. The mathematical flaws in Theorem 1 (identifying expected Kolmogorov complexity with entropy despite only an asymptotic inequality) and Theorem 3 (calling H(Z) a 'higher bound' when Eq. 13 places it on the lower side) are correctness concerns, not circularity. Self-citations to [48], [61], [62], [66], and [111] appear only as related work or illustrative context and are not load-bearing for the empirical claims. No parameter fitted to a subset of downstream data is renamed as a prediction, and no uniqueness or forced-choice conclusion depends on an author-cited prior theorem. The paper is therefore best described as self-contained against external benchmarks, with the principal risks being theoretical overclaim and disconnect between stated mechanism and implemented loss, not circularity.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The theory pulls in several assumptions: asymptotic equivalence of Kolmogorov complexity and entropy, Gaussian covariance approximation, a dropped dimension-change term, a homogeneity assumption for temperature estimation, and the conflation of the feature matrix Z with a positive definite Hermitian matrix. The conditional ratio beta is an ad hoc free parameter with no reported value. No new physical or ontological entities are introduced; the 'temperature' kT(h) is a re-used statistical mechanics concept.

free parameters (3)
  • conditional ratio beta
    Introduced in Eq. 12 to estimate temperature h = Ec(Z0, beta) by partial elements of the latent; no value or fitting procedure is reported, and the downstream results depend on it via the exploration decoder.
  • Langevin step size epsilon
    Appears in Eq. 9-10 for entropy-maximizing sampling; no schedule or value is reported, and the decoder is never shown to implement these equations.
  • gradient step size beta in Eq. 4
    Serves as the step size for the iterative compression update; no value or schedule is given.
assumptions (5)
  • standard math Kolmogorov complexity of image embeddings asymptotically approaches entropy (Cover and Thomas Thm 14.3.1)
    Invoked in Theorem 1 to equate complexity maximization with entropy maximization; the theorem is used as an asymptotic relation but stated as equality.
  • domain assumption The covariance of embeddings is approximated by Sigma = (1/m) Z Z^T after zero-mean normalization
    Used in Section II-B to express entropy as ln|Sigma|; this assumes the latent distribution is approximately Gaussian and centered.
  • ad hoc to paper Dimension-change term (1/2) Delta M can be dropped because dimension changes are constant
    Used to simplify Eq. 3 to -1/2 ln|Sigma_Z|; the paper later enforces dimension reduction, so the constancy assumption is not justified.
  • ad hoc to paper Assumption 1 in Theorem 6: the distribution of particles z is homogeneous across all subspaces
    Explicitly assumed so that E[e(z)] = E[e(z_subi)] and the temperature can be estimated from partial hidden states; no empirical validation is provided.
  • ad hoc to paper Z is a positive definite Hermitian matrix
    Theorems 2 and 5 rely on concavity of ln|Sigma| over positive-definite Hermitian matrices, but Z is a rectangular feature matrix, not Hermitian; the paper conflates Z with its covariance.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance." pith.science (2026). https://pith.science/paper/LRVYWDCU

@misc{pith2026250601980,
  author       = {Pith},
  title        = {Pith review of: Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LRVYWDCU}},
  note         = {Machine review of arXiv:2506.01980}
}
read the original abstract

Real-time video understanding is critical to guide procedures in minimally invasive surgery (MIS). However, supervised learning approaches require large, annotated datasets that are scarce due to annotation efforts that are prohibitive, e.g., in medical fields. Although self-supervision methods can address such limitations, current self-supervised methods often fail to capture structural and physical information in a form that generalizes across tasks. We propose Compress-to-Explore (C2E), a novel self-supervised framework that leverages Kolmogorov complexity to learn compact, informative representations from surgical videos. C2E uses entropy-maximizing decoders to compress images while preserving clinically relevant details, improving encoder performance without labeled data. Trained on large-scale unlabeled surgical datasets, C2E demonstrates strong generalization across a variety of surgical ML tasks, such as workflow classification, tool-tissue interaction classification, segmentation, and diagnosis tasks, providing improved performance as a surgical visual foundation model. As we further show in the paper, the model's internal compact representation better disentangles features from different structural parts of images. The resulting performance improvements highlight the yet untapped potential of self-supervised learning to enhance surgical AI and improve outcomes in MIS.

Figures

Figures reproduced from arXiv: 2506.01980 by the authors.

Figure 1
Figure 1. C2E General structure. In pretraining, the encoder compresses the raw images patches I into hidden state Z0 using n layers of proposed encoder structure.The output of hidden state of each layer is named as Zi, where i ∈ [0, n]. The decoder reconstruct the images patches and the the output of each layer are as estimate of hidden state Zˆi, where i ∈ [0, n]. The error of the reconstructed images and the raw images are… view at source ↗
Figure 2
Figure 2. The generalization of the pretrained encoder was evaluated across various downstream tasks, such as (a) workflow classification, (b) action classification, (c) segmentation, and (d) diagnosis Tab. IV) using the CholecSeg8k dataset [52], adopting the downstream decoder from the Dense Prediction Transformer (DPT) decoder. A demonstration ( [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Comparison of segmentation using VIT and C2E. a: The original input image. b: The ground truth of the segmentation. Subfigure c and d indicate the difference of ground truth of and prediction Segmentation using VIT(c) and C2E(d). The blue parts are the difference. Our pro￾posed method has smaller difference (small blue areas) and has better semantic segmentation on the area between tool, liver and background TABLE I… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: T-SNE feature visualization of our C2E approach (a) vs MAE (b) for frames from a variety of procedures. Notice how each procedure is mapped to a different cluster in the embedded space, in a more pronounced way, compared to the entangled clusters in MAE. a b [PITH_FUL…
Figure 5
Figure 5. Figure 5: T-SNE feature visualization of our C2E approach (a) vs MAE (b), for phases from Laparoscopic Cholecystectomy procedures. Each phase is mapped to a more seperated clustered group in the embedded space compared to the entangled clusters in MAE. a b [PITH_FULL_IMAGE:figu…
Figure 6
Figure 6. Figure 6: T-SNE feature visualization of our C2E approach (a) vs MAE (b), for frames from Laparoscopic Cholecystectomy procedures. The embeddings of different instruments are disentangled in reduced dimensions. IV. CONCLUSION In conclusion, we leverage the compression method and…
Figure 7
Figure 7. Figure 7: Saliency map comparison of C2E and VIT head. The propsed method C2E has a relative sparsity of the activation maps indicating a more compressed representation. The high attention region of each heads with C2E are complementary and not overlapped indicating that the inf…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

112 extracted references · 67 canonical work pages

  1. [17]

    Masked autoencoders are scalable vision learners,

    K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pp. 16000–16009, 2022

  2. [1]

    Automated 3d liver segmentation from hepatobiliary phase mri for enhanced preoperative planning,

    N. Oh, J.-H. Kim, J. Rhu, W. K. Jeong, G.-s. Choi, J. M. Kim, and J.-W. Joh, “Automated 3d liver segmentation from hepatobiliary phase mri for enhanced preoperative planning,” Scientific Reports , vol. 13, no. 1, p. 17605, 2023

  3. [2]

    Ai-based chest ct semantic segmentation algorithm enables semi-automated lung cancer surgery planning by recognizing anatomical variants of pulmonary vessels,

    X. Chen, H. Xu, Q. Qi, C. Sun, J. Jin, H. Zhao, X. Wang, W. Weng, S. Wang, X. Sui, et al. , “Ai-based chest ct semantic segmentation algorithm enables semi-automated lung cancer surgery planning by recognizing anatomical variants of pulmonary vessels,” Frontiers in Oncology, vol. 12, p. 1021084, 2022

  4. [3]

    Image-guided simulation of heterogeneous tissue deforma- tion for augmented reality during hepatic surgery,

    N. Haouchine, J. Dequidt, I. Peterl ´ık, E. Kerrien, M. Berger, and S. Cotin, “Image-guided simulation of heterogeneous tissue deforma- tion for augmented reality during hepatic surgery,” 2013

  5. [4]

    Computer vision in surgery: from potential to clinical value,

    P. Mascagni, D. Alapatt, L. Sestini, M. S. Altieri, A. Madani, Y . Watan- abe, A. Alseidi, J. A. Redan, S. Alfieri, G. Costamagna, et al. , “Computer vision in surgery: from potential to clinical value,” npj Digital Medicine, vol. 5, no. 1, p. 163, 2022

  6. [5]

    Laparoscopic image-based critical action recognition and anticipation with explainable features,

    J. Zhang, S. Zhou, Y . Wang, S. Shi, C. Wan, H. Zhao, X. Cai, and H. Ding, “Laparoscopic image-based critical action recognition and anticipation with explainable features,” IEEE J. Biomed. Health Inform., vol. 27, pp. 5393–5404, Nov. 2023

  7. [6]

    Current and future applications of artificial intelligence in surgery: implications for clinical practice and research,

    M. X. Morris, D. Fiocco, T. Caneva, P. Yiapanis, and D. P. Orgill, “Current and future applications of artificial intelligence in surgery: implications for clinical practice and research,” Front. Surg., vol. 11, p. 1393898, May 2024

  8. [7]

    Artificial intelligence in surgery: the future is now,

    H. Ashrafiana, “Artificial intelligence in surgery: the future is now,” Eur Surg Res , vol. 65, pp. 22–39, 2024

Show all 112 references
  1. [8]

    Artificial in- telligence in improving the outcome of surgical treatment in colorectal cancer,

    M. F. Avram, D. C. Laz ˘ar, M. I. Maris ¸, and S. Olariu, “Artificial in- telligence in improving the outcome of surgical treatment in colorectal cancer,” Front. Oncol., vol. 13, p. 1116761, Jan. 2023

  2. [9]

    Shalev-Shwartz and S

    S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: From theory to algorithms . Cambridge university press, 2014

  3. [10]

    SAGES consensus recommendations on an annotation framework for surgical video,

    O. R. Meireles, G. Rosman, M. S. Altieri, L. Carin, G. Hager, A. Madani, N. Padoy, C. M. Pugh, P. Sylla, T. M. Ward, D. A. Hashimoto, and SAGES Video Annotation for AI Working Groups, “SAGES consensus recommendations on an annotation framework for surgical video,” Surg. Endosc...

  4. [11]

    J. A. Eckhoff, G. Rosman, M. S. Altieri, S. Speidel, D. Stoyanov, M. Anvari, L. Meier-Hein, K. M ¨arz, P. Jannin, C. Pugh, M. Wagner, E. Witkowski, P. Shaw, A. Madani, Y . Ban, T. Ward, F. Filicori, N. Padoy, M. Talamini, and O. R. Meireles, “SAGES consensus recommendations on...

  5. [12]

    What does classifying more than 10,000 image categories tell us?,

    J. Deng, A. C. Berg, K. Li, and L. Fei-Fei, “What does classifying more than 10,000 image categories tell us?,” inComputer Vision–ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, September 5-11, 2010, Proceedings, Part V 11 , pp. 71– 84, Springer, 2010

  6. [13]

    MT4MTL- KD: A multi-teacher knowledge distillation framework for triplet recognition,

    S. Gui, Z. Wang, J. Chen, X. Zhou, C. Zhang, and Y . Cao, “MT4MTL- KD: A multi-teacher knowledge distillation framework for triplet recognition,” IEEE Trans. Med. Imaging, vol. 43, pp. 1628–1639, Apr. 2024

  7. [14]

    Foundation model for endoscopy video analysis via large-scale self-supervised pre-train,

    Z. Wang, C. Liu, S. Zhang, and Q. Dou, “Foundation model for endoscopy video analysis via large-scale self-supervised pre-train,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2023, pp. 101–111, Springer Nature Switzerland, 2023

  8. [15]

    The effect of image resolution on deep learning in radiography,

    C. F. Sabottke and B. M. Spieler, “The effect of image resolution on deep learning in radiography,” Radiol. Artif. Intell., vol. 2, p. e190015, Jan. 2020

  9. [16]

    Ultra-high resolution, multi-scale, context- aware approach for detection of small cancers on mammography,

    K. Rangarajan, A. Gupta, S. Dasgupta, U. Marri, A. K. Gupta, S. Hari, S. Banerjee, and C. Arora, “Ultra-high resolution, multi-scale, context- aware approach for detection of small cancers on mammography,” Sci. Rep., vol. 12, p. 11622, July 2022

  10. [18]

    Exploring the effect of dataset diversity in self-supervised learning for surgical computer vision,

    T. J. Jaspers, R. L. de Jong, Y . Al Khalil, T. Zeelenberg, C. H. Kusters, Y . Li, R. C. van Jaarsveld, F. H. Bakker, J. P. Ruurda, W. M. Brinkman, et al. , “Exploring the effect of dataset diversity in self-supervised learning for surgical computer vision,” in MICCAI Workshop...

  11. [19]

    Endovit: pretraining vision transformers on a large collection of endoscopic images,

    D. Bati ´c, F. Holm, E. ¨Ozsoy, T. Czempiel, and N. Navab, “Endovit: pretraining vision transformers on a large collection of endoscopic images,” International Journal of Computer Assisted Radiology and Surgery, vol. 19, no. 6, pp. 1085–1091, 2024

  12. [20]

    High-resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, and others, “High-resolution image synthesis with latent diffusion models,” Proceedings of the , 2022

  13. [21]

    On the opportunities and risks of foundation models,

    R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al. , “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021

  14. [22]

    Exploring simple siamese representation learn- ing,

    X. Chen and K. He, “Exploring simple siamese representation learn- ing,” arXiv [cs.CV], pp. 15750–15758, Nov. 2020

  15. [23]

    Self-supervised learning for endoscopic video analysis,

    R. Hirsch, M. Caron, R. Cohen, A. Livne, R. Shapiro, T. Golany, R. Goldenberg, D. Freedman, and E. Rivlin, “Self-supervised learning for endoscopic video analysis,” arXiv [cs.CV], Aug. 2023

  16. [24]

    Pseudo- label guided cross-video pixel contrast for robotic surgical scene seg- mentation with limited annotations,

    Y . Yu, Z. Zhao, Y . Jin, G. Chen, Q. Dou, and P.-A. Heng, “Pseudo- label guided cross-video pixel contrast for robotic surgical scene seg- mentation with limited annotations,” in 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 10857– 1086...

  17. [25]

    Less is more: Sur- gical phase recognition with less annotations through self-supervised pre-training of cnn-lstm networks,

    G. Yengera, D. Mutter, J. Marescaux, and N. Padoy, “Less is more: Sur- gical phase recognition with less annotations through self-supervised pre-training of cnn-lstm networks,” arXiv preprint arXiv:1805.08569 , 2018

  18. [26]

    “train one, classify one, teach one

    D. Neimark, O. Bar, M. Zohar, G. D. Hager, and D. Asselmann, ““train one, classify one, teach one”-cross-surgery transfer learning for surgical step recognition,” in Medical Imaging with Deep Learning , pp. 532– 544, PMLR, 2021

  19. [27]

    Dissecting self- supervised learning methods for surgical computer vision,

    S. Ramesh, V . Srivastav, D. Alapatt, T. Yu, A. Murali, L. Sestini, C. I. Nwoye, I. Hamoud, S. Sharma, A. Fleurentin, et al., “Dissecting self- supervised learning methods for surgical computer vision,” Medical Image Analysis, vol. 88, p. 102844, 2023

  20. [28]

    Self-supervised learning from images with a joint-embedding predictive architecture,

    M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, Y . LeCun, and N. Ballas, “Self-supervised learning from images with a joint-embedding predictive architecture,” arXiv [cs.CV], pp. 15619– 15629, Jan. 2023

  21. [29]

    VICRegL: Self-supervised learning of local visual features,

    A. Bardes, J. Ponce, and Y . LeCun, “VICRegL: Self-supervised learning of local visual features,” in Advances in Neural Information Processing Systems (S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, eds.), vol. 35, pp. 8799–8810, Curran Associates, Inc., 2022

  22. [30]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” arXiv [cs.LG], pp. 6840–6851, June 2020

  23. [31]

    Diffusion models for medical image analysis: A comprehensive survey,

    A. Kazerouni, E. K. Aghdam, M. Heidari, R. Azad, M. Fayyaz, I. Hacihaliloglu, and D. Merhof, “Diffusion models for medical image analysis: A comprehensive survey,” arXiv [eess.IV], Nov. 2022

  24. [32]

    A fast karhunen-loeve transform for a class of random processes,

    A. K. Jain, “A fast karhunen-loeve transform for a class of random processes,” IEEE Transactions on Communications , vol. 24, no. 9, pp. 1023–1029, 1976

  25. [33]

    Segmentation of multivariate mixed data via lossy data coding and compression,

    Y . Ma, H. Derksen, W. Hong, and J. Wright, “Segmentation of multivariate mixed data via lossy data coding and compression,” IEEE transactions on pattern analysis and machine intelligence , vol. 29, no. 9, pp. 1546–1562, 2007

  26. [34]

    Provable bounds for learning some deep representations,

    S. Arora, A. Bhaskara, R. Ge, and T. Ma, “Provable bounds for learning some deep representations,” in International conference on machine learning, pp. 584–592, PMLR, 2014

  27. [35]

    White-box transformers via sparse rate reduction,

    Y . Yu, S. Buchanan, D. Pai, T. Chu, Z. Wu, S. Tong, B. Haeffele, and Y . Ma, “White-box transformers via sparse rate reduction,”Advances in Neural Information Processing Systems, vol. 36, pp. 9422–9457, 2023

  28. [36]

    Deep learning and the information bottleneck principle,

    N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in 2015 ieee information theory workshop (itw) , pp. 1–5, Ieee, 2015. 10

  29. [37]

    Direct validation of the information bottleneck principle for deep nets,

    A. Elad, D. Haviv, Y . Blau, and T. Michaeli, “Direct validation of the information bottleneck principle for deep nets,” 2019

  30. [38]

    The information bottleneck problem and its applications in machine learning,

    Z. Goldfeld and Y . Polyanskiy, “The information bottleneck problem and its applications in machine learning,” 2020

  31. [39]

    On neural networks fitting, compression, and generalization behavior via information-bottleneck- like approaches,

    Z. Lyu, G. Aminian, and M. Rodrigues, “On neural networks fitting, compression, and generalization behavior via information-bottleneck- like approaches,” 2023

  32. [40]

    Neural manifold clustering and embedding,

    Z. Li, Y . Chen, Y . LeCun, and F. T. Sommer, “Neural manifold clustering and embedding,” arXiv preprint arXiv:2201.10000 , 2022

  33. [41]

    Self-supervised learning via maximum entropy coding,

    X. Liu, Z. Wang, Y .-L. Li, and S. Wang, “Self-supervised learning via maximum entropy coding,” Advances in neural information processing systems, vol. 35, pp. 34091–34105, 2022

  34. [42]

    Segmentation of multivariate mixed data via lossy data coding and compression,

    Y . Ma, H. Derksen, and W. Hong, “Segmentation of multivariate mixed data via lossy data coding and compression,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 29, pp. 1546–1562, Sept. 2007

  35. [43]

    Fast: Efficient action tokenization for vision-language-action models,

    K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine, “Fast: Efficient action tokenization for vision-language-action models,” arXiv preprint arXiv:2501.09747 , 2025

  36. [44]

    Human-like systematic generalization through a meta-learning neural network,

    B. M. Lake and M. Baroni, “Human-like systematic generalization through a meta-learning neural network,” Nature, vol. 623, no. 7985, pp. 115–121, 2023

  37. [45]

    Unsuper- vised learning of compositional energy concepts,

    Y . Du, S. Li, Y . Sharma, J. Tenenbaum, and I. Mordatch, “Unsuper- vised learning of compositional energy concepts,” Advances in Neural Information Processing Systems , vol. 34, pp. 15608–15620, 2021

  38. [46]

    Energy-based models are zero-shot planners for compositional scene rearrangement,

    N. Gkanatsios, A. Jain, Z. Xian, Y . Zhang, C. Atkeson, and K. Fragki- adaki, “Energy-based models are zero-shot planners for compositional scene rearrangement,” arXiv preprint arXiv:2304.14391 , 2023

  39. [47]

    Compositional visual generation with energy based models,

    Y . Du, S. Li, and I. Mordatch, “Compositional visual generation with energy based models,” Advances in Neural Information Processing Systems, vol. 33, pp. 6637–6647, 2020

  40. [48]

    Artificial intelligence in surgery: Promises and perils,

    D. A. Hashimoto, G. Rosman, D. Rus, and O. R. Meireles, “Artificial intelligence in surgery: Promises and perils,” Ann. Surg. , vol. 268, pp. 70–76, July 2018

  41. [49]

    Superpixel-based structure classification for laparoscopic surgery,

    S. Bodenstedt, J. G ¨ortler, M. Wagner, H. Kenngott, B. M ¨uller-Stich, R. Dillmann, and S. Speidel, “Superpixel-based structure classification for laparoscopic surgery,” 2016

  42. [50]

    ToolNet: Holistically-nested real-time segmentation of robotic surgical tools,

    L. C. Garc ´ıa-Peraza-Herrera, W. Li, L. Fidon, C. Gruijthuijsen, A. Devreker, G. Attilakos, J. Deprest, E. V . Poorten, D. Stoyanov, T. Vercauteren, and S. Ourselin, “ToolNet: Holistically-nested real-time segmentation of robotic surgical tools,” in2017 IEEE/RSJ International...

  43. [51]

    2018 robotic scene segmentation challenge,

    M. Allan, S. Kondo, S. Bodenstedt, S. Leger, R. Kadkhodamoham- madi, I. Luengo, F. Fuentes-Hurtado, E. Flouty, A. K. Mohammed, M. Pedersen, A. Kori, A. Varghese, G. Krishnamurthi, D. Rauber, R. Mendel, C. Palm, S. Bano, G. Saibro, C. Shih, H. Chiang, J. Zhuang, J. Yang, V . Ig...

  44. [52]

    Cholecseg8k: a semantic segmentation dataset for laparoscopic cholecystectomy based on cholec80,

    W.-Y . Hong, C.-L. Kao, Y .-H. Kuo, J.-R. Wang, W.-L. Chang, and C.-S. Shih, “Cholecseg8k: a semantic segmentation dataset for laparoscopic cholecystectomy based on cholec80,” arXiv preprint arXiv:2012.12453, 2020

  45. [53]

    Detection and localization of robotic tools in robot-assisted surgery videos using deep neural net- works for region proposal and detection,

    D. Sarikaya, J. J. Corso, and K. A. Guru, “Detection and localization of robotic tools in robot-assisted surgery videos using deep neural net- works for region proposal and detection,” IEEE Trans. Med. Imaging , vol. 36, pp. 1542–1549, July 2017

  46. [54]

    Tool detection and operative skill assessment in surgical videos using region-based convolutional neural networks,

    A. Jin, S. Yeung, J. Jopling, J. Krause, D. Azagury, A. Milstein, and L. Fei-Fei, “Tool detection and operative skill assessment in surgical videos using region-based convolutional neural networks,” in 2018 IEEE Winter Conference on Applications of Computer Vision (WACV) , pp....

  47. [55]

    LapFormer: surgical tool detection in laparoscopic surgical video using transformer architecture,

    S. Kondo, “LapFormer: surgical tool detection in laparoscopic surgical video using transformer architecture,” Computer Methods in Biome- chanics and Biomedical Engineering: Imaging & Visualization , vol. 9, pp. 302–307, May 2021

  48. [56]

    Vision-based and marker-less surgical tool detection and tracking: a review of the literature,

    D. Bouget, M. Allan, D. Stoyanov, and P. Jannin, “Vision-based and marker-less surgical tool detection and tracking: a review of the literature,” Medical image analysis , vol. 35, pp. 633–654, 2017

  49. [57]

    Tracking of in- struments in minimally invasive surgery for surgical skill analysis,

    S. Speidel, M. Delles, C. Gutt, and R. Dillmann, “Tracking of in- struments in minimally invasive surgery for surgical skill analysis,” in Medical Imaging and Augmented Reality: Third International Work- shop, Shanghai, China, August 17-18, 2006 Proceedings 3 , pp. 148– 155, S...

  50. [58]

    Tracking-by- detection of surgical instruments in minimally invasive surgery via the convolutional neural network deep learning-based method,

    Z. Zhao, S. V oros, Y . Weng, F. Chang, and R. Li, “Tracking-by- detection of surgical instruments in minimally invasive surgery via the convolutional neural network deep learning-based method,” Computer Assisted Surgery, vol. 22, no. sup1, pp. 26–35, 2017

  51. [59]

    Statistical modeling and recognition of surgical workflow,

    N. Padoy, T. Blum, S.-A. Ahmadi, H. Feussner, M.-O. Berger, and N. Navab, “Statistical modeling and recognition of surgical workflow,” Medical image analysis , vol. 16, no. 3, pp. 632–641, 2012

  52. [60]

    Predicting surgical phases using cnn-narx neural network,

    N. A. Jalal, T. A. Alshirbaji, and K. M ¨oller, “Predicting surgical phases using cnn-narx neural network,” Current Directions in Biomedical Engineering, vol. 5, no. 1, pp. 405–407, 2019

  53. [61]

    Supr-gan: Surgical prediction gan for event anticipation in laparoscopic and robotic surgery,

    Y . Ban, G. Rosman, J. A. Eckhoff, T. M. Ward, D. A. Hashimoto, T. Kondo, H. Iwaki, O. R. Meireles, and D. Rus, “Supr-gan: Surgical prediction gan for event anticipation in laparoscopic and robotic surgery,” IEEE Robotics and Automation Letters , vol. 7, no. 2, pp. 5741–5748, 2022

  54. [62]

    Hypergraph-transformer (hgt) for interactive event prediction in la- paroscopic and robotic surgery,

    L. Yin, Y . Ban, J. Eckhoff, O. Meireles, D. Rus, and G. Rosman, “Hypergraph-transformer (hgt) for interactive event prediction in la- paroscopic and robotic surgery,” arXiv preprint arXiv:2402.01974 , 2024

  55. [63]

    Surgical SAM 2: Real- time segment anything in surgical video by efficient frame pruning,

    H. Liu, E. Zhang, J. Wu, M. Hong, and Y . Jin, “Surgical SAM 2: Real- time segment anything in surgical video by efficient frame pruning,” arXiv preprint arXiv:2408.07931 , 2024

  56. [64]

    T. M. Cover and J. A. Thomas, Elements of Information Theory 2nd Edition (Wiley Series in Telecommunications and Signal Processing) . Wiley-Interscience, 2 ed., July 2006

  57. [65]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , pp. 770–778, 2016

  58. [66]

    Fast regularization of matrix-valued images,

    G. Rosman, Y . Wang, X.-C. Tai, R. Kimmel, and A. M. Bruckstein, “Fast regularization of matrix-valued images,” in Efficient Algorithms for Global Optimization Methods in Computer Vision: International Dagstuhl Seminar, Dagstuhl Castle, Germany, November 20-25, 2011, Revised S...

  59. [67]

    Information theory and statistical mechanics,

    E. T. Jaynes, “Information theory and statistical mechanics,” Physical review, 1957

  60. [68]

    Imagededup

    T. Jain, C. Lennan, Z. John, and D. Tran, “Imagededup.” https: //github.com/idealo/imagededup, 2019

  61. [69]

    Exploring CLIP for assessing the look and feel of images,

    J. Wang, K. C. Chan, and C. C. Loy, “Exploring CLIP for assessing the look and feel of images,” in AAAI, pp. 2555–2563, 2023

  62. [70]

    EndoNet: A deep architecture for recognition tasks on laparoscopic videos,

    A. P. Twinanda, S. Shehata, D. Mutter, J. Marescaux, M. de Mathelin, and N. Padoy, “EndoNet: A deep architecture for recognition tasks on laparoscopic videos,” IEEE Trans. Med. Imaging , vol. 36, pp. 86–97, Jan. 2017

  63. [71]

    hsdb-instrument: Instrument localization database for laparoscopic and robotic surgeries,

    J. Yoon, J. Lee, S. Heo, H. Yu, J. Lim, C. H. Song, S. Hong, S. Hong, B. Park, S. Park, et al. , “hsdb-instrument: Instrument localization database for laparoscopic and robotic surgeries,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th Internat...

  64. [72]

    Heidelberg colorectal data set for surgical data science in the sensor operating room,

    L. Maier-Hein, M. Wagner, T. Ross, A. Reinke, S. Bodenstedt, P. M. Full, H. Hempe, D. Mindroc-Filimon, P. Scholz, T. N. Tran, et al. , “Heidelberg colorectal data set for surgical data science in the sensor operating room,” Scientific data, vol. 8, no. 1, p. 101, 2021

  65. [73]

    The dresden surgical anatomy dataset for abdominal organ segmentation in surgical data science,

    M. Carstens, F. M. Rinner, S. Bodenstedt, A. C. Jenke, J. Weitz, M. Distler, S. Speidel, and F. R. Kolbinger, “The dresden surgical anatomy dataset for abdominal organ segmentation in surgical data science,” Scientific Data, vol. 10, no. 1, pp. 1–8, 2023

  66. [74]

    Lapgyn4: a dataset for 4 automatic content analysis problems in the domain of laparoscopic gynecol- ogy,

    A. Leibetseder, S. Petscharnig, M. J. Primus, S. Kletz, B. M ¨unzer, K. Schoeffmann, and J. Keckstein, “Lapgyn4: a dataset for 4 automatic content analysis problems in the domain of laparoscopic gynecol- ogy,” in Proceedings of the 9th ACM multimedia systems conference , pp. 3...

  67. [75]

    Video retrieval in laparoscopic video recordings with dynamic content descriptors,

    K. Schoeffmann, H. Husslein, S. Kletz, S. Petscharnig, B. Muenzer, and C. Beecks, “Video retrieval in laparoscopic video recordings with dynamic content descriptors,” Multimedia Tools and Applications, vol. 77, pp. 16813–16832, 2018

  68. [76]

    Glenda: gynecologic laparoscopy endometriosis dataset,

    A. Leibetseder, S. Kletz, K. Schoeffmann, S. Keckstein, and J. Keck- stein, “Glenda: gynecologic laparoscopy endometriosis dataset,” in International Conference on Multimedia Modeling , pp. 439–450, Springer, 2019

  69. [77]

    The saras endoscopic surgeon action detection (esad) dataset: Challenges and methods,

    V . S. Bawa, G. Singh, F. KapingA, I. Skarga-Bandurova, E. Oleari, A. Leporini, C. Landolfo, P. Zhao, X. Xiang, G. Luo, et al. , “The saras endoscopic surgeon action detection (esad) dataset: Challenges and methods,” arXiv preprint arXiv:2104.03178 , 2021

  70. [78]

    A. P. Twinanda, Vision-based approaches for surgical activity recogni- tion using laparoscopic and RBGD videos . PhD thesis, University of Strasbourg, Strasbourg, France, Jan. 2017. 11

  71. [79]

    Multi-task recurrent convolutional network with correlation loss for surgical video analysis,

    Y . Jin, H. Li, Q. Dou, H. Chen, J. Qin, C.-W. Fu, and P.-A. Heng, “Multi-task recurrent convolutional network with correlation loss for surgical video analysis,” Med. Image Anal. , vol. 59, p. 101572, Jan. 2020

  72. [80]

    Single- and multi-task architectures for surgical workflow challenge at M2CAI 2016,

    A. P. Twinanda, D. Mutter, J. Marescaux, M. de Mathelin, and N. Padoy, “Single- and multi-task architectures for surgical workflow challenge at M2CAI 2016,” arXiv [cs.CV], Oct. 2016

  73. [81]

    SV-RCNet: Workflow recognition from surgical videos using recurrent convolutional network,

    Y . Jin, Q. Dou, H. Chen, L. Yu, J. Qin, C.-W. Fu, and P.-A. Heng, “SV-RCNet: Workflow recognition from surgical videos using recurrent convolutional network,” IEEE Trans. Med. Imaging, vol. 37, pp. 1114– 1126, May 2018

  74. [82]

    TeCNO: Surgical phase recognition with multi- stage temporal convolutional networks,

    T. Czempiel, M. Paschali, M. Keicher, W. Simson, H. Feussner, S. T. Kim, and N. Navab, “TeCNO: Surgical phase recognition with multi- stage temporal convolutional networks,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2020 , pp. 343–352, Springer Int...

  75. [83]

    Trans-SVNet: Accurate phase recognition from surgical videos via hybrid embedding aggregation transformer,

    X. Gao, Y . Jin, Y . Long, Q. Dou, and P.-A. Heng, “Trans-SVNet: Accurate phase recognition from surgical videos via hybrid embedding aggregation transformer,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2021 , pp. 593–603, Springer Interna- tional P...

  76. [84]

    LoViT: Long video transformer for surgical phase recognition,

    Y . Liu, M. Boels, L. C. Garcia-Peraza-Herrera, T. Vercauteren, P. Das- gupta, A. Granados, and S. Ourselin, “LoViT: Long video transformer for surgical phase recognition,” Med. Image Anal. , vol. 99, p. 103366, Oct. 2024

  77. [85]

    Surgformer: Surgical transformer with hierarchical temporal attention for surgical phase recognition,

    S. Yang, L. Luo, Q. Wang, and H. Chen, “Surgformer: Surgical transformer with hierarchical temporal attention for surgical phase recognition,” in International Conference on Medical Image Computing and Computer-Assisted Intervention , pp. 606–616, Springer, 2024

  78. [86]

    CholecTriplet2021: A benchmark challenge for surgical action triplet recognition,

    C. I. Nwoye, D. Alapatt, T. Yu, A. Vardazaryan, F. Xia, Z. Zhao, T. Xia, F. Jia, Y . Yang, H. Wang, D. Yu, G. Zheng, X. Duan, N. Getty, R. Sanchez-Matilla, M. Robu, L. Zhang, H. Chen, J. Wang, L. Wang, B. Zhang, B. Gerats, S. Raviteja, R. Sathish, R. Tao, S. Kondo, W. Pang, H....

  79. [87]

    Recognition of instrument-tissue interactions in endo- scopic videos via action triplets,

    C. I. Nwoye, C. Gonzalez, T. Yu, P. Mascagni, D. Mutter, J. Marescaux, and N. Padoy, “Recognition of instrument-tissue interactions in endo- scopic videos via action triplets,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2020, pp. 364–374, Springer I...

  80. [88]

    Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos,

    C. I. Nwoye, T. Yu, C. Gonzalez, B. Seeliger, P. Mascagni, D. Mutter, J. Marescaux, and N. Padoy, “Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos,” Med. Image Anal., vol. 78, p. 102433, May 2022

  81. [89]

    Concept graph neural networks for surgical video understanding,

    Y . Ban, J. A. Eckhoff, T. M. Ward, D. A. Hashimoto, O. R. Meireles, D. Rus, and G. Rosman, “Concept graph neural networks for surgical video understanding,” IEEE Trans. Med. Imaging , vol. PP, July 2023

  82. [90]

    Rendezvous in time: an attention-based temporal fusion approach for surgical triplet recognition,

    S. Sharma, C. I. Nwoye, D. Mutter, and N. Padoy, “Rendezvous in time: an attention-based temporal fusion approach for surgical triplet recognition,” Int. J. Comput. Assist. Radiol. Surg. , vol. 18, pp. 1053– 1059, June 2023

  83. [91]

    Parameter-efficient framework for surgical action triplet recognition,

    Y . Li, B. Bai, and F. Jia, “Parameter-efficient framework for surgical action triplet recognition,” International Journal of Computer Assisted Radiology and Surgery , vol. 19, no. 7, pp. 1291–1299, 2024

  84. [92]

    AdaptiveSAM: Towards efficient tuning of SAM for surgical scene segmentation,

    J. N. Paranjape, N. G. Nair, S. Sikder, S. S. Vedula, and V . M. Patel, “AdaptiveSAM: Towards efficient tuning of SAM for surgical scene segmentation,” in Medical Image Understanding and Analysis , Lecture notes in computer science, pp. 187–201, Cham: Springer Nature Switzerland, 2024

  85. [93]

    Medical transformer: Gated axial-attention for medical image segmentation,

    J. M. J. Valanarasu, P. Oza, I. Hacihaliloglu, and V . M. Patel, “Medical transformer: Gated axial-attention for medical image segmentation,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2021, pp. 36–46, Springer International Publishing, 2021

  86. [94]

    TransUNet: Transformers make strong encoders for medical image segmentation,

    J. Chen, Y . Lu, Q. Yu, X. Luo, E. Adeli, Y . Wang, L. Lu, A. L. Yuille, and Y . Zhou, “TransUNet: Transformers make strong encoders for medical image segmentation,” arXiv [cs.CV], Feb. 2021

  87. [95]

    Encoder-decoder with atrous separable convolution for semantic im- age segmentation,

    L.-C. Chen, Y . Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic im- age segmentation,” arXiv [cs.CV], Feb. 2018

  88. [96]

    UNETR: Transformers for 3D medical image segmentation,

    A. Hatamizadeh, Y . Tang, V . Nath, D. Yang, A. Myronenko, B. Land- man, H. R. Roth, and D. Xu, “UNETR: Transformers for 3D medical image segmentation,” in 2022 IEEE/CVF Winter Conference on Appli- cations of Computer Vision (WACV) , pp. 574–584, IEEE, Jan. 2022

  89. [97]

    nnU-net: a self-configuring method for deep learning-based biomedical image segmentation,

    F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, and K. H. Maier- Hein, “nnU-net: a self-configuring method for deep learning-based biomedical image segmentation,” Nature methods , vol. 18, no. 2, pp. 203–211, 2021

  90. [98]

    UNet++: A nested U-net architecture for medical image segmentation,

    Z. Zhou, M. M. R. Siddiquee, N. Tajbakhsh, and J. Liang, “UNet++: A nested U-net architecture for medical image segmentation,” Deep Learn. Med. Image Anal. Multimodal Learn. Clin. Decis. Support , vol. 11045, pp. 3–11, Sept. 2018

  91. [99]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Lecture Notes in Computer Science , Lecture notes in computer science, pp. 234–241, Cham: Springer International Publishing, 2015

  92. [100]

    Customized segment anything model for medical image segmentation,

    K. Zhang and D. Liu, “Customized segment anything model for medical image segmentation,” arXiv [cs.CV], Apr. 2023

  93. [101]

    LoRA: Low-rank adaptation of large language models,

    E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” arXiv [cs.CL], June 2021

  94. [102]

    S-SAM: SVD-based fine-tuning of segment anything model for medical image segmentation,

    J. N. Paranjape, S. Sikder, S. S. Vedula, and V . M. Patel, “S-SAM: SVD-based fine-tuning of segment anything model for medical image segmentation,” arXiv [cs.CV], Aug. 2024

  95. [103]

    Rethinking RGB-D fusion for semantic segmentation in surgical datasets,

    M. A. Jamal and O. Mohareri, “Rethinking RGB-D fusion for semantic segmentation in surgical datasets,” arXiv [cs.CV], July 2024

  96. [104]

    DDA: Dimensionality driven augmentation search for contrastive learning in laparoscopic surgery,

    Y . Zhou, H. Badgery, M. Read, J. Bailey, and C. E. Davey, “DDA: Dimensionality driven augmentation search for contrastive learning in laparoscopic surgery,” arXiv [cs.CV], June 2024

  97. [105]

    Revisiting surgical instrument segmentation without human intervention: A graph partitioning view,

    M. Sheng, J. Fan, D. Liu, R. Kikinis, and W. Cai, “Revisiting surgical instrument segmentation without human intervention: A graph partitioning view,” arXiv [cs.CV], Aug. 2024

  98. [106]

    Contrastive transformer-based multiple instance learning for weakly supervised polyp frame detection,

    Y . Tian, G. Pang, F. Liu, Y . Liu, C. Wang, Y . Chen, J. Verjans, and G. Carneiro, “Contrastive transformer-based multiple instance learning for weakly supervised polyp frame detection,” in International Conference on Medical Image Computing and Computer-Assisted Intervention...

  99. [107]

    Probabilistic representations for video contrastive learning,

    J. Park, J. Lee, I.-J. Kim, and K. Sohn, “Probabilistic representations for video contrastive learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pp. 14711– 14721, 2022

  100. [108]

    Static and dynamic concepts for self-supervised video representation learning,

    R. Qian, S. Ding, X. Liu, and D. Lin, “Static and dynamic concepts for self-supervised video representation learning,” in European conference on computer vision , pp. 145–164, Springer, 2022

  101. [109]

    Parameter- efficient image-to-video transfer learning for action recognition,

    J. Pan, Z. Lin, X. Zhu, J. Shao, and H. L. ST-Adapter, “Parameter- efficient image-to-video transfer learning for action recognition,” Preprint at https://arxiv. org/abs/2206.13559 , 2022

  102. [110]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv [cs.CV], Oct. 2020

  103. [111]

    TEsoNet: knowledge transfer in surgical phase recognition from laparoscopic sleeve gastrectomy to the laparoscopic part of Ivor–Lewis esophagectomy,

    J. A. Eckhoff, Y . Ban, G. Rosman, D. T. M ¨uller, D. A. Hashimoto, E. Witkowski, B. Babic, D. Rus, C. Bruns, H. F. Fuchs, and O. Meire- les, “TEsoNet: knowledge transfer in surgical phase recognition from laparoscopic sleeve gastrectomy to the laparoscopic part of Ivor–Lewis ...

  104. [112]

    C. R. Johnson and R. A. Horn, Matrix analysis. Cambridge university press Cambridge, 1985

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.