Pith. sign in

REVIEW 4 major objections 5 minor 115 references

Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read GRAD claims that disentangling video into spatial, wavelet, and Fourier views, fusing them with kinematics through a graph attention network, and aligning the modalities adversarially produces more accurate and corruption-resistant…

desk verdict Competent architecture, but test-set-tuned hyperparameters and an impossible mean in the robustness table undermine the headline claims. read the letter →

arxiv 2505.01766 v1 pith:XFLSZKS6 submitted 2025-05-03 cs.CV cs.RO

classification cs.CVcs.RO
keywords surgicaldatascienceworkflowrecognitionmultimodalfusionadversariallearningrobustnessgraphrepresentationfeaturedisentanglementnetworkcalibration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes GRAD, a multimodal network for surgical workflow recognition that combines video and kinematic signals. It claims that explicitly separating visual content into spatial, wavelet, and Fourier representations, modeling cross-modal relations with an attention graph, and adversarially aligning the vision and kinematics feature spaces yields higher accuracy than single-modality or prior multimodal methods. The authors report 86.87% accuracy on MISAW and 92.38% on CUHK-MRG, and argue that under 18 types of image corruption the method degrades more slowly than the previous graph-based baseline. If true, this would make automated phase recognition in robotic surgery more reliable under bleeding, smoke, blur, and data-transmission errors.

What carries the argument

The central object is the four-node graph attention network with nodes $\vec{x}^i_t$, $\vec{x}^w_t$, $\vec{x}^f_t$, and $\vec{x}^k_t$ (spatial, wavelet, Fourier, kinematic), updated by multi-head attention coefficients $\alpha_{ij}$. The Vision-Kinematic Adversarial loss $\mathcal{L}_{AL} = 0.5[\log(1-D(\vec{x}^k_t)) + \log D(\vec{x}^i_t)+\log D(\vec{x}^w_t)+\log D(\vec{x}^f_t)]$ pushes the kinematic and visual distributions together. The Contextual Calibrated Decoder blends pre-graph embeddings $E_{v-k}$ with post-graph embeddings $E_g$ as $E=\alpha E_{v-k} + \beta E_g$ and adds a minimal-entropy term to cross-entropy. The graph attention layer is what lets the model weight which modality to trust per frame, and the calibration term is what keeps confidence high when inputs are corrupted.

What would settle it

Recompute the severity-level-4 mean in the robustness table from the 18 per-corruption GRAD values; if it does not equal 88.15, the robustness summary is wrong. Separately, retrain GRAD with all five hyperparameters ($\alpha,\beta,\gamma,\delta,\lambda$) selected on a validation split that is disjoint from the MISAW test split and report the test accuracy; if the margin over MRG-Net narrows to near zero, the state-of-the-art claim is not established.

Watch

Extended reading notes

Core claim

The paper claims that the bottleneck in multimodal surgical phase recognition is not fusion itself but the poverty of each modality's internal representation. By feeding the graph attention network four node types—spatial, wavelet, and Fourier views of the video plus a joint left/right kinematic embedding—the model can read both local texture and global spectral structure, and by aligning the three visual distributions to the kinematic distribution with a discriminator, the model learns a common space in which corruption of any one view is compensated by the others. The authors report that this design reaches 86.87% accuracy on MISAW and 92.38% on CUHK-MRG, and that in the 18-type corruption protocol it degrades more slowly than MRG-Net at every severity level.

Load-bearing premise

The performance and robustness claims assume the evaluation is a true held-out test, but the paper does not report a separate validation split for the five grid-searched hyperparameters, and the robustness table's severity-4 mean is inconsistent with its own column entries.

Editorial extensions

If this is right

  • On the two public benchmarks, GRAD's frame-level accuracy (86.87% on MISAW, 92.38% on CUHK-MRG) exceeds the previous graph-based multimodal baseline by about 3.1 and 2.5 percentage points, with the largest gains in per-class recall and F1.
  • The ablation tables attribute most of the gain to the wavelet-plus-Fourier visual disentanglement; Fourier alone slightly hurts, wavelet alone helps, and both together give the best accuracy.
  • The adversarial alignment works better when the three visual views are the source and kinematics is the target; reversing the roles lowers accuracy.
  • The calibrated decoder's weighted mix of pre-graph and post-graph embeddings plus the entropy penalty is what the robustness experiments credit for slower degradation under 18 corruption types at five severity levels.
  • If the reported robustness holds, multimodal surgical workflow recognition with frequency-domain visual branches remains usable under occlusion from bleeding, smoke, blur, and compression artifacts where a vision-only model fails.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same recipe—splitting a fragile modality into complementary domains, aligning them adversarially to a more reliable modality, and calibrating the fused decoder—is a general pattern that could transfer to audio-visual or sensor-fusion tasks with asymmetric noise.
  • A direct test the authors did not run: corrupt or drop the kinematics channel entirely and check whether the adversarial alignment lets the model fall back on the three visual domains; if the advantage persists with kinematics missing, the compensation story is real.
  • The grid-searched hyperparameters on a single dataset suggest the architecture's advantage may be partly dataset-specific; reporting the same five settings on both benchmarks would clarify how much of the gain is structural rather than tuned.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes GRAD, a multimodal graph representation network with adversarial feature disentanglement for surgical workflow recognition. It combines visual representations in spatial, wavelet, and Fourier domains with kinematic embeddings via a graph attention network, aligns modalities adversarially, and adds a calibrated prediction decoder with a modified cross-entropy loss. The authors report state-of-the-art accuracy on MISAW (86.87%) and CUHK-MRG (92.38%) and claim excellent robustness to 18 image corruptions, comparing against MRG-Net. The contributions are presented as a novel architecture and a robustness evaluation protocol.

Significance. If the reported results were valid, GRAD would be a meaningful advance in multimodal surgical workflow recognition: the idea of fusing three visual domains with kinematic data through graph attention and adversarial alignment is technically interesting, and the robustness evaluation on a standard corruption benchmark is a welcome step beyond clean-accuracy comparisons. The paper also includes extensive ablations of each module and comparisons with many baseline methods. However, the central claims are not supported by the evidence as presented: the headline accuracy is produced by selecting multiple hyperparameters on the test set, and the main robustness table contains an internal inconsistency. The methodological value of the proposed components cannot be assessed from the current experiments.

major comments (4)
  1. [§4.3.5, §4.3.7, Tables 7/9/10] The hyperparameters alpha, beta, gamma, delta, and lambda are selected by grid search on the MISAW test split, and the final results in Tables 1 and 12 are reported on that same split. Sections 4.3.5 and 4.3.7 explicitly state that the chosen values (alpha=0.3, beta=0.7; gamma=0.9, delta=0.1; lambda=0.02) are the ones that achieve the highest test accuracy. No held-out validation split, nested cross-validation, or repeated-seed error bars are described. This makes the reported 86.87% accuracy and the robustness margins over MRG-Net partly fitted values, not independent predictions. The state-of-the-art claim is therefore not established and requires a proper evaluation protocol.
  2. [Table 12] The severity-level-4 mean accuracy for GRAD is reported as 88.15, but all 18 per-corruption entries in that column range from 73.23 to 81.18. No average of these values can equal 88.15, so the table is internally inconsistent. Additionally, the entries for 'Gauss' under the Noise group and under the Blur group are identical for both MRG-Net and GRAD, suggesting a duplicated column that should presumably be Gaussian noise and Gaussian blur. As presented, Table 12 cannot support the conclusion of 'excellent stability and robustness.' The per-corruption comparisons may still be salvageable, but the summary table and the narrative built on it need to be corrected and recomputed.
  3. [§3.2.3, Eq. (15)] The derivation of the calibrated loss L_CCE is mathematically incorrect. In Eq. (15), the term P*H(P_hat) is treated as a sum over classes, but H(P_hat) is a scalar, so P*H(P_hat) is a scalar multiple of the target vector, not a per-class term. The expansion to -sum p_i p_hat_i log p_hat_i does not follow. For a one-hot target, the added regularization -lambda * p_i * p_hat_i * log p_hat_i equals -lambda * p_hat_c * log p_hat_c for the correct class; this expression is minimized when p_hat_c approaches e^{-1}, not 1. Thus the loss does not 'amplify confidence' as claimed, and the mechanism underlying the calibration module is not established.
  4. [Table 3 and §4.3.2] The text in Section 4.3.2 states that 'introducing any new technique results in a performance boost,' but Table 3 shows that row 2 (CAL only, no VRD, no VKA) achieves accuracy 82.12 and row 4 (VKA only, no CAL, no VRD) achieves 83.01, both below the baseline of 83.75 in row 1. The ablation table is therefore internally inconsistent with the accompanying narrative, and the claimed effectiveness of the individual modules is not demonstrated as stated.
minor comments (5)
  1. [§4.3.1] The text says 'Our GRAD framework achieves the best overall accuracy of 92.83%,' but Table 2 reports accuracy 92.38% for GRAD on CUHK-MRG. The numbers should be reconciled.
  2. [Table 12] The corruption type 'Gauss' appears twice, once under Noise and once under Blur, with identical numerical values. These should be labeled distinctly (e.g., Gaussian noise and Gaussian blur) and the values should be independently computed.
  3. [§4.1] The MISAW dataset description mentions only a training/test split; please clarify whether any validation split was used during development and how the final model was selected.
  4. [Eq. (6)] The attention update in Eq. (6) includes self-coefficients alpha_ii, but the attention computation in Eq. (7) only defines coefficients for neighbors j in {i, w, f, k}. Please clarify how self-attention is handled in the graph layer.
  5. [Figure 2] Figure 2 is very dense; the labels for the calibrated decoder and the adversarial branch are hard to read. A larger figure or a separate diagram for the decoder would improve clarity.

Circularity Check

1 steps flagged · score 6.0 of 10

MISAW headline accuracy is selected on the MISAW test set via hyperparameter grid search, making the central MISAW 'prediction' partly fitted; CUHK-MRG cross-validation provides independent support.

  1. fitted input called prediction [Section 4.3.5 (Eq. 12, Table 7), Section 4.3.7 (Eq. 16, Tables 9-10), reported in Table 1]
    "The dataset was divided into a training set with 17 cases and a testing set with 10 cases. ... Finally, 𝛼 = 0.3 and 𝛽 = 0.7 achieve the highest performance across multiple metrics; therefore, we use this configuration as the final parameters for our model. ... Ultimately, by analyzing the results, we find the best performance with a ratio of 0.9 for 𝐶𝐶𝐸 and 0.1 for adversarial loss 𝐴𝐿. ... Table 10 shows that we achieve the best results when 𝜆 = 0.02."

    All five hyperparameters (𝛼, 𝛽, 𝛾, 𝛿, 𝜆) are selected by grid search on the MISAW test split: Tables 7, 9, and 10 are labelled as experiments on the MISAW dataset, and Section 4.1 states that this dataset has only a 17-case training set and a 10-case testing set, with no validation split described. The final 86.87% accuracy in Table 1 is then reported on the same 10-case test set used to choose those hyperparameters. The reported MISAW result is therefore a test-set-selected maximum rather than an independent prediction of the method; the 'superior performance' claim for MISAW is partly constructed from the test labels. The CUHK-MRG 3-fold cross-validation results provide some independent grounding, which is why this is partial rather than total circularity.

full rationale

The only clear circularity is evaluation-based: the paper tunes its five loss and fusion hyperparameters on the MISAW test set and then reports accuracy on that same test set as the headline result. Tables 7, 9, and 10 are ablation sweeps over 𝛼, 𝛽, 𝛾, 𝛿, and 𝜆 conducted on MISAW, and Section 4.3.5/4.3.7 explicitly adopt the values that perform best on those sweeps. Table 1 then cites 86.87% for GRAD on MISAW, so the MISAW 'prediction' is partly selected, not independently predicted. This fits the 'fitted input called prediction' pattern. No load-bearing self-citation was found: the MRG-Net baseline [59] is an external prior work, not by the present authors, and no uniqueness theorem or ansatz is imported from the authors' own prior work. The CUHK-MRG results use 3-fold cross-validation and provide independent evidence for the method, which prevents the circularity from being total. Separately, but not as a circularity, Table 12 is internally inconsistent: the GRAD severity-level-4 mean of 88.15 exceeds every per-corruption GRAD value in that column (the largest is about 81.18), so the robustness claim based on that table is not arithmetically supported and should be re-examined on correctness grounds. Overall, the central MISAW superiority claim is partially compromised by test-set selection, while the CUHK-MRG evaluation gives the paper some independent content, yielding a score of 6.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The model is built from standard deep learning components; the only novel numerical choices are the fusion and loss weights that are fitted to the MISAW test set. The paper introduces no new physical or ontological entities. The domain assumptions concern dataset representativeness and the transferability of synthetic image corruptions to surgical scenes.

free parameters (5)
  • alpha (vision-kinematic embedding weight) = 0.3
    Weight of the pre-graph vision-kinematic embedding in Equation 12; selected by grid search on the MISAW test set (Table 7), affecting the final prediction input.
  • beta (graph embedding weight) = 0.7
    Weight of the graph output embedding in Equation 12; selected together with alpha on the MISAW test set (Table 7).
  • gamma (calibrated cross-entropy loss weight) = 0.9
    Coefficient for L_CCE in Equation 16; grid-searched on the MISAW test set (Table 9).
  • delta (adversarial loss weight) = 0.1
    Coefficient for L_AL in Equation 16; grid-searched on the MISAW test set (Table 9).
  • lambda (calibration regularization coefficient) = 0.02
    Regularization strength in Equation 15; grid-searched on the MISAW test set (Table 10).
assumptions (5)
  • standard math Standard definitions and implementations of GAT, TCN, LSTM, ResNet, and wavelet and Fourier transforms are correct and suitable as building blocks.
    Used without formal proof; these are well-established techniques from the cited literature.
  • domain assumption The MISAW and CUHK-MRG datasets provide reliable ground-truth gesture labels and are representative of surgical workflow recognition tasks.
    The evaluation trusts the dataset splits and annotations from Huaulme et al. and Long et al.; no independent verification is provided.
  • ad hoc to paper Adversarial alignment of vision and kinematic feature distributions improves accuracy and robustness.
    Modeling assumption behind the Vision-Kinematic Adversarial module; supported only by the paper's own ablations, not by external theory or benchmarks.
  • ad hoc to paper The confidence-amplifying loss L_CCE improves model calibration and robustness.
    The paper never measures calibration directly, and the derivation in Section 3.2.3 contains a sign inconsistency; the claim is asserted rather than demonstrated.
  • domain assumption ImageNet-style corruptions from Hendrycks and Dietterich are representative of real surgical data corruption.
    Robustness experiments apply 18 synthetic corruption types to MISAW video frames; the transfer of these corruptions to clinical surgical degradation is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement." pith.science (2026). https://pith.science/paper/XFLSZKS6

@misc{pith2026250501766,
  author       = {Pith},
  title        = {Pith review of: Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XFLSZKS6}},
  note         = {Machine review of arXiv:2505.01766}
}
read the original abstract

Surgical workflow recognition is vital for automating tasks, supporting decision-making, and training novice surgeons, ultimately improving patient safety and standardizing procedures. However, data corruption can lead to performance degradation due to issues like occlusion from bleeding or smoke in surgical scenes and problems with data storage and transmission. In this case, we explore a robust graph-based multimodal approach to integrating vision and kinematic data to enhance accuracy and reliability. Vision data captures dynamic surgical scenes, while kinematic data provides precise movement information, overcoming limitations of visual recognition under adverse conditions. We propose a multimodal Graph Representation network with Adversarial feature Disentanglement (GRAD) for robust surgical workflow recognition in challenging scenarios with domain shifts or corrupted data. Specifically, we introduce a Multimodal Disentanglement Graph Network that captures fine-grained visual information while explicitly modeling the complex relationships between vision and kinematic embeddings through graph-based message modeling. To align feature spaces across modalities, we propose a Vision-Kinematic Adversarial framework that leverages adversarial training to reduce modality gaps and improve feature consistency. Furthermore, we design a Contextual Calibrated Decoder, incorporating temporal and contextual priors to enhance robustness against domain shifts and corrupted data. Extensive comparative and ablation experiments demonstrate the effectiveness of our model and proposed modules. Moreover, our robustness experiments show that our method effectively handles data corruption during storage and transmission, exhibiting excellent stability and robustness. Our approach aims to advance automated surgical workflow recognition, addressing the complexities and dynamism inherent in surgical procedures.

Figures

Figures reproduced from arXiv: 2505.01766 by the authors.

Figure 1
Figure 1. Comparison between the vision-based workflow recognition framework and the multimodal workflow recogni￾tion framework. sequence and structure of surgical tasks, workflow recogni￾tion systems can provide real-time feedback, facilitate error correction, and contribute to the standardization of surgical procedures, ultimately leading to improved patient safety and better training methodologies [10]. Vision-based surgic… view at source ↗
Figure 2
Figure 2. The overall architecture of our GRAD framework, consisting of the Multimodal Disentanglement Graph Network (Visual Representation Disentanglement module, the Kinematic Temporal Representation Extraction module, the Multimodal Graph Learning module), the Vision-Kinematic Adversarial Training, and the Calibrated Prediction Decoder. Our GRAD framework extracts visual and kinematic features through intra-modal feature m… view at source ↗
Figure 3
Figure 3. Visualization of GRAD compared to MRG-Net (Vis+Kin), Trans-SVNet (Vis), and RL-TCN (Kin) model on MISAW dataset. The different color bands represent each category, which are needle holding, suture making, suture handling, 1 ◦ knot, 2 ◦ knot, and 3 ◦ knot (with order). term, the new loss function can be formulated as: 𝐶𝐶𝐸 = 𝐶𝐸 − 𝜆𝑃 (𝑃̂) = −∑ 𝐾 𝑖=1 𝑝𝑖 log ̂𝑝𝑖 − 𝜆 ∑ 𝐾 𝑖=1 𝑝𝑖 ̂𝑝𝑖 log ̂𝑝𝑖 = −∑ 𝐾 𝑖=1 ( 1 + 𝜆 ̂𝑝𝑖 ) 𝑝𝑖 l… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Figure of CUHK-MRG Dataset illustrates five different gesture steps, which are idle (no action performed), reach for peg (with left hand), lift peg (with left hand), exchange (transfer the peg to right hand), place peg (with right hand) [PITH_FULL_IMAGE:figures/full_f…
Figure 5
Figure 5. Figure 5: An example of corrupted data visualization for robustness assessment, featuring four different types ranging from noise to digital damage. 82.53 80.42 78.82 76.40 74.94 79.73 75.61 72.27 69.55 60.22 60.00 65.00 70.00 75.00 80.00 85.00 1 2 3 4 5 Accuracy Severity Level …
Figure 6
Figure 6. Figure 6: Robustness experiments on the MISAW dataset under 18 different corruptions. L. Bai et al.: Preprint submitted to Elsevier Page 15 of 20 [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: Robustness experiments on the MISAW dataset at each severity level. queries; K2V: kinematics as attention queries) [51], Self￾Attention [82], Gated Fusion [6], BAN [40], TDA [2], AFF & iAFF [17]. Specifically, we replace our graph-based fusion with these different meth…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

115 extracted references · 57 canonical work pages

  1. [1]

    Adatasetand benchmarks for segmentation and recognition of gestures in robotic surgery

    Ahmidi, N., Tao, L., Sefati, S., Gao, Y., Lea, C., Haro, B.B., Zap- pella,L.,Khudanpur,S.,Vidal,R.,Hager,G.D.,2017. Adatasetand benchmarks for segmentation and recognition of gestures in robotic surgery. IEEE Transactions on Biomedical Engineering 64, 2025– 2041

  2. [2]

    Bottom-up and top-down attention for image captioning and visual question answering, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

    Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., Zhang, L., 2018. Bottom-up and top-down attention for image captioning and visual question answering, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6077–6086

  3. [3]

    Pixel-wise recognition for holistic surgical scene understanding

    Ayobi, N., Rodríguez, S., Pérez, A., Hernández, I., Aparicio, N., Dessevres, E., Peña, S., Santander, J., Caicedo, J.I., Fernández, N., et al., 2024. Pixel-wise recognition for holistic surgical scene understanding. arXiv preprint arXiv:2401.11174

  4. [4]

    Bai, L., Chen, T., Wu, Y., Wang, A., Islam, M., Ren, H., 2023a. Llcaps: Learning to illuminate low-light capsule endoscopy with curved wavelet attention and reverse diffusion, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 34–44

  5. [5]

    Bai, L., Islam, M., Ren, H., 2023b. Cat-vil: co-attention gated vision-language embedding for visual question localized-answering in robotic surgery, in: International Conference on Medical Image ComputingandComputer-AssistedIntervention,Springer.pp.397– 407

  6. [6]

    Bai, L., Islam, M., Seenivasan, L., Ren, H., 2023c. Surgical- vqla: Transformer with gated vision-language embedding for visual question localized-answering in robotic surgery, in: 2023 IEEE International Conference on Robotics and Automation (ICRA), pp. 6859–6865. doi:10.1109/ICRA48891.2023.10160403

  7. [7]

    Ossar: Towards open-set surgi- cal activity recognition in robot-assisted surgery

    Bai, L., Wang, G., Wang, J., Yang, X., Gao, H., Liang, X., Wang, A., Islam, M., Ren, H., 2024. Ossar: Towards open-set surgi- cal activity recognition in robot-assisted surgery. arXiv preprint arXiv:2402.06985

  8. [8]

    Classifier calibration with roc-regularized isotonic regression, in: International Conference on Artificial Intelligence and Statistics, PMLR

    Berta, E., Bach, F., Jordan, M., 2024. Classifier calibration with roc-regularized isotonic regression, in: International Conference on Artificial Intelligence and Statistics, PMLR. pp. 1972–1980

Show all 115 references
  1. [9]

    Weight uncertainty in neural network, in: International conference on machine learning, PMLR

    Blundell, C., Cornebise, J., Kavukcuoglu, K., Wierstra, D., 2015. Weight uncertainty in neural network, in: International conference on machine learning, PMLR. pp. 1613–1622

  2. [10]

    Intelligent surgicalworkflowrecognitionforendoscopicsubmucosaldissection with real-time animal study

    Cao, J., Yip, H.C., Chen, Y., Scheppach, M., Luo, X., Yang, H., Cheng,M.K.,Long,Y.,Jin,Y.,Chiu,P.W.Y.,etal.,2023. Intelligent surgicalworkflowrecognitionforendoscopicsubmucosaldissection with real-time animal study. Nature Communications 14, 6676

  3. [11]

    Multimodal visual-textual object graph attention network for propaganda detec- tion in memes

    Chen, P., Zhao, L., Piao, Y., Ding, H., Cui, X., 2023. Multimodal visual-textual object graph attention network for propaganda detec- tion in memes. Multimedia Tools and Applications , 1–16. L. Bai et al.:Preprint submitted to Elsevier Page 17 of 20 Multimodal Graph Representa...

  4. [12]

    Lightdiff:Surgicalendoscopicimagelow-lighten- hancementwitht-diffusion,in:InternationalConferenceonMedical Image Computing and Computer-Assisted Intervention, Springer

    Chen, T., Lyu, Q., Bai, L., Guo, E., Gao, H., Yang, X., Ren, H., Zhou,L.,2024a. Lightdiff:Surgicalendoscopicimagelow-lighten- hancementwitht-diffusion,in:InternationalConferenceonMedical Image Computing and Computer-Assisted Intervention, Springer. pp. 369–379

  5. [13]

    Surgsora: Decoupled rgbd-flow diffusion model for controllable surgical video generation

    Chen, T., Yang, S., Wang, J., Bai, L., Ren, H., Zhou, L., 2024b. Surgsora: Decoupled rgbd-flow diffusion model for controllable surgical video generation. arXiv preprint arXiv:2412.14018

  6. [14]

    On the properties of neural machine translation: Encoder-decoder approaches

    Cho, K., Van Merriënboer, B., Bahdanau, D., Bengio, Y., 2014. On the properties of neural machine translation: Encoder-decoder approaches. arXiv preprint arXiv:1409.1259

  7. [15]

    Surgical-dino: Adapter learning of foundation model for depth estimation in endoscopic surgery

    Cui, B., Islam, M., Bai, L., Ren, H., 2024. Surgical-dino: Adapter learning of foundation model for depth estimation in endoscopic surgery. arXiv preprint arXiv:2401.06013

  8. [16]

    Czempiel, T., Paschali, M., Keicher, M., Simson, W., Feussner, H., Kim,S.T.,Navab,N.,2020. Tecno:Surgicalphaserecognitionwith multi-stage temporal convolutional networks, in: Medical Image Computing and Computer Assisted Intervention–MICCAI 2020: 23rd International Conference,...

  9. [17]

    Attentional feature fusion, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp

    Dai, Y., Gieseke, F., Oehmcke, S., Wu, Y., Barnard, K., 2021. Attentional feature fusion, in: Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 3560–3569

  10. [18]

    Dergachyova,O.,Bouget,D.,Huaulmé,A.,Morandi,X.,Jannin,P.,

  11. [19]

    Understanding how image quality af- fectsdeepneuralnetworks,in:2016eighthinternationalconference on quality of multimedia experience (QoMEX), IEEE

    Dodge, S., Karam, L., 2016. Understanding how image quality af- fectsdeepneuralnetworks,in:2016eighthinternationalconference on quality of multimedia experience (QoMEX), IEEE. pp. 1–6

  12. [20]

    Surgicalmotionanalysisusingdiscriminativeinterpretablepatterns

    Forestier, G., Petitjean, F., Senin, P., Despinoy, F., Huaulmé, A., Fawaz,H.I.,Weber,J.,Idoumghar,L.,Muller,P.A.,Jannin,P.,2018. Surgicalmotionanalysisusingdiscriminativeinterpretablepatterns. Artificial intelligence in medicine 91, 3–11

  13. [21]

    Bayesian recurrent neural networks

    Fortunato, M., Blundell, C., Vinyals, O., 2017. Bayesian recurrent neural networks. arXiv preprint arXiv:1704.02798

  14. [22]

    Transendoscopic flexible parallel continuum robotic mechanism for bimanual endo- scopicsubmucosaldissection

    Gao, H., Yang, X., Xiao, X., Zhu, X., Zhang, T., Hou, C., Liu, H., Meng, M.Q.H., Sun, L., Zuo, X., et al., 2023. Transendoscopic flexible parallel continuum robotic mechanism for bimanual endo- scopicsubmucosaldissection. TheInternationalJournalofRobotics Research , 02783649231209338

  15. [23]

    Generativeadversarial networks

    Goodfellow,I., Pouget-Abadie,J.,Mirza, M.,Xu,B., Warde-Farley, D.,Ozair,S.,Courville,A.,Bengio,Y.,2020. Generativeadversarial networks. Communications of the ACM 63, 139–144

  16. [24]

    Longshort-termmemory

    Graves,A.,Graves,A.,2012. Longshort-termmemory. Supervised sequence labelling with recurrent neural networks , 37–45

  17. [25]

    Oncalibration ofmodernneuralnetworks,in:Internationalconferenceonmachine learning, PMLR

    Guo,C.,Pleiss,G.,Sun,Y.,Weinberger,K.Q.,2017. Oncalibration ofmodernneuralnetworks,in:Internationalconferenceonmachine learning, PMLR. pp. 1321–1330

  18. [26]

    A systematic review of robotic surgery: From supervised paradigms to fully autonomous robotic approaches

    Han,J.,Davids,J.,Ashrafian,H.,Darzi,A.,Elson,D.S.,Sodergren, M., 2022. A systematic review of robotic surgery: From supervised paradigms to fully autonomous robotic approaches. The Interna- tional Journal of Medical Robotics and Computer Assisted Surgery 18, e2358

  19. [27]

    Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

    He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778

  20. [28]

    Benchmarking neural network robustnesstocommoncorruptionsandperturbations

    Hendrycks, D., Dietterich, T., 2019. Benchmarking neural network robustnesstocommoncorruptionsandperturbations. arXivpreprint arXiv:1903.12261

  21. [29]

    Long short-term memory

    Hochreiter, S., Schmidhuber, J., 1997. Long short-term memory. Neural computation 9, 1735–80. doi:10.1162/neco.1997.9.8.1735

  22. [30]

    Hybrid graph convolutional network with online masked autoencoder for robust multimodal cancer survival prediction

    Hou, W., Lin, C., Yu, L., Qin, J., Yu, R., Wang, L., 2023. Hybrid graph convolutional network with online masked autoencoder for robust multimodal cancer survival prediction. IEEE Transactions on Medical Imaging 42, 2462–2473. doi:10.1109/TMI.2023.3253760

  23. [31]

    What makes graph neural networks miscalibrated? Advances in Neural Information Processing Systems 35, 13775–13786

    Hsu, H.H.H., Shen, Y., Tomani, C., Cremers, D., 2022. What makes graph neural networks miscalibrated? Advances in Neural Information Processing Systems 35, 13775–13786

  24. [32]

    Post- hoc reward calibration: A case study on length bias

    Huang, Z., Qiu, Z., Wang, Z., Ponti, E.M., Titov, I., 2024. Post- hoc reward calibration: A case study on length bias. arXiv preprint arXiv:2409.17407

  25. [33]

    Micro-surgical anastomose workflow recognition challenge report

    Huaulmé,A.,Sarikaya,D.,LeMut,K.,Despinoy,F.,Long,Y.,Dou, Q., Chng, C.B., Lin, W., Kondo, S., Bravo-Sánchez, L., et al., 2021. Micro-surgical anastomose workflow recognition challenge report. Computer Methods and Programs in Biomedicine 212, 106452

  26. [34]

    Class- distribution-aware calibration for long-tailed visual recognition

    Islam, M., Seenivasan, L., Ren, H., Glocker, B., 2021. Class- distribution-aware calibration for long-tailed visual recognition. arXiv preprint arXiv:2109.05263

  27. [35]

    Categorical reparameterization with gumbel-softmax

    Jang, E., Gu, S., Poole, B., 2016. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144

  28. [36]

    Sv-rcnet:workflowrecognitionfromsurgicalvideosusingrecurrent convolutional network

    Jin,Y.,Dou,Q.,Chen,H.,Yu,L.,Qin,J.,Fu,C.W.,Heng,P.A.,2017. Sv-rcnet:workflowrecognitionfromsurgicalvideosusingrecurrent convolutional network. IEEE transactions on medical imaging 37, 1114–1126

  29. [37]

    Multi-task recurrent convolutional network with correlation loss for surgical video analysis

    Jin,Y.,Li,H.,Dou,Q.,Chen,H.,Qin,J.,Fu,C.W.,Heng,P.A.,2020. Multi-task recurrent convolutional network with correlation loss for surgical video analysis. Medical image analysis 59, 101572

  30. [38]

    Temporal memory relation network for workflow recognition from surgical video

    Jin, Y., Long, Y., Chen, C., Zhao, Z., Dou, Q., Heng, P.A., 2021. Temporal memory relation network for workflow recognition from surgical video. IEEE Transactions on Medical Imaging 40, 1911– 1923

  31. [39]

    Trans-svnet:hybridembeddingaggregationtransformerforsurgical workflow analysis

    Jin, Y., Long, Y., Gao, X., Stoyanov, D., Dou, Q., Heng, P.A., 2022. Trans-svnet:hybridembeddingaggregationtransformerforsurgical workflow analysis. International Journal of Computer Assisted Radiology and Surgery 17, 2193–2202

  32. [40]

    Bilinear attention networks

    Kim, J.H., Jun, J., Zhang, B.T., 2018. Bilinear attention networks. Advances in neural information processing systems 31

  33. [41]

    Uncertaintycalibrationwithenergybased instance-wise scaling in the wild dataset, in: European Conference on Computer Vision, Springer

    Kim,M.,Kwon,J.,2025. Uncertaintycalibrationwithenergybased instance-wise scaling in the wild dataset, in: European Conference on Computer Vision, Springer. pp. 232–248

  34. [42]

    Semi-supervised classification with graph convolutional networks

    Kipf, T.N., Welling, M., 2016. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907

  35. [43]

    Simple and scalable predictive uncertainty estimation using deep ensembles

    Lakshminarayanan, B., Pritzel, A., Blundell, C., 2017. Simple and scalable predictive uncertainty estimation using deep ensembles. Advances in neural information processing systems 30

  36. [44]

    Tem- poralconvolutionalnetworksforactionsegmentationanddetection, in: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp

    Lea,C.,Flynn,M.D.,Vidal,R.,Reiter,A.,Hager,G.D.,2017. Tem- poralconvolutionalnetworksforactionsegmentationanddetection, in: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 156–165

  37. [45]

    Lea, C., Reiter, A., Vidal, R., Hager, G.D., 2016a. Segmental spatiotemporal cnns for fine-grained action segmentation, in: Com- puter Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part III 14, Springer. pp. 36–52

  38. [46]

    Learning convolutional action primitives for fine-grained action recognition, in: 2016 IEEE international conference on robotics and automation (ICRA), IEEE

    Lea, C., Vidal, R., Hager, G.D., 2016b. Learning convolutional action primitives for fine-grained action recognition, in: 2016 IEEE international conference on robotics and automation (ICRA), IEEE. pp. 1642–1649

  39. [47]

    Lea, C., Vidal, R., Reiter, A., Hager, G.D., 2016c. Temporal convolutional networks: A unified approach to action segmentation, in: Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part III 14, Springer. pp. 47–54

  40. [48]

    Three-dimensional collision avoidance method for robot-assisted minimally invasive surgery

    Li, L., Li, X., Ouyang, B., Mo, H., Ren, H., Yang, S., 2023. Three-dimensional collision avoidance method for robot-assisted minimally invasive surgery. Cyborg and Bionic Systems 4, 0042

  41. [49]

    Multimodal graph learning based on 3d haar semi-tight framelet for student engage- ment prediction

    Li, M., Zhuang, X., Bai, L., Ding, W., 2024a. Multimodal graph learning based on 3d haar semi-tight framelet for student engage- ment prediction. Information Fusion 105, 102224

  42. [50]

    Efficiency calibration of implicit regularization in deep networks via self-paced curriculum- driven singular value selection, in: Larson, K

    Li, Z., Chen, S., Yang, J., Luo, L., 2024b. Efficiency calibration of implicit regularization in deep networks via self-paced curriculum- driven singular value selection, in: Larson, K. (Ed.), Proceedings of the Thirty-Third International Joint Conference on Artificial In- tel...

  43. [51]

    Cat: Cross attention in vision transformer, in: 2022 IEEE international conference on multimedia and expo (ICME), IEEE

    Lin, H., Cheng, X., Wu, X., Shen, D., 2022. Cat: Cross attention in vision transformer, in: 2022 IEEE international conference on multimedia and expo (ICME), IEEE. pp. 1–6

  44. [52]

    Focal loss for dense object detection

    Lin, T., 2017. Focal loss for dense object detection. arXiv preprint arXiv:1708.02002

  45. [53]

    Liu, D., Jiang, T., 2018. Deep reinforcement learning for surgical gesture segmentation and classification, in: Medical Image Com- puting and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, September 16-20, 2018, Proceedings, Part IV ...

  46. [54]

    Liu,J.,Fan,X.,Huang,Z.,Wu,G.,Liu,R.,Zhong,W.,Luo,Z.,2022. Target-aware dual adversarial learning and a multi-scenario multi- modalitybenchmarktofuseinfraredandvisibleforobjectdetection, in:ProceedingsoftheIEEE/CVFconferenceoncomputervisionand pattern recognition, pp. 5802–5811

  47. [55]

    Visual-kinematics graph learning for procedure-agnostic instrumenttipsegmentationinroboticsurgeries,in:2023IEEE/RSJ InternationalConferenceonIntelligentRobotsandSystems(IROS), IEEE

    Liu, J., Long, Y., Chen, K., Leung, C.H., Wang, Z., Dou, Q., 2023a. Visual-kinematics graph learning for procedure-agnostic instrumenttipsegmentationinroboticsurgeries,in:2023IEEE/RSJ InternationalConferenceonIntelligentRobotsandSystems(IROS), IEEE. pp. 4633–4639

  48. [56]

    Uncertainty estimation and quantification for llms: A simple supervised approach

    Liu, L., Pan, Y., Li, X., Chen, G., 2024. Uncertainty estimation and quantification for llms: A simple supervised approach. arXiv preprint arXiv:2404.15993

  49. [57]

    Lovit: Long video transformerforsurgicalphaserecognition

    Liu, Y., Boels, M., Garcia-Peraza-Herrera, L.C., Vercauteren, T., Dasgupta, P., Granados, A., Ourselin, S., 2025. Lovit: Long video transformerforsurgicalphaserecognition. MedicalImageAnalysis 99, 103366

  50. [58]

    Skit: a fast key information video trans- former for online surgical phase recognition, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

    Liu, Y., Huo, J., Peng, J., Sparks, R., Dasgupta, P., Granados, A., Ourselin, S., 2023b. Skit: a fast key information video trans- former for online surgical phase recognition, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 21074–21084

  51. [59]

    Long, Y., Wu, J.Y., Lu, B., Jin, Y., Unberath, M., Liu, Y.H., Heng, P.A., Dou, Q., 2021. Relational graph learning on visual and kinematics embeddings for accurate gesture recognition in robotic surgery, in: 2021 IEEE International Conference on Robotics and Automation (ICRA),...

  52. [60]

    Mai, S., Hu, H., Xing, S., 2020. Modality to modality translation: An adversarial representation learning and graph fusion network for multimodal fusion, in: proceedings of the AAAI conference on artificial intelligence, pp. 164–172

  53. [61]

    Multimodal graph for unaligned multimodal sequence analysis via graph convolution and graph pooling

    Mai, S., Xing, S., He, J., Zeng, Y., Hu, H., 2023. Multimodal graph for unaligned multimodal sequence analysis via graph convolution and graph pooling. ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS 19. doi:10.1145/3542927

  54. [62]

    IEEE/ACM Transactions on Audio, Speech, and Language Processing 32, 4298–4312

    Meng,T.,Zhang,F.,Shou,Y.,Shao,H.,Ai,W.,Li,K.,2024.Masked graph learning with recurrent alignment for multimodal emotion recognition in conversation. IEEE/ACM Transactions on Audio, Speech, and Language Processing 32, 4298–4312

  55. [63]

    Moglia,A.,Georgiou,K.,Georgiou,E.,Satava,R.M.,Cuschieri,A.,

  56. [64]

    Obtaining well calibrated probabilities using bayesian binning, in: Proceedings of the AAAI conference on artificial intelligence

    Naeini, M.P., Cooper, G., Hauskrecht, M., 2015. Obtaining well calibrated probabilities using bayesian binning, in: Proceedings of the AAAI conference on artificial intelligence

  57. [65]

    Adaptive multi-hypergraph convolutional networks for 3d object classification

    Nong, L., Peng, J., Zhang, W., Lin, J., Qiu, H., Wang, J., 2022. Adaptive multi-hypergraph convolutional networks for 3d object classification. IEEE Transactions on Multimedia

  58. [66]

    Impact of regularization on cal- ibration and robustness: from the representation space perspective

    Park, J., Kim, J., Lee, J.S., 2024. Impact of regularization on cal- ibration and robustness: from the representation space perspective. arXiv preprint arXiv:2410.03999

  59. [67]

    Transformer uncertainty estimation with hierarchical stochastic attention, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp

    Pei, J., Wang, C., Szarvas, G., 2022. Transformer uncertainty estimation with hierarchical stochastic attention, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 11147–11155

  60. [68]

    Pereyra, G., Tucker, G., Chorowski, J., Kaiser, Ł., Hinton, G.,

  61. [69]

    Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods

    Platt, J., et al., 1999. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Ad- vances in large margin classifiers 10, 61–74

  62. [70]

    Popordanoska, T., Tiulpin, A., Blaschko, M.B., 2024. Beyond classification:Definitionanddensity-basedestimationofcalibration in object detection, in: Proceedings of the IEEE/CVF Winter Con- ference on Applications of Computer Vision, pp. 585–594

  63. [71]

    Sar-rarp50: Segmentation of surgical instrumentation and action recognitiononrobot-assistedradicalprostatectomychallenge

    Psychogyios,D.,Colleoni,E.,VanAmsterdam,B.,Li,C.Y.,Huang, S.Y., Li, Y., Jia, F., Zou, B., Wang, G., Liu, Y., et al., 2023. Sar-rarp50: Segmentation of surgical instrumentation and action recognitiononrobot-assistedradicalprostatectomychallenge. arXiv preprint arXiv:2401.00496

  64. [72]

    History of robotic surgery: from aesop® and zeus® to da vinci®

    Pugin, F., Bucher, P., Morel, P., 2011. History of robotic surgery: from aesop® and zeus® to da vinci®. Journal of visceral surgery 148, e3–e8

  65. [73]

    Qin, Y., Pedram, S.A., Feyzabadi, S., Allan, M., McLeod, A.J., Burdick,J.W.,Azizian,M.,2020.Temporalsegmentationofsurgical sub-tasks through deep learning with multiple data sources, in: 2020 IEEE International Conference on Robotics and Automation (ICRA), IEEE. pp. 371–377

  66. [74]

    Ensemble manifold regularized multi-modal graph convolutional network for cognitive ability prediction

    Qu, G., Xiao, L., Hu, W., Wang, J., Zhang, K., Calhoun, V.D., Wang, Y.P., 2021. Ensemble manifold regularized multi-modal graph convolutional network for cognitive ability prediction. IEEE Transactions on Biomedical Engineering 68, 3564–3573

  67. [75]

    On the pitfalls of batch normalization for end-to-end video learning: a study on surgical workflow analysis

    Rivoir, D., Funke, I., Speidel, S., 2024. On the pitfalls of batch normalization for end-to-end video learning: a study on surgical workflow analysis. Medical Image Analysis 94, 103126

  68. [76]

    Schlichtkrull, M., Kipf, T.N., Bloem, P., Van Den Berg, R., Titov, I., Welling, M., 2018. Modeling relational data with graph con- volutional networks, in: The semantic web: 15th international con- ference, ESWC 2018, Heraklion, Crete, Greece, June 3–7, 2018, proceedings 15, S...

  69. [77]

    Enhancing adversarialrobustnessofmulti-modalrecommendationviamodality balancing, in: Proceedings of the 31st ACM International Confer- ence on Multimedia, pp

    Shang,Y.,Gao,C.,Chen,J.,Jin,D.,Ma,H.,Li,Y.,2023. Enhancing adversarialrobustnessofmulti-modalrecommendationviamodality balancing, in: Proceedings of the 31st ACM International Confer- ence on Multimedia, pp. 6274–6282

  70. [78]

    Wavelet transforms versus fourier transforms

    Strang, G., 1993. Wavelet transforms versus fourier transforms. Bulletin of the American Mathematical Society 28, 288–305

  71. [79]

    Wavelet- based mamba with fourier adjustment for low-light image enhance- ment,in:ProceedingsoftheAsianConferenceonComputerVision, pp

    Tan, J., Pei, S., Qin, W., Fu, B., Li, X., Huang, L., 2024. Wavelet- based mamba with fourier adjustment for low-light image enhance- ment,in:ProceedingsoftheAsianConferenceonComputerVision, pp. 3449–3464

  72. [80]

    Are graph neural networks miscalibrated? arXiv preprint arXiv:1905.02296

    Teixeira, L., Jalaian, B., Ribeiro, B., 2019. Are graph neural networks miscalibrated? arXiv preprint arXiv:1905.02296

  73. [81]

    Endonet: a deep architecture for recognition tasksonlaparoscopicvideos

    Twinanda,A.P.,Shehata,S.,Mutter,D.,Marescaux,J.,DeMathelin, M., Padoy, N., 2016. Endonet: a deep architecture for recognition tasksonlaparoscopicvideos. IEEEtransactionsonmedicalimaging 36, 86–97

  74. [82]

    Attention is all you need

    Vaswani, A., 2017. Attention is all you need. Advances in Neural Information Processing Systems

  75. [83]

    Graph attention networks

    Veličković, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., Bengio, Y., 2017. Graph attention networks. arXiv preprint arXiv:1710.10903

  76. [84]

    Graph attention networks

    Velickovic, P., Cucurull, G., Casanova, A., Romero, A., Lio, P., Bengio, Y., et al., 2017. Graph attention networks. stat 1050, 10– 48550

  77. [85]

    Calibrationindeeplearning:Asurveyofthestate- of-the-art

    Wang,C.,2023. Calibrationindeeplearning:Asurveyofthestate- of-the-art. arXiv preprint arXiv:2308.01222

  78. [86]

    Magneticsoftmicrorobotdesignforcell grasping and transportation

    Wang, F., Zhang, Y., Jin, D., Jiang, Z., Liu, Y., Knoll, A., Jiang, H., Ying,Y.,Zhou,M.,2024a. Magneticsoftmicrorobotdesignforcell grasping and transportation. Cyborg and Bionic Systems 5, 0109

  79. [87]

    Surgical-lvlm:Learningtoadapt largevision-languagemodelforgroundedvisualquestionanswering in robotic surgery

    Wang, G., Bai, L., Nah, W.J., Wang, J., Zhang, Z., Chen, Z., Wu, J., Islam,M.,Liu,H.,Ren,H.,2024b. Surgical-lvlm:Learningtoadapt largevision-languagemodelforgroundedvisualquestionanswering in robotic surgery. arXiv preprint arXiv:2405.10948

  80. [88]

    Domain adaptivesim-to-realsegmentationoforopharyngealorgans.Medical & Biological Engineering & Computing 61, 2745–2755

    Wang, G., Ren, T.A., Lai, J., Bai, L., Ren, H., 2023. Domain adaptivesim-to-realsegmentationoforopharyngealorgans.Medical & Biological Engineering & Computing 61, 2745–2755. L. Bai et al.:Preprint submitted to Elsevier Page 19 of 20 Multimodal Graph Representation Learning for...

  81. [89]

    Copesd: A multi-level surgical motion dataset for training large vision-language models to co-pilot endoscopic submucosal dissection

    Wang,G.,Xiao,H.,Gao,H.,Zhang,R.,Bai,L.,Yang,X.,Li,Z.,Li, H., Ren, H., 2024c. Copesd: A multi-level surgical motion dataset for training large vision-language models to co-pilot endoscopic submucosal dissection. arXiv preprint arXiv:2410.07540

  82. [90]

    Video-instrumentsynergisticnetworkforreferringvideo instrument segmentation in robotic surgery

    Wang,H.,Yang,G.,Zhang,S.,Qin,J.,Guo,Y.,Xu,B.,Jin,Y.,Zhu, L.,2024d. Video-instrumentsynergisticnetworkforreferringvideo instrument segmentation in robotic surgery. IEEE Transactions on Medical Imaging

  83. [91]

    Gcl: Graph calibration loss for trustworthy graph neural network, in: Proceedings of the 30th ACM International Conference on Multimedia, pp

    Wang, M., Yang, H., Cheng, Q., 2022a. Gcl: Graph calibration loss for trustworthy graph neural network, in: Proceedings of the 30th ACM International Conference on Multimedia, pp. 988–996

  84. [92]

    Wang, M., Yang, H., Huang, J., Cheng, Q., 2024e. Moderate message passing improves calibration: A universal way to mitigate confidence bias in graph neural networks, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 21681–21689

  85. [93]

    IEEE Transactions on Multimedia

    Wang,Q.,Wei,Y.,Yin,J.,Wu,J.,Song,X.,Nie,L.,2021a.Dualgnn: Dual graph neural network for multimedia recommendation. IEEE Transactions on Multimedia

  86. [94]

    Be confident! towards trustworthy graph neural networks via confidence calibration

    Wang, X., Liu, H., Shi, C., Yang, C., 2021b. Be confident! towards trustworthy graph neural networks via confidence calibration. Ad- vancesinNeuralInformationProcessingSystems34,23768–23779

  87. [95]

    Adversarial multimodal fusion with attention mechanism for skin lesion classification using clinical and dermo- scopic images

    Wang, Y., Feng, Y., Zhang, L., Zhou, J.T., Liu, Y., Goh, R.S.M., Zhen, L., 2022b. Adversarial multimodal fusion with attention mechanism for skin lesion classification using clinical and dermo- scopic images. Medical Image Analysis 81, 102535

  88. [96]

    Exploring het- erophily in calibration of graph neural networks

    Xie, X., Luo, B., Li, Y., Yang, C., Gui, W., 2024. Exploring het- erophily in calibration of graph neural networks. Neurocomputing 604, 128294

  89. [97]

    A fourier- based framework for domain generalization, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

    Xu, Q., Zhang, R., Zhang, Y., Wang, Y., Tian, Q., 2021. A fourier- based framework for domain generalization, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14383–14392

  90. [98]

    Multimodal semi-supervisedlearningforonlinerecognitionofmulti-granularity surgical workflows

    Yamada, Y., Colan, J., Davila, A., Hasegawa, Y., 2024. Multimodal semi-supervisedlearningforonlinerecognitionofmulti-granularity surgical workflows. International Journal of Computer Assisted Radiology and Surgery 19, 1075–1083

  91. [99]

    Calibrating graph neural networks from a data-centric perspective, in:ProceedingsoftheACMonWebConference2024,pp.745–755

    Yang, C., Yang, C., Shi, C., Li, Y., Zhang, Z., Zhou, J., 2024a. Calibrating graph neural networks from a data-centric perspective, in:ProceedingsoftheACMonWebConference2024,pp.745–755

  92. [100]

    Yang,S.,Luo,L.,Wang,Q.,Chen,H.,2024b. Surgformer:Surgical transformer with hierarchical temporal attention for surgical phase recognition, in: International Conference on Medical Image Com- puting and Computer-Assisted Intervention, Springer. pp. 606–616

  93. [101]

    Fda: Fourier domain adaptation for se- mantic segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

    Yang, Y., Soatto, S., 2020. Fda: Fourier domain adaptation for se- mantic segmentation, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4085–4095

  94. [102]

    Anticipation for surgical workflow through instrument interaction and recognized signals

    Yuan, K., Holden, M., Gao, S., Lee, W., 2022. Anticipation for surgical workflow through instrument interaction and recognized signals. Medical Image Analysis 82, 102611

  95. [103]

    Yuan, K., Holden, M., Gao, S., Lee, W.S., 2021. Surgical work- flow anticipation using instrument interaction, in: Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27– October 1, 2021, Proceedi...

  96. [104]

    Yuan, K., Srivastav, V., Navab, N., Padoy, N., 2024. Hecvl: Hierar- chicalvideo-languagepretrainingforzero-shotsurgicalphaserecog- nition, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 306–316

  97. [105]

    Learning multi-modal representations by watching hundreds of surgical video lectures

    Yuan,K.,Srivastav,V.,Yu,T.,Lavanchy,J.L.,Mascagni,P.,Navab, N., Padoy, N., 2023. Learning multi-modal representations by watching hundreds of surgical video lectures. arXiv preprint arXiv:2307.15220

  98. [106]

    Multimodal in- formationfusionapproachfornoncontactheartrateestimationusing facial videos and graph convolutional network

    Yue, Z., Ding, S., Yang, S., Wang, L., Li, Y., 2021. Multimodal in- formationfusionapproachfornoncontactheartrateestimationusing facial videos and graph convolutional network. IEEE Transactions on Instrumentation and Measurement 71, 1–13

  99. [107]

    Zadrozny, B., Elkan, C., 2002. Transforming classifier scores into accurate multiclass probability estimates, in: Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 694–699

  100. [108]

    Multimodal interaction and fused graph convolution network for sentiment clas- sification of online reviews

    Zeng, D., Chen, X., Song, Z., Xue, Y., Cai, Q., 2023. Multimodal interaction and fused graph convolution network for sentiment clas- sification of online reviews. Mathematics 11, 2335

  101. [109]

    Dbfft: Adversarial- robust dual-branch frequency domain feature fusion in vision trans- formers

    Zeng, J., Huang, L., Bai, X., Wang, K., 2024. Dbfft: Adversarial- robust dual-branch frequency domain feature fusion in vision trans- formers. Information Fusion 108, 102387

  102. [110]

    Graph structure enhanced pre-training language model for knowledge graph completion

    Zhu, H., Xu, D., Huang, Y., Jin, Z., Ding, W., Tong, J., Chong, G., 2024. Graph structure enhanced pre-training language model for knowledge graph completion. IEEE Transactions on Emerging Topics in Computational Intelligence

  103. [111]

    Dyadic relational graph convolutional networks for skeleton-based human interaction recognition

    Zhu, L., Wan, B., Li, C., Tian, G., Hou, Y., Yuan, K., 2021. Dyadic relational graph convolutional networks for skeleton-based human interaction recognition. Pattern Recognition 115, 107920

  104. [112]

    Brain tumorsegmentationbasedonthefusionofdeepsemanticsandedge information in multimodal mri

    Zhu, Z., He, X., Qi, G., Li, Y., Cong, B., Liu, Y., 2023. Brain tumorsegmentationbasedonthefusionofdeepsemanticsandedge information in multimodal mri. Information Fusion 91, 376–387. L. Bai et al.:Preprint submitted to Elsevier Page 20 of 20

  105. [2016]

    Internationaljournalofcomputerassisted radiology and surgery 11, 1081–1089

    Automatic data-driven real-time segmentation and recogni- tionofsurgicalworkflow. Internationaljournalofcomputerassisted radiology and surgery 11, 1081–1089

  106. [2017]

    arXiv preprint arXiv:1701.06548

    Regularizing neural networks by penalizing confident output distributions. arXiv preprint arXiv:1701.06548

  107. [2021]

    International Journal of Surgery 95, 106151

    Asystematicreviewonartificialintelligenceinrobot-assisted surgery. International Journal of Surgery 95, 106151

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.