REVIEW 4 major objections 5 minor 71 references
Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper argues that a counterfactual score—total effect minus the visual modality's direct effect—removes image-matching shortcut bias in multi-modal entity alignment and beats 14 prior methods on nine benchmarks.
desk verdict A useful inference-time reweighting trick for MMEA, wrapped in a causal story the implementation doesn't actually support; engage with it for the empirical results, not for the causal claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the causal graph over $V$ (visual input), $G$ (graph input), $M$ (fused modality, a mediator), and $Y$ (prediction), together with the counterfactual subtraction $TIE = Y_{v,g,m} - \beta Y_{v,g^*,m^*}$. Here $Y_{v,g^*,m^*}$ is obtained by setting the graph and fused similarity scores to zero while keeping the visual branch's learned weights, implementing the counterfactual world in which the visual direct path is the only contributor; subtracting a fraction $\beta=0.2$ of it from the factual prediction is meant to cancel the visual shortcut while preserving the indirect path $V \to M \to Y$. The framework is implemented with per-modality encoders (VGG/ResNet for images, a relational reflection graph attention network for graph structure and for fusion) and an attentive weighted sum over the three score branches.
What would settle it
Compare CDMEA's debiased score with the counterfactual graph and fused scores set to zero against a version that sets them to random baseline values. If the ranking gains depend on which baseline is used, the improvement is tied to the zero convention rather than to a causal estimate. In the opposite direction, on a test set where equivalent entities already share highly similar images, the causal account predicts little or no gain from debiasing; a large gain there would indicate the method is adjusting score weights rather than removing a shortcut.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that visual-modality bias in multi-modal entity alignment can be diagnosed and removed causally. The authors define the factual prediction $Y_{v,g,m}$ from visual, graph, and fused scores, and a counterfactual prediction $Y_{v,g^*,m^*}$ in which the graph and fused similarity scores are set to zero while the visual branch stays active. The debiased score is $TIE = Y_{v,g,m} - \beta Y_{v,g^*,m^*}$, which they interpret as the Total Effect minus the Natural Direct Effect of the visual modality; the optimal $\beta$ is $0.2$. With this score, CDMEA reports the best Hits@1, Hits@10, and MRR on all nine benchmark settings, with the largest margins at low training ratios (e.g., FB-DB15K 20% H@1 $0.674$ vs. runner-up $0.631$).
Load-bearing premise
The load-bearing premise is that making the graph and fused-modality similarity scores zero in the imagined alternative world is a faithful way to switch those pathways off, and that nothing unobserved drives both the images and the predictions; if either assumption fails, the subtracted quantity is not really the visual modality's direct effect.
Editorial extensions
If this is right
- Existing MMEA models can be debiased at inference time simply by subtracting a scaled counterfactual visual score; the paper shows this improves MCLEA, DESAlign, and MEAformer without retraining them.
- The benefit should be largest when image similarity is low, image noise is high, or alignment seeds are scarce; the paper's experiments report gains concentrated in those regimes.
- Partial subtraction is necessary: $\beta=0.2$ beats both no subtraction ($\beta=0$) and full subtraction ($\beta=1$), so the visual modality still contributes useful signal through the fused path.
- The debiasing also speeds up training: CDMEA converges faster and reports lower training time than four strong baselines on the same hardware.
Reading between the lines
- Beyond the paper, the same total-effect-minus-direct-effect recipe could be ported to any multimodal retrieval task with a suspected shortcut modality, with a per-modality $\beta$ tuned on validation data.
- The reported $\beta$ sensitivity suggests $\beta$ could be made adaptive per entity or per image-similarity bin, subtracting more visual effect precisely where image similarity is low.
- The framework could be extended to more than two modalities by building a multi-level causal graph and subtracting each suspect modality's direct effect in turn.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CDMEA, a framework for multi-modal entity alignment (MMEA) that aims to reduce visual modality bias. The authors set up a causal graph with visual (V), graph (G), fused (M), and prediction (Y) nodes, define TE, NDE, and TIE, and propose a debiasing inference that computes TIE = Y_{v,g,m} - beta * Y_{v,g*,m*}. The implementation consists of three encoders (visual, graph RRGAT, fused RRGAT), a weighted fusion of per-modality similarity scores, InfoNCE training, and a counterfactual inference that zeroes the graph and fused scores. The paper reports state-of-the-art results on nine benchmarks, with the largest gains under low-resource, noisy, and low-similarity-image settings.
Significance. If the causal interpretation held, the paper would contribute a principled and general debiasing principle for MMEA, with a reusable implementation and extensive empirical validation. The strengths are the breadth of the evaluation (9 benchmarks, 14 baselines), the module-ablation and robustness studies, the code release, and the observation that the visual modality can hurt performance. However, the central causal quantity reduces algebraically to a linear reweighting of the same three scores, so the reported gains currently demonstrate a tuned weighted fusion rather than a causal debiasing effect. The empirical contribution may still be useful, but the title, abstract, and framing over-claim what is established.
major comments (4)
- [3.2 / 3.3.4 / 3.4.2 (Eqs. 20, 18, 23)] The counterfactual branch as implemented collapses to alpha_v * Y_v. With the attentive fusion in Eq. (20), Y_{v,g,m} = alpha_v Y_v + alpha_g Y_g + alpha_m Y_m, and Eq. (18) sets Y_{g*}=Y_{m*}=0, so Y_{v,g*,m*} = alpha_v Y_v. Substituting into Eq. (23) gives TIE = (1-beta)*alpha_v*Y_v + alpha_g*Y_g + alpha_m*Y_m. This is a one-parameter reweighting of the existing scores; the NDE/TIE vocabulary and the causal graph impose no additional constraint. The claim that the model predicts based on the Total Indirect Effect is therefore not supported by the implementation. Either implement a genuine counterfactual (for example, recomputing M(V=v*,G=g*) and the fusion weights in the blocked world) or reframe the method as a heuristic reweighted fusion.
- [3.2 (Eqs. 7-10)] Setting a blocked modality score to zero is not Pearl's do-operator and does not correspond to a no-treatment counterfactual in the stated graph. In the graph of Fig. 4(a), G is not a descendant of V, yet the NDE computation zeroes both G and M; and the 'blocked' direct effect is only damped by beta=0.2, not blocked. The paper should justify the zero-imputation as an intervention or remove the causal-effect terminology.
- [4.1.3 / 4.4.5 (beta tuning)] The hyperparameter beta is grid-searched from 0.0 to 0.9 on H@1 (Section 4.1.3, Fig. 9), so the final debiasing strength is selected on the target metric. This makes the TIE prediction a calibrated interpolation rather than a parameter-free causal estimate. At minimum, report a validation-set selection or a sensitivity analysis with beta chosen without test feedback.
- [4.2 (Tables 1-2)] All reported numbers are single runs without error bars. On the bilingual benchmarks the gains over IBMEA are 0.6-1.3% H@1, which is within the typical run-to-run variance of such models. Without multiple seeds or significance testing, the claim of consistent state-of-the-art performance on all nine benchmarks is not firmly established.
minor comments (5)
- [Section 1 and Section 2.2.2] 'generalCasual Debiasing framework' should be 'general Causal Debiasing framework', and the Section 2.2.2 heading 'Casual Effects' should be 'Causal Effects'.
- [Eq. (13)] The attention denominator is malformed ('exp(q^T h_rk) ... exp(q^T h_rk'))); please fix the formula and the summation notation.
- [Eq. (14)] The orthogonality proof should explicitly use the stated normalization ||h_rk||=1 inside the displayed equation; as written, the second equality hides this assumption.
- [Table 4 and Section 4.2] The Table 4 header 'FB-YG15K ((50%)' has a double parenthesis, and the improvement percentages in Section 4.2 should state whether they are absolute or relative.
- [Section 4.1.3 and Section 4.4.4] Entities without images are assigned random visual vectors; given the paper's focus on visual reliability, please ablate this choice or discuss its effect on the low-similarity improvements. Also, 'the results improve as image similarity decreases' in Section 4.4.4 is ambiguous: presumably the relative advantage over baselines grows as similarity decreases, not the absolute H@1.
Circularity Check
The TIE scoring rule reduces by construction to a reweighted fusion of the three modality scores; the causal NDE estimate is the visual branch's own contribution, so the central debiasing 'prediction' is equivalent to tuning a visual downweighting.
-
self definitional
[Section 3.2 Eq. (10); Section 3.3.4 Eq. (18); Section 3.3.5 Eq. (20); Section 3.4.2 Eq. (23)]
"𝑇𝐼𝐸 =𝑌𝑣,𝑔,𝑚−𝛽·𝑌𝑣,𝑔∗,𝑚∗ ... we set the 𝑌𝑣∗,𝑌𝑔∗, and𝑌𝑚∗ as zero, since if the corresponding input modality is blocked, indicating no meaningful similarity ... 𝑌𝑣,𝑔,𝑚 =∑_{𝑘∈𝑣,𝑔,𝑚} exp(𝜙𝑘)˝_{𝑘′∈𝑣,𝑔,𝑚} exp(𝜙𝑘′)𝑌𝑘"
Under the attentive fusion of Eq. (20), the counterfactual score Y_{v,g*,m*} is obtained by fusing (Y_v, 0, 0), which equals alpha_v Y_v because the softmax weights remain unchanged while Y_g* and Y_m* are zeroed. Substituting into Eq. (23) gives TIE = (1-beta) alpha_v Y_v + alpha_g Y_g + alpha_m Y_m, which is exactly the original fused score with the visual coefficient rescaled by (1-beta). Thus the 'Natural Direct Effect of the visual modality' is defined to be the visual branch's own contribution to the fused score, and subtracting it is merely reweighting the same three input scores.
full rationale
The central derivation chain contains one load-bearing reduction-by-construction. Because the factual and counterfactual fusions share the same learned attention weights, setting Y_g*=Y_m*=0 makes Y_{v,g*,m*} = alpha_v Y_v; therefore the TIE of Eq. (23) is algebraically identical to the original attentive fusion with the visual branch rescaled by (1-beta). The causal machinery is not needed to obtain this score, and the 'blocked' counterfactual world is an implementation-level zeroing of two scores rather than Pearl's do-operator. Since beta is grid-searched on H@1, the reported 'debiased prediction' is a calibrated interpolation of the same three modality scores rather than an independently estimated causal effect. This makes the causal-debiasing claim partially circular: the quantity advertised as a causally derived prediction is, by construction, the original score with one tuned scalar weight adjustment. The empirical evaluation against external baselines, ablations, and nine benchmarks remains self-contained and externally meaningful, so the paper is not a pure restatement or a self-citation artifact. The self-citations present (e.g., IBMEA [37] and LoginMEA [38]) are used only as baselines or related work and are not load-bearing for the central derivation.
Assumptions & free parameters
free parameters (2)
- beta (debiasing coefficient) =
0.2
- modality attention weights phi_v, phi_g, phi_m =
learned, not reported
assumptions (3)
- domain assumption The causal graph V->M->Y, G->M->Y, V->Y, G->Y has no unobserved confounders and is correctly specified.
- ad hoc to paper Setting a blocked modality's similarity score to zero is a valid no-treatment counterfactual baseline.
- standard math Relation reflection weights W_{rk}=I-2 h_{rk} h_{rk}^T preserve norm because ||h_{rk}||=1.
Cite this review
Pith. "Pith review of Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective." pith.science (2026). https://pith.science/paper/AZOBVRH2
@misc{pith2026250419458,
author = {Pith},
title = {Pith review of: Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/AZOBVRH2}},
note = {Machine review of arXiv:2504.19458}
}
read the original abstract
Multi-Modal Entity Alignment (MMEA) aims to retrieve equivalent entities from different Multi-Modal Knowledge Graphs (MMKGs), a critical information retrieval task. Existing studies have explored various fusion paradigms and consistency constraints to improve the alignment of equivalent entities, while overlooking that the visual modality may not always contribute positively. Empirically, entities with low-similarity images usually generate unsatisfactory performance, highlighting the limitation of overly relying on visual features. We believe the model can be biased toward the visual modality, leading to a shortcut image-matching task. To address this, we propose a counterfactual debiasing framework for MMEA, termed CDMEA, which investigates visual modality bias from a causal perspective. Our approach aims to leverage both visual and graph modalities to enhance MMEA while suppressing the direct causal effect of the visual modality on model predictions. By estimating the Total Effect (TE) of both modalities and excluding the Natural Direct Effect (NDE) of the visual modality, we ensure that the model predicts based on the Total Indirect Effect (TIE), effectively utilizing both modalities and reducing visual modality bias. Extensive experiments on 9 benchmark datasets show that CDMEA outperforms 14 state-of-the-art methods, especially in low-similarity, high-noise, and low-resource data scenarios.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Antoine Bordes, Nicolas Usunier, Alberto García-Durán, Jason Weston, and Ok- sana Yakhnenko. 2013. Translating Embeddings for Modeling Multi-relational Data. In Proceedings of NeurIPS. 2787–2795
work page 2013
-
[2]
Yixin Cao, Zhiyuan Liu, Chengjiang Li, Zhiyuan Liu, Juanzi Li, and Tat-Seng Chua. 2019. Multi-Channel Graph Neural Network for Entity Alignment. In Proceedings of ACL. 1452–1461
work page 2019
-
[3]
Catherine Chen, Jack Merullo, and Carsten Eickhoff. 2024. Axiomatic Causal In- terventions for Reverse Engineering Relevance Computation in Neural Retrieval Models. In Proceedings of SIGIR , Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, and Yi Zhang (Eds.). ACM, 1401–1410
work page 2024
-
[4]
Liyi Chen, Zhi Li, Yijun Wang, Tong Xu, Zhefeng Wang, and Enhong Chen. 2020. MMEA: entity alignment for multi-modal knowledge graph. In Proceedings of KSEM. Springer, 134–147
work page 2020
-
[5]
Liyi Chen, Zhi Li, Tong Xu, Han Wu, Zhefeng Wang, Nicholas Jing Yuan, and Enhong Chen. 2022. Multi-modal siamese network for entity alignment. In Proceedings of KDD. 118–126
work page 2022
-
[6]
Muhao Chen, Yingtao Tian, Kai-Wei Chang, Steven Skiena, and Carlo Zaniolo
-
[7]
Muhao Chen, Yingtao Tian, Mohan Yang, and Carlo Zaniolo. 2017. Multilin- gual Knowledge Graph Embeddings for Cross-lingual Knowledge Alignment. In Proceedings of IJCAI. 1511–1517
work page 2017
-
[8]
Zhuo Chen, Jiaoyan Chen, Wen Zhang, Lingbing Guo, Yin Fang, Yufeng Huang, Yichi Zhang, Yuxia Geng, Jeff Z Pan, Wenting Song, et al. 2023. Meaformer: Multi- modal entity alignment transformer for meta modality hybrid. In Proceedings of ACM MM. 3317–3327
work page 2023
Show all 71 references
-
[9]
Pan, Yangning Li, Huajun Chen, and Wen Zhang
Zhuo Chen, Lingbing Guo, Yin Fang, Yichi Zhang, Jiaoyan Chen, Jeff Z. Pan, Yangning Li, Huajun Chen, and Wen Zhang. 2023. Rethinking Uncertainly Missing and Ambiguous Visual Modality in Multi-Modal Entity Alignment. In Proceedings of ISWC. 121–139
2023
-
[10]
Congcong Ge, Xiaoze Liu, Lu Chen, Baihua Zheng, and Yunjun Gao. 2021. Make It Easy: An Effective End-to-End Entity Alignment Framework. In Proceedings of SIGIR. ACM, 777–786
2021
-
[11]
Zemel, Wieland Brendel, Matthias Bethge, and Felix A
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard S. Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020. Shortcut learning in deep neural networks. Nat. Mach. Intell. 2, 11 (2020), 665–673
2020
-
[12]
Hao Guo, Jiuyang Tang, Weixin Zeng, Xiang Zhao, and Li Liu. 2021. Multi-modal entity alignment in hyperbolic space. Neurocomputing 461 (2021), 598–607
2021
-
[13]
Lingbing Guo, Zhuo Chen, Jiaoyan Chen, Yin Fang, Wen Zhang, and Huajun Chen. 2024. Revisit and Outstrip Entity Alignment: A Perspective of Generative Models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenR...
2024
-
[14]
Muskan Gupta, Priyanka Gupta, Jyoti Narwariya, Lovekesh Vig, and Gautam Shroff. 2024. SCM4SR: Structural Causal Model-based Data Augmentation for Robust Session-based Recommendation. In Proceedings of SIGIR, Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, ...
2024
-
[15]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In Proceedings of CVPR. 770–778
2016
-
[16]
Jiacheng Huang, Wei Hu, Haoxuan Li, and Yuzhong Qu. 2018. Automated Com- parative Table Generation for Facilitating Human Intervention in Multi-Entity Resolution. In The 41st International ACM SIGIR Conference on Research & Devel- opment in Information Retrieval, SIGIR 2018, A...
2018
-
[17]
Yani Huang, Xuefeng Zhang, Richong Zhang, Junfan Chen, and Jaein Kim. 2024. Progressively Modality Freezing for Multi-Modal Entity Alignment. In Proceed- ings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok...
2024
-
[18]
Kusner, and Ricardo Silva
Jean Kaddour, Aengus Lynch, Qi Liu, Matt J. Kusner, and Ricardo Silva. 2022. Causal Machine Learning: A Survey and Open Problems. CoRR abs/2206.15475 (2022)
2022 arXiv
-
[19]
Malawade, and Mohammad Abdullah Al Faruque
Amar Viswanathan Kannan, Dmitriy Fradkin, Ioannis Akrotirianakis, Tugba Kulahcioglu, Arquimedes Canedo, Aditi Roy, Shih-Yuan Yu, Arnav V. Malawade, and Mohammad Abdullah Al Faruque. 2020. Multimodal Knowledge Graph for Deep Learning Papers and Code. In CIKM ’20: The 29th ACM I...
2020
-
[20]
Kipf and Max Welling
Thomas N. Kipf and Max Welling. 2017. Semi-Supervised Classification with Graph Convolutional Networks. In Proceedings of ICLR
2017
-
[21]
Chengjiang Li, Yixin Cao, Lei Hou, Jiaxin Shi, Juanzi Li, and Tat-Seng Chua. 2019. Semi-supervised Entity Alignment via Joint Knowledge Embedding Model and Cross-graph Model. In Proceedings of EMNLP. 2723–2732
2019
-
[22]
Qian Li, Shu Guo, Yangyifei Luo, Cheng Ji, Lihong Wang, Jiawei Sheng, and Jianxin Li. 2023. Attribute-Consistent Knowledge Graph Representation Learning for Multi-Modal Entity Alignment. In Proceedings of WWW . 2499–2508
2023
-
[23]
Ran Li, Shimin Di, Lei Chen, and Xiaofang Zhou. 2024. SimDiff: Simple Denoising Probabilistic Latent Diffusion Model for Data Augmentation on Multi-modal Knowledge Graph. In Proceedings of KDD, Ricardo Baeza-Yates and Francesco Bonchi (Eds.). ACM, 1631–1642
2024
-
[24]
Zhenxi Lin, Ziheng Zhang, Meng Wang, Yinghui Shi, Xian Wu, and Yefeng Zheng
-
[25]
Fangyu Liu, Muhao Chen, Dan Roth, and Nigel Collier. 2021. Visual Pivoting for (Unsupervised) Entity Alignment. In Proceedings of AAAI. 4257–4266
2021
-
[26]
Rosenblum
Ye Liu, Hui Li, Alberto García-Durán, Mathias Niepert, Daniel Oñoro-Rubio, and David S. Rosenblum. 2019. MMKG: Multi-modal Knowledge Graphs. In Proceedings of ESWC, Vol. 11503. 459–474
2019
-
[27]
Zhiyuan Liu, Yixin Cao, Liangming Pan, Juanzi Li, and Tat-Seng Chua. 2020. Exploring and Evaluating Attributes, Values, and Structures for Entity Alignment. In Proceedings of EMNLP. 6355–6364
2020
-
[28]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In Proceedings of ICLR
2019
-
[29]
Junyu Lu, Bo Xu, Xiaokun Zhang, Kaiyuan Liu, Dongyu Zhang, Liang Yang, and Hongfei Lin. 2024. Take Its Essence, Discard Its Dross! Debiasing for Toxic Language Detection via Counterfactual Causal Effect. In Proceedings of COLING, Nicoletta Calzolari, Min-Yen Kan, Véronique Hos...
2024
-
[30]
Yangyifei Luo, Zhuo Chen, Lingbing Guo, Qian Li, Wenxuan Zeng, Zhixin Cai, and Jianxin Li. 2024. ASGEA: Exploiting Logic Rules from Align-Subgraphs for Entity Alignment. CoRR abs/2402.11000 (2024)
2024 arXiv
-
[31]
Xin Mao, Wenting Wang, Yuanbin Wu, and Man Lan. 2021. Boosting the Speed of Entity Alignment 10×: Dual Attention Matching Network with Normalized Hard Sample Mining. In Proceedings of WWW . 821–832
2021
-
[32]
Xin Mao, Wenting Wang, Huimin Xu, Yuanbin Wu, and Man Lan. 2020. Relational Reflection Entity Alignment. In Proceedings of CIKM. 1095–1104
2020
-
[33]
Wenxin Ni, Qianqian Xu, Yangbangyan Jiang, Zongsheng Cao, Xiaochun Cao, and Qingming Huang. 2023. PSNEA: Pseudo-Siamese Network for Entity Alignment between Multi-modal Knowledge Graphs. InProceedings of ACM MM. 3489–3497
2023
-
[34]
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji- Rong Wen. 2021. Counterfactual VQA: A Cause-Effect Look at Language Bias. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021. Computer Vision Foundation / ...
2021
-
[35]
Judea Pearl. 2022. Direct and Indirect Effects. InProbabilistic and Causal Inference: The Works of Judea Pearl, Hector Geffner, Rina Dechter, and Joseph Y. Halpern (Eds.). ACM Books, Vol. 36. ACM, 373–392
2022
-
[36]
Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Net- works for Large-Scale Image Recognition. In Proceedings of ICLR, Yoshua Bengio and Yann LeCun (Eds.)
2015
-
[37]
Taoyu Su, Jiawei Sheng, Shicheng Wang, Xinghua Zhang, Hongbo Xu, and Tingwen Liu. 2024. IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment. In Proceedings of ACM MM , Jianfei Cai, Mohan S. Kankanhalli, Balakrishnan Prabhakaran, Susanne Boll, R...
2024
-
[38]
Taoyu Su, Xinghua Zhang, Jiawei Sheng, Zhenyu Zhang, and Tingwen Liu
-
[39]
Pengzhan Sun, Bo Wu, Xunsong Li, Wen Li, Lixin Duan, and Chuang Gan. 2021. Counterfactual Debiasing Inference for Compositional Action Recognition. In MM ’21: ACM Multimedia Conference, Virtual Event, China, October 20 - 24, 2021 , Heng Tao Shen, Yueting Zhuang, John R. Smith,...
2021
-
[40]
Rui Sun, Xuezhi Cao, Yan Zhao, Junchen Wan, Kun Zhou, Fuzheng Zhang, Zhongyuan Wang, and Kai Zheng. 2020. Multi-modal Knowledge Graphs for Recommender Systems. In Proceedings of CIKM. 1405–1414
2020
-
[41]
Zequn Sun, Wei Hu, and Chengkai Li. 2017. Cross-Lingual Entity Alignment via Joint Attribute-Preserving Embedding. In Proceedings of ISWC. 628–644
2017
-
[42]
Zequn Sun, Wei Hu, Qingheng Zhang, and Yuzhong Qu. 2018. Bootstrapping Entity Alignment with Knowledge Graph Embedding. In Proceedings of IJCAI. 4396–4402
2018
-
[43]
Zequn Sun, Jiacheng Huang, Wei Hu, Muhao Chen, Lingbing Guo, and Yuzhong Qu. 2019. TransEdge: Translating Relation-Contextualized Embeddings for Knowledge Graphs. In Proceedings of ISWC. 612–629
2019
-
[44]
Zequn Sun, Chengming Wang, Wei Hu, Muhao Chen, Jian Dai, Wei Zhang, and Yuzhong Qu. 2020. Knowledge Graph Alignment Network with Gated Multi-Hop Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective SIGIR ’25, July 13–18, 2025, Padua, Italy Neighbo...
2020
-
[45]
Zhongxiang Sun, Jun Xu, Xiao Zhang, Zhenhua Dong, and Ji-Rong Wen. 2023. Law Article-Enhanced Legal Case Matching: A Causal Learning Approach. In Proceedings of SIGIR, Hsin-Hsi Chen, Wei-Jou (Edward) Duh, Hen-Hsen Huang, Makoto P. Kato, Josiane Mothe, and Barbara Poblete (Eds....
2023
-
[46]
Bayu Distiawan Trisedya, Jianzhong Qi, and Rui Zhang. 2019. Entity alignment between knowledge graphs using attribute embeddings. In Proceedings of AAAI, Vol. 33. 297–304
2019
-
[47]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2019. Representation Learning with Contrastive Predictive Coding. arXiv:1807.03748 [cs.LG]
2019 arXiv
-
[48]
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018. Graph Attention Networks. InProceedings of ICLR
2018
-
[49]
Jihu Wang, Yuliang Shi, Han Yu, Xinjun Wang, Zhongmin Yan, and Fanyu Kong
-
[50]
Luyao Wang, Pengnian Qi, Xigang Bao, Chunlai Zhou, and Biao Qin. 2024. Pseudo- Label Calibration Semi-supervised Multi-Modal Entity Alignment. In Proceedings of AAAI. 9116–9124
2024
-
[51]
Yuanyi Wang, Haifeng Sun, Jiabo Wang, Jingyu Wang, Wei Tang, Qi Qi, Shaoling Sun, and Jianxin Liao. 2024. Towards Semantic Consistency: Dirichlet Energy Driven Robust Multi-Modal Entity Alignment. In 40th IEEE International Confer- ence on Data Engineering, ICDE 2024, Utrecht,...
2024
-
[52]
Zhichun Wang, Qingsong Lv, Xiaohan Lan, and Yu Zhang. 2018. Cross-lingual Knowledge Graph Alignment via Graph Convolutional Networks. In Proceedings of EMNLP. 349–357
2018
-
[53]
Haokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei, Liqiang Nie, and Tat- Seng Chua. 2024. Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval. In Proceedings of SIGIR. ACM, 229–239
2024
-
[54]
Junyang Wu, Tianyi Li, Lu Chen, Yunjun Gao, and Ziheng Wei. 2023. SEA: A Scalable Entity Alignment System. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023, Taipei, Taiwan, July 23-27, 2023 , Hsin-Hs...
2023
-
[55]
Yuting Wu, Xiao Liu, Yansong Feng, Zheng Wang, Rui Yan, and Dongyan Zhao
-
[56]
Baogui Xu, Chengjin Xu, and Bing Su. 2023. Cross-Modal Graph Attention Network for Entity Alignment. In Proceedings of ACM MM . 3715–3723
2023
-
[57]
Guohai Xu, Hehong Chen, Feng-Lin Li, Fu Sun, Yunzhou Shi, Zhixiong Zeng, Wei Zhou, Zhongzhou Zhao, and Ji Zhang. 2021. Alime mkg: A multi-modal knowledge graph for live-streaming e-commerce. In Proceedings of CIKM. 4808– 4812
2021
-
[58]
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2024. Diffusion Models: A Comprehensive Survey of Methods and Applications. ACM Comput. Surv. 56, 4 (2024), 105:1–105:39
2024
-
[59]
Yawen Zeng, Qin Jin, Tengfei Bao, and Wenfeng Li. 2023. Multi-Modal Knowledge Hypergraph for Diverse Image Retrieval. In Proceedings of AAAI, Brian Williams, Yiling Chen, and Jennifer Neville (Eds.). AAAI Press, 3376–3383
2023
-
[60]
Qingheng Zhang, Zequn Sun, Wei Hu, Muhao Chen, Lingbing Guo, and Yuzhong Qu. 2019. Multi-view Knowledge Graph Embedding for Entity Alignment. In Proceedings of IJCAI. 5429–5435
2019
-
[61]
Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Binbin Hu, Ziqi Liu, Wen Zhang, and Huajun Chen. 2024. NativE: Multi-modal Knowledge Graph Comple- tion in the Wild. In Proceedings of SIGIR, Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, and Yi Zhang (Eds...
2024
-
[62]
Yingying Zhang, Shengsheng Qian, Quan Fang, and Changsheng Xu. 2019. Multi- modal Knowledge-aware Hierarchical Attention Network for Explainable Medical Question Answering. In Proceedings of ACM MM . 1089–1097
2019
-
[63]
Yu Zhao, Ying Zhang, Baohang Zhou, Xinying Qian, Kehui Song, and Xiangrui Cai. 2024. Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph Completion. In Proceedings of SIGIR , Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, G...
2024
-
[64]
Hao Zhu, Ruobing Xie, Zhiyuan Liu, and Maosong Sun. 2017. Iterative Entity Alignment via Joint Knowledge Embeddings. In Proceedings of IJCAI. 4258–4264
2017
-
[65]
Qiannan Zhu, Xiaofei Zhou, Jia Wu, Jianlong Tan, and Li Guo. 2019. Neighborhood-Aware Attentional Representation for Multilingual Knowledge Graphs. In Proceedings of IJCAI. 1943–1949
2019
-
[66]
Xiangru Zhu, Zhixu Li, Xiaodan Wang, Xueyao Jiang, Penglei Sun, Xuwu Wang, Yanghua Xiao, and Nicholas Jing Yuan. 2024. Multi-Modal Knowledge Graph Construction and Application: A Survey. IEEE Trans. Knowl. Data Eng. 36, 2 (2024), 715–735
2024
-
[2018]
In Proceedings of IJCAI
Co-training Embeddings of Knowledge Graphs and Entity Descriptions for Cross-lingual Entity Alignment. In Proceedings of IJCAI. 3998–4004
-
[2019]
In Proceedings of IJCAI, Sarit Kraus (Ed.)
Relation-Aware Entity Alignment for Heterogeneous Knowledge Graphs. In Proceedings of IJCAI, Sarit Kraus (Ed.). 5278–5284
-
[2022]
In Proceedings of COLING
Multi-modal Contrastive Representation Learning for Entity Alignment. In Proceedings of COLING. 2572–2584
-
[2023]
In Proceedings of SIGIR, Hsin-Hsi Chen, Wei-Jou (Ed- ward) Duh, Hen-Hsen Huang, Makoto P
Mixed-Curvature Manifolds Interaction Learning for Knowledge Graph- aware Recommendation. In Proceedings of SIGIR, Hsin-Hsi Chen, Wei-Jou (Ed- ward) Duh, Hen-Hsen Huang, Makoto P. Kato, Josiane Mothe, and Barbara Poblete (Eds.). ACM, 372–382
-
[2024]
LoginMEA: Local-to-Global Interaction Network for Multi-Modal Entity Alignment. In ECAI 2024 - 27th European Conference on Artificial Intelligence, 19-24 October 2024, Santiago de Compostela, Spain - Including 13th Conference on Prestigious Applications of Intelligent Systems ...
2024
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.