Pith. sign in

REVIEW 3 major objections 5 minor 59 references

Treating time as an entity-level modality in multi-modal knowledge graphs can disambiguate entities whose text and images look alike, with large gains on the hardest cases.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 15:51 UTC pith:BQ3ADCRK

load-bearing objection Solid MMKG engineering paper: entity-level time as a modality with real ablations; the 58% hard-subset number is the softest claim, not the whole case. the 3 major comments →

arxiv 2607.09777 v1 pith:BQ3ADCRK submitted 2026-07-08 cs.CV

Time Imprint: Learning Time-Aware Representations in Multi-Modal Knowledge Graphs

classification cs.CV
keywords Multi-modal Knowledge GraphsTemporal Representation LearningContrastive LearningModality FusionLink PredictionTimestamp SelectionAttention Pooling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Multi-modal knowledge graphs attach text and images to entities, but many entities still share nearly identical descriptions and pictures, so models cannot tell them apart. This paper argues that an entity's timestamps can act as a third modality that breaks those ties, if time is encoded, fused, and aligned with text and images rather than treated only as a label on a triple. The authors introduce Time Imprint: it selects a compact set of years per entity (median window plus earliest-year anchor), pools them with attention into one temporal embedding, injects that embedding at encoding and scoring time, and pulls the temporal, textual, and visual views of the same entity together with a multi-view contrastive loss. On three standard benchmarks the method reaches state-of-the-art link-prediction scores, with the largest lifts on the top-1% most confusable entities. The work therefore claims that time-as-modality is most useful precisely when visual and textual cues alone fail, provided timestamps are selected and pooled carefully.

Core claim

Time Imprint shows that modeling time as an entity-level modality, jointly aligned with text and images through a three-view contrastive objective and a compact multi-timestamp attention pool, produces more discriminative multi-modal entity representations and yields state-of-the-art link prediction, especially on entities whose text and image features are highly similar.

What carries the argument

Time Imprint: year-level timestamps are subset-selected (median-K with earliest-year anchor), attention-pooled into a temporal vector, injected as a prefix token and gated into scoring, and aligned with visual and textual entity views via multi-view InfoNCE contrastive loss.

Load-bearing premise

Year-level timestamps scraped from entity descriptions and image metadata are accurate enough and representative enough that a median-plus-anchor selection plus attention pooling yields a clean entity signal rather than noise.

What would settle it

On the same three benchmarks, replace the extracted years with random years or drop them for the top-1% ambiguity entities; if Hits@1 gains of the claimed magnitude disappear while other modalities stay fixed, the time-as-modality claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes Time Imprint, a multi-modal KG framework that treats entity timestamps as a first-class entity-level modality alongside text and images. It encodes years with a sinusoidal temporal encoder, selects a compact median-centered subset of timestamps (with an earliest-year anchor), aggregates them by cosine-attention pooling, injects the resulting temporal embedding at three stages (prefix token in a cross-modal Transformer, gated scoring/relation modulation, and multi-view contrastive alignment), and optimizes a joint link-prediction plus InfoNCE objective. On DB15K, MKG-W, and MKG-Y the method reports competitive or best link-prediction metrics, with the largest claimed gains on an author-defined top-1% multi-modal ambiguity subset (up to +58% Hits@1) and supporting ablations over selection strategy, K, aggregation, and timestamp noise.

Significance. If the results hold under tighter controls, the work is a clear contribution to multi-modal KG representation learning: it reframes time from a triple-level label to an alignable entity modality, systematically maps a multi-timestamp design space, and shows that temporal cues help most when text and image features are confusable. Strengths include thorough ablations of the three injection stages (Table 3), inverted-U analysis of K, selection/aggregation sweeps, noise robustness curves, and a public code release. These make the design choices falsifiable and reusable even if headline numbers are moderated.

major comments (3)
  1. Section 4.5 and Table 2: the top-1% ambiguity subset is defined by highest average nearest-neighbor cosine similarity in the same pretrained text/image embedding spaces whose source documents (DBpedia descriptions and image metadata) also supply the rule-extracted years (Section 4.2). This construction can preferentially retain Napoleon-style cases whose year ranges already separate entities, so the +58.21% Hits@1 figure does not isolate the pure contribution of time-as-modality. Please report (i) the fraction of subset pairs that are already year-disjoint, (ii) a control subset of high text-image similarity that remains temporally overlapping, and/or (iii) gains after ablating temporal features only on that control. Without this, the central disambiguation claim is overstated relative to overall Table 1 gains.
  2. Table 1: the SOTA claim is uneven and unsupported by variance. On DB15K, Time Imprint trails MOMOK on Hits@1 (31.76 vs 32.38) and Hits@3/H@10; on MKG-Y Hits@10 trails SNAG. All numbers are single-run point estimates with no error bars, seeds, or significance tests. Given free parameters (K, tau_a, lambda, gamma, alpha) and modest absolute gains outside MKG-W, either report multi-seed means/stds or qualify the abstract/conclusion language to “best or second-best on most metrics, with largest gains on MKG-W and the ambiguity slice.”
  3. Section 3 (year-granularity paragraph) and Section 4.2: the claim that finer than year resolution helps <1% of pairs rests on a boundary-case count and a 100-entity manual audit (89% with at least one correct year). The audit does not report precision/recall of extracted years, multi-year noise rates, or whether incorrect years systematically bias median-K selection. Because the weakest modeling assumption is timestamp quality/representativeness, please expand the audit (error types, per-dataset rates) and, if possible, show performance when only high-confidence years are kept versus the current full extraction pipeline.
minor comments (5)
  1. Abstract and introduction advertise a “three-view contrastive objective,” but Section 3.4 defines five views C(e). Align the wording.
  2. Placeholder metadata remains throughout (Conference acronym ’XX, Woodstock, NY; 2018 copyright/ACM Reference Format). Clean for camera-ready.
  3. Figure 4/5 axis labels appear as garbled Unicode glyph sequences in the manuscript PDF text; ensure vector fonts render correctly.
  4. Equation (12): clarify whether the temporal view’s extra weight gamma is applied inside the sum or as a multiplier on pairs involving the temporal view; the prose mentions gamma but the displayed formula does not show it.
  5. Related Work 2.2: a short explicit comparison to literal-time MMKG methods (e.g., Wilcke et al.) would help readers see what is new beyond “time as features.”

Circularity Check

0 steps flagged

No significant circularity; empirical architecture + standard filtered link-prediction evaluation with open ablations, not a derivation that reduces to its inputs.

full rationale

Time Imprint is an empirical multi-modal KG completion method. Its claimed contributions (entity-level temporal modality, median-K + earliest-anchor selection, attention pooling, three-view contrastive alignment, gated temporal injection into Tucker scoring) are architectural choices whose parameters are learned end-to-end from the ordinary link-prediction loss plus a standard InfoNCE term; none of the equations define a quantity from the evaluation metric or from a fitted target and then re-present it as a prediction. Table 1 reports ordinary filtered MRR/Hits against public baselines on the full test sets; Table 2 is a diagnostic hard-slice defined solely by pretrained text/image nearest-neighbor cosine similarity (independent of the model’s temporal embeddings at selection time). Timestamp extraction is a fixed preprocessing step whose quality is audited and ablated (corruption curves, K-sensitivity, selection strategies), not a circular fit. There are no uniqueness theorems, self-citation load-bearing premises, or renamed known results that force the central claim. Minor experimental-design caveats (shared source documents for text and years) affect claim strength, not circularity of any derivation chain. Score 0 is therefore the correct, proportionate finding.

Axiom & Free-Parameter Ledger

5 free parameters · 4 axioms · 3 invented entities

The central empirical claim rests on a handful of architectural free parameters (K, temperatures, loss weights) chosen by validation sweeps, on the domain assumption that year-level extracted timestamps are valid entity-level signals, and on several invented architectural modules (time prefix token, relation-aware gate, five-view contrastive set) that have no independent existence outside this paper. No deep mathematical axioms are required; the load-bearing content is design choices plus data-quality assumptions.

free parameters (5)
  • K (number of selected timestamps) = 4–6 (default 6)
    Chosen by sweep; performance peaks at K=4–6 and is used as default. Directly controls specificity–robustness trade-off claimed in the abstract.
  • attention temperature tau_a = 0.5–0.6
    Controls sharpness of timestamp attention; paper reports 0.5–0.6 as stable. Affects how outliers are suppressed.
  • contrastive loss weight lambda and temporal view weight gamma
    Balance link-prediction versus multi-view alignment; values not exhaustively reported but required for the three-view objective to function as claimed.
  • gate scale alpha and zero-init projections
    Control strength of temporal injection into scoring; zero-init is a training-stability choice that affects whether time is used early.
  • embedding dimension / epochs / batch size = 256 / 1500 / 1024
    Standard training hyper-parameters fixed at 256 / 1500 / 1024; results depend on them remaining in a workable regime.
axioms (4)
  • domain assumption Year-level timestamps extracted from entity text and image metadata constitute a valid entity-level modality comparable to text and images.
    Stated in §3 and §4.2; finer granularity is argued to help <1% of pairs. If extraction is systematically biased, the modality is noise.
  • domain assumption Sinusoidal year encoding with learnable frequencies produces smoothly varying temporal features suitable for attention pooling and Transformer fusion.
    Eq. (1) and surrounding text; borrowed from temporal KG practice but not re-proved here.
  • ad hoc to paper Median-centered selection plus earliest-year anchor plus cosine-attention pooling balances specificity and robustness better than earliest/latest/random alternatives.
    Design choice validated only by the paper’s own sweeps (Fig. 4–5, Table 4–5).
  • ad hoc to paper The top-1% entities ranked by average nearest-neighbor cosine similarity in pretrained text/image space are the correct operationalization of multi-modal ambiguity.
    Defines the subset that produces the headline 58% gain (Table 2); alternative ambiguity definitions could shrink the reported benefit.
invented entities (3)
  • Temporal prefix token t_e (zero-init projected time embedding prepended after [ENT]) no independent evidence
    purpose: Inject sparse temporal signal into the cross-modal Transformer so self-attention can condition entity representations on time.
    Architectural invention of §3.2; no external evidence beyond ablation (ii).
  • Relation-aware temporal gate and temporal relation modulation in scoring no independent evidence
    purpose: Allow the model to decide which relations benefit from time and to shift relation embeddings by the known entity’s time.
    Eqs. (8)–(9); validated only by ablation (iii) inside this paper.
  • Five-view multi-view contrastive set C(e) that includes the aggregated temporal embedding as a first-class view no independent evidence
    purpose: Force sparse temporal embeddings into the same space as denser visual/textual views, addressing sparse temporal semantics.
    §3.4; the temporal view and its gamma weight are paper-specific.

pith-pipeline@v1.1.0-grok45 · 29722 in / 3508 out tokens · 37049 ms · 2026-07-14T15:51:36.451566+00:00 · methodology

0 comments
read the original abstract

Multi-Modal Knowledge Graphs (MMKGs) enrich entities with multiple modalities such as text and images, yet entities with highly similar multi-modal features remain difficult to distinguish. Temporal information of an entity can serve as an additional modality to disambiguate such entities, but existing approaches rarely treat time as a separate modality alongside text and images due to two major challenges: (1) sparse temporal semantics, which hinder alignment with richer modalities, and (2) multiple timestamps, which introduce noise or reduce robustness in representation learning. To address these challenges, we propose Time Imprint, a framework that treats time as an entity-level modality and jointly aligns temporal, textual, and visual representations via a three-view contrastive objective. Additionally, to mitigate multi-timestamp ambiguity, Time Imprint studies a compact timestamp subset selection design space and aggregates the selected timestamps into a discriminative temporal embedding with attention pooling, balancing temporal specificity and robustness. Experiments on three MMKG benchmarks demonstrate that Time Imprint achieves state-of-the-art link prediction performance, improving Hits@1 by up to 6.07\% overall and yielding up to 58\% gains on the subset of the top-1\% ambiguity samples. We further examine different fusion strategies and the sensitivity to timestamp availability and quality, clarifying when and why time-as-modality is most beneficial, while adding only modest training overhead. We release our code at https://anonymous.4open.science/r/Time-Imprint.

Figures

Figures reproduced from arXiv: 2607.09777 by Congfeng Cao, Jia-Hong Huang, Klim Zaporojets, Paul Groth, Pengyu Zhang.

Figure 1
Figure 1. Figure 1: “Napoleon Bonaparte” (left) and “Napoleon (2023 film)” (right) share similar images and textual descriptions, which can mislead models. In contrast, their temporal in￾formation is different (1769-1821 vs. 2020-2024); therefore, incorporating time as a modality helps disambiguate them. entity-level representations in MMKGs [32], yielding more infor￾mative and semantically grounded features. This mismatch, w… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of Time Imprint. (1) Modality encoders extract visual and textual tokens and aggregate selected timestamps [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Illustration of timestamp extraction and selection [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Effect of timestamp selection strategy ( [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Effect of the number of selected timestamps [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

59 extracted references · 7 linked inside Pith

  1. [1]

    Ivana Balazevic, Carl Allen, and Timothy Hospedales. 2019. TuckER: Tensor Fac- torization for Knowledge Graph Completion. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (E...

  2. [2]

    Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Ok- sana Yakhnenko. 2013. Translating Embeddings for Modeling Multi-relational Data. InAdvances in Neural Information Processing Systems, C.J. Burges, L. Bot- tou, M. Welling, Z. Ghahramani, and K.Q. Weinberger (Eds.), Vol. 26. Curran Associates, Inc

  3. [3]

    Li Cai, Xin Mao, Yuhao Zhou, Zhaoguang Long, Changxu Wu, and Man Lan

  4. [4]

    arXiv:2403.04782 [cs.CL]

    A Survey on Temporal Knowledge Graph: Representation Learning and Applications. arXiv:2403.04782 [cs.CL]

  5. [5]

    Zongsheng Cao, Qianqian Xu, Zhiyong Yang, Yuan He, Xiaochun Cao, and Qingming Huang. 2022. OTKGE: Multi-modal Knowledge Graph Embeddings via Optimal Transport. InAdvances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Associates, Inc., 39090–39102

  6. [6]

    Pan, Hua- jun Chen, and Wen Zhang

    Zhuo Chen, Yin Fang, Yichi Zhang, Lingbing Guo, Jiaoyan Chen, Jeff Z. Pan, Hua- jun Chen, and Wen Zhang. 2025. Noise-powered Multi-modal Knowledge Graph Representation Framework. InProceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Sch...

  7. [7]

    Shib Sankar Dasgupta, Swayambhu Nath Ray, and Partha Talukdar. 2018. HyTE: Hyperplane-based Temporally aware Knowledge Graph Embedding. InProceed- ings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Associa- tion for Computational Linguistics, Brussels,...

  8. [8]

    Daniel Daza, Michael Cochez, and Paul Groth. 2021. Inductive Entity Represen- tations from Text via Link Prediction. InProceedings of the Web Conference 2021 (Ljubljana, Slovenia)(WWW ’21). Association for Computing Machinery, New York, NY, USA, 798–808

  9. [9]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Jill Burstein, Christy...

  10. [10]

    Rishab Goel, Seyed Mehran Kazemi, Marcus Brubaker, and Pascal Poupart. 2020. Diachronic Embedding for Temporal Knowledge Graph Completion.Proceedings of the AAAI Conference on Artificial Intelligence34, 04 (Apr. 2020), 3988–3995

  11. [11]

    Hao Guo, Jiuyang Tang, Weixin Zeng, Xiang Zhao, and Li Liu. 2021. Multi-modal entity alignment in hyperbolic space.Neurocomputing461 (2021), 598–607

  12. [12]

    Yani Huang, Xuefeng Zhang, Richong Zhang, Junfan Chen, and Jaein Kim. 2024. Progressively Modality Freezing for Multi-Modal Entity Alignment. InProceed- ings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, B...

  13. [13]

    Tingsong Jiang, Tianyu Liu, Tao Ge, Lei Sha, Baobao Chang, Sujian Li, and Zhifang Sui. 2016. Towards Time-Aware Knowledge Graph Completion. InProceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, Yuji Matsumoto and Rashmi Prasad (Eds.). The COLING 2016 Organizing Committee, Osaka, Japan, 1715–1724

  14. [14]

    Woojeong Jin, Meng Qu, Xisen Jin, and Xiang Ren. 2020. Recurrent Event Network: Autoregressive Structure Inferenceover Temporal Knowledge Graphs. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 6669–6683

  15. [15]

    Jaejun Lee, Chanyoung Chung, Hochang Lee, Sungho Jo, and Joyce Whang. 2023. VISTA: Visual-Textual Knowledge Graph Representation Learning. InFindings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 7314–7328

  16. [16]

    Xinhang Li, Xiangyu Zhao, Jiaxing Xu, Yong Zhang, and Chunxiao Xing. 2023. IMF: Interactive Multimodal Fusion Model for Link Prediction. InProceedings of the ACM Web Conference 2023(Austin, TX, USA)(WWW ’23). Association for Computing Machinery, New York, NY, USA, 2572–2580

  17. [17]

    Zixuan Li, Xiaolong Jin, Wei Li, Saiping Guan, Jiafeng Guo, Huawei Shen, Yuanzhuo Wang, and Xueqi Cheng. 2021. Temporal Knowledge Graph Rea- soning Based on Evolutional Representation Learning. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Informa- tion Retrieval(Virtual Event, Canada)(SIGIR ’21). Association ...

  18. [18]

    Wanying Liang, Pasquale De Meo, Yong Tang, and Jia Zhu. 2024. A Survey of Multi-modal Knowledge Graphs: Technologies and Trends.ACM Comput. Surv. 56, 11, Article 273 (June 2024), 41 pages

  19. [19]

    Qika Lin, Jun Liu, Rui Mao, Fangzhi Xu, and Erik Cambria. 2023. TECHS: Temporal Logical Graph Networks for Explainable Extrapolation Reasoning. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguist...

  20. [20]

    Zhenxi Lin, Ziheng Zhang, Meng Wang, Yinghui Shi, Xian Wu, and Yefeng Zheng

  21. [21]

    Multi-modal Contrastive Representation Learning for Entity Alignment. InProceedings of the 29th International Conference on Computational Linguistics, Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wan- ner, Key-Sun Choi, Pum-Mo Ryu, Hsin-Hsi Chen, Lucia Donatelli, Heng Ji, Sadao Kurohashi, Patrizia Paggio, Nianwen Xue, Seokhwan K...

  22. [22]

    Fangyu Liu, Muhao Chen, Dan Roth, and Nigel Collier. 2021. Visual Pivoting for (Unsupervised) Entity Alignment.Proceedings of the AAAI Conference on Artificial Intelligence35, 5 (May 2021), 4257–4266

  23. [23]

    Rosenblum

    Ye Liu, Hui Li, Alberto Garcia-Duran, Mathias Niepert, Daniel Onoro-Rubio, and David S. Rosenblum. 2019. MMKG: Multi-modal Knowledge Graphs. InThe Semantic Web: 16th International Conference, ESWC 2019, Portorož, Slovenia, June 2–6, 2019, Proceedings(Portoroz, Slovenia). Springer-Verlag, Berlin, Heidelberg, 459–474

  24. [24]

    Xinyu Lu, Lifang Wang, Zejun Jiang, Shichang He, and Shizhong Liu. 2022. MMKRL: A robust embedding approach for multi-modal knowledge graph repre- sentation learning.Applied Intelligence52, 7 (2022), 7480–7497

  25. [25]

    Haodi Ma, Dzmitry Kasinets, and Daisy Zhe Wang. 2025. Transformer- Based Multimodal Knowledge Graph Completion with Link-Aware Contexts. arXiv:2501.15688 [cs.CL]

  26. [26]

    Hatem Mousselly-Sergieh, Teresa Botschen, Iryna Gurevych, and Stefan Roth

  27. [27]

    InProceedings of the seventh joint conference on lexical and computational semantics

    A multimodal translation-based approach for knowledge graph repre- sentation learning. InProceedings of the seventh joint conference on lexical and computational semantics. 225–234

  28. [28]

    Vardaan Pahuja, Weidi Luo, Yu Gu, Cheng-Hao Tu, Hong-You Chen, Tanya Berger- Wolf, Charles Stewart, Song Gao, Wei-Lun Chao, and Yu Su. 2024. Reviving the Context: Camera Trap Species Classification as Link Prediction on Multimodal Knowledge Graphs. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management(Boise, ID, US...

  29. [29]

    Zhiliang Peng, Li Dong, Hangbo Bao, Qixiang Ye, and Furu Wei. 2022. BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers. arXiv:2208.06366 [cs.CV]

  30. [30]

    Bin Shang, Yinliang Zhao, Jun Liu, and Di Wang. 2024. LAFA: Multimodal Knowl- edge Graph Completion with Link Aware Fusion and Aggregation.Proceedings of the AAAI Conference on Artificial Intelligence38, 8 (Mar. 2024), 8957–8965

  31. [31]

    Jinqing Shen, Chengjin Xu, Yingqi Liu, Xuhui Jiang, Jiaming Li, Zhenxin Huang, Jens Lehmann, and Xuesong Chen. 2025. Learning Temporal Knowledge Graphs via Time-Sensitive Graph Attention.IEEE Access13 (2025), 178517–178526

  32. [32]

    Bowen Song, Kossi Amouzouvi, Chengjin Xu, Maocai Wang, Jens Lehmann, and Sahar Vahdati. 2024. Temporal relevance for representing learning over temporal knowledge graphs.Semantic Web15, 6 (2024), 2695–2711

  33. [33]

    Zhiqing Sun, Zhi-Hong Deng, Jian-Yun Nie, and Jian Tang. 2019. Ro- tatE: Knowledge Graph Embedding by Relational Rotation in Complex Space. arXiv:1902.10197 [cs.LG]

  34. [34]

    Théo Trouillon, Johannes Welbl, Sebastian Riedel, Eric Gaussier, and Guillaume Bouchard. 2016. Complex Embeddings for Simple Link Prediction. InProceedings of The 33rd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 48), Maria Florina Balcan and Kilian Q. Weinberger (Eds.). PMLR, New York, New York, USA, 2071–2080

  35. [35]

    Jiapu Wang, Boyue Wang, Meikang Qiu, Shirui Pan, Bo Xiong, Heng Liu, Linhao Luo, Tengfei Liu, Yongli Hu, Baocai Yin, et al . 2023. A survey on temporal knowledge graph completion: Taxonomy, progress, and prospects.arXiv preprint arXiv:2308.02457(2023)

  36. [36]

    Meng Wang, Sen Wang, Han Yang, Zheng Zhang, Xi Chen, and Guilin Qi. 2021. Is visual context really helpful for knowledge graph? A representation learning perspective. InProceedings of the 29th ACM international conference on multimedia. 2735–2743

  37. [37]

    Xin Wang, Benyuan Meng, Hong Chen, Yuan Meng, Ke Lv, and Wenwu Zhu. 2023. TIVA-KG: A multimodal knowledge graph with text, image, video and audio. In Proceedings of the 31st ACM international conference on multimedia. 2391–2399

  38. [38]

    Yunpeng Wang, Bo Ning, Xin Wang, Chengfei Liu, and Guanyu Li. 2025. Seg- mentation Similarity Enhanced Semantic Related Entity Fusion for Multi-modal Knowledge Graph Completion. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval(Padua, Italy)(SIGIR ’25). Association for Computing Machinery, Ne...

  39. [39]

    Zikang Wang, Linjing Li, Qiudan Li, and Daniel Zeng. 2019. Multimodal Data Enhanced Representation Learning for Knowledge Graphs. In2019 International Time Imprint: Learning Time-Aware Representations in Multi-Modal Knowledge Graphs Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Joint Conference on Neural Networks (IJCNN). 1–8

  40. [40]

    W. X. Wilcke, P. Bloem, V. de Boer, and R. H. van t Veer. 2023. End-to-End Learning on Multimodal Knowledge Graphs. arXiv:2309.01169 [cs.LG]

  41. [41]

    Ruobing Xie, Zhiyuan Liu, Huanbo Luan, and Maosong Sun. 2017. Image- embodied knowledge representation learning. InProceedings of the 26th Interna- tional Joint Conference on Artificial Intelligence(Melbourne, Australia)(IJCAI’17). AAAI Press, 3140–3146

  42. [42]

    Derong Xu, Tong Xu, Shiwei Wu, Jingbo Zhou, and Enhong Chen. 2022. Relation- enhanced Negative Sampling for Multimodal Knowledge Graph Completion. InProceedings of the 30th ACM International Conference on Multimedia(Lisboa, Portugal)(MM ’22). Association for Computing Machinery, New York, NY, USA, 3857–3866

  43. [43]

    Bishan Yang, Wen tau Yih, Xiaodong He, Jianfeng Gao, and Li Deng. 2015. Em- bedding Entities and Relations for Learning and Inference in Knowledge Bases. arXiv:1412.6575 [cs.CL]

  44. [44]

    Yang, Ruikun Luo, and Jieming Yang

    Shundong Yang, Jing Yang, Xiaowen Jiang, Yuan Gao, Laurence T. Yang, Ruikun Luo, and Jieming Yang. 2025. Towards Multimodal Inductive Learning: Adaptively Embedding MMKG via Prototypes. InProceedings of the ACM on Web Conference 2025(Sydney NSW, Australia)(WWW ’25). Association for Computing Machinery, New York, NY, USA, 109–118

  45. [45]

    Jinchuan Zhang, Tianqi Wan, Chong Mu, Guangxi Lu, and Ling Tian. 2024. Learning Granularity Representation for Temporal Knowledge Graph Completion. InNeural Information Processing: 31th International Conference, ICONIP 2024, Auckland, New Zealand, December 2–6, 2024. Springer

  46. [46]

    Ming Zhang, Ke Chang, and Yunfang Wu. 2024. Multi-modal Semantic Under- standing with Contrastive Cross-modal Feature Alignment. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, an...

  47. [47]

    Yichi Zhang, Mingyang Chen, and Wen Zhang. 2023. Modality-Aware Negative Sampling for Multi-modal Knowledge Graph Embedding. In2023 International Joint Conference on Neural Networks (IJCNN). 1–8

  48. [48]

    Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Binbin Hu, Ziqi Liu, Wen Zhang, and Huajun Chen. 2024. NativE: Multi-modal Knowledge Graph Comple- tion in the Wild. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval(Washington DC, USA)(SIGIR ’24). Association for Computing Machinery, New York...

  49. [49]

    Yichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu, Binbin Hu, Ziqi Liu, Wen Zhang, and Huajun Chen. 2025. Tokenization, Fusion, and Augmentation: To- wards Fine-grained Multi-modal Entity Representation. InAAAI. AAAI Press, 13322–13330

  50. [50]

    Yichi Zhang, Zhuo Chen, Lingbing Guo, yajing Xu, Binbin Hu, Ziqi Liu, Wen Zhang, and Huajun Chen. 2025. Multiple Heads are Better than One: Mixture of Modality Knowledge Experts for Entity Representation Learning. InThe Thir- teenth International Conference on Learning Representations

  51. [51]

    Yichi Zhang, Zhuo Chen, Lei Liang, Huajun Chen, and Wen Zhang. 2024. Unleash- ing the Power of Imbalanced Modality Information for Multi-modal Knowledge Graph Completion. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Nicoletta Calzolari, Min-Yen Kan, Veronique H...

  52. [52]

    Yichi Zhang, Zhuo Chen, and Wen Zhang. 2023. MACO: A Modality Adversarial and Contrastive Framework for Modality-Missing Multi-modal Knowledge Graph Completion. InNatural Language Processing and Chinese Computing, Fei Liu, Nan Duan, Qingting Xu, and Yu Hong (Eds.). Springer Nature Switzerland, Cham, 123–134

  53. [53]

    Yichi Zhang and Wen Zhang. 2022. Knowledge graph completion with pre- trained multimodal transformer and twins negative sampling.arXiv preprint arXiv:2209.07084(2022)

  54. [54]

    Yu Zhao, Xiangrui Cai, Yike Wu, Haiwei Zhang, Ying Zhang, Guoqing Zhao, and Ning Jiang. 2022. MoSE: Modality Split and Ensemble for Multimodal Knowledge Graph Completion. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistic...

  55. [55]

    Yu Zhao, Ying Zhang, Baohang Zhou, Xinying Qian, Kehui Song, and Xiangrui Cai

  56. [56]

    InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA)(SIGIR ’24)

    Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph Completion. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington DC, USA)(SIGIR ’24). Association for Computing Machinery, New York, NY, USA, 102–111

  57. [57]

    Cunchao Zhu, Muhao Chen, Changjun Fan, Guangquan Cheng, and Yan Zhang

  58. [58]

    Learning from History: Modeling Temporal Knowledge Graphs with Sequential Copy-Generation Networks.Proceedings of the AAAI Conference on Artificial Intelligence35, 5 (May 2021), 4732–4740

  59. [59]

    Xiangru Zhu, Zhixu Li, Xiaodan Wang, Xueyao Jiang, Penglei Sun, Xuwu Wang, Yanghua Xiao, and Nicholas Jing Yuan. 2024. Multi-Modal Knowledge Graph Construction and Application: A Survey.IEEE Transactions on Knowledge and Data Engineering36, 2 (2024), 715–735. Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009