REVIEW 3 major objections 4 minor 66 references
Multi-level Mixture of Experts for Multimodal Entity Linking
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A multi-level mixture-of-experts model with LLM-selected entity descriptions sets new state-of-the-art scores on three multimodal entity linking benchmarks.
desk verdict The SOTA claim is carried by the LLM description-selection module, and until the authors control for answer leakage the architecture's independent contribution is not proven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Switch Mixture of Experts (SMoE), applied at two levels: intra-level SMoE selects the most informative textual tokens and visual patches within each modality, while inter-level SMoE dynamically combines cross-modal features in both text-to-visual and visual-to-text directions. Around this, the Description-aware Mention Enhancement (DME) module uses an LLM (GPT-3.5 in the default configuration) to pick the WikiData description that best fits the mention's textual context and appends it to the context, increasing semantic content. The SMoE layers use a sparse router that activates k of K expert feed-forward networks per token or patch, and the output is split back into coarse- and fine-grained features for matching.
What would settle it
Replace the LLM-selected description in DME with a random description from the candidate list while keeping MMoE unchanged; if performance drops back to the no-DME level, the gains come from the LLM's selection rather than the model's architecture.
Extended reading notes
Core claim
The central claim is that combining description-aware mention enhancement with sparse expert routing yields the best reported performance on three standard MEL datasets. Concretely, MMoE with the DME module achieves 93.75 MRR and 90.77 Hits@1 on WikiMEL, 89.86 MRR on RichpediaMEL, and 84.23 MRR on WikiDiverse, each above the best baseline. The paper further shows that adding the DME module alone improves existing strong models such as MIMIC and M3EL, and that the full model retains its advantage when training on only 10% or 20% of the data. The authors interpret these results as evidence that both enriching mention semantics with LLM-selected descriptions and dynamically selecting fine-grained intra- and inter-modal features are effective, complementary strategies for MEL.
Load-bearing premise
The model assumes the LLM in the DME module reliably identifies the true entity's WikiData description, and that adding that description to the mention context does not leak the target, so the reported gains reflect the architecture rather than the answer being present in the input.
Editorial extensions
If this is right
- If the reported results hold, MMoE+DME becomes the new state of the art on WikiMEL, RichpediaMEL, and WikiDiverse, with a single architecture across datasets.
- The improvement pattern suggests that enriching short mention contexts with LLM-selected entity descriptions is a transferable fix for mention ambiguity, since it also lifts the performance of existing baselines.
- The ablation results indicate that the textual intra-level SMoE is the component with the largest impact, pointing to fine-grained text feature selection as the key driver of MEL accuracy.
- The model's gains persist when training on only 10% or 20% of the data, which matters for low-resource applications.
Reading between the lines
- The DME module's reliance on an LLM to pick the correct description may mean part of the measured gain comes from answer leakage: if the LLM often selects the true entity's description, the mention representation already contains the answer text, and the remaining architecture's contribution is unclear without a controlled comparison to random or gold description selection.
- The approach could be adapted to other ambiguous multimodal tasks, such as visual grounding or multimodal question answering, where short queries and informative image regions matter.
- The appendix's finding that LLaMA2 and LLaMA3.1 underperform GPT-3.5 in DME suggests that the quality of the LLM's description ranking is a reproducibility bottleneck; open-weight models may not deliver the same gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Multi-level Mixture of Experts (MMoE), a model for Multimodal Entity Linking (MEL) with four components: a description-aware mention enhancement (DME) module that uses GPT-3.5 to select a WikiData description to append to the mention context, a multimodal feature extraction module based on CLIP, an intra-level mixture-of-experts module, and an inter-level mixture-of-experts module. The authors evaluate on WikiMEL, RichpediaMEL, and WikiDiverse, reporting state-of-the-art MRR and Hits@k results, and they include ablations, parameter sensitivity experiments, and complexity analysis. The code is publicly available.
Significance. If the central claim is correct, the paper makes a useful empirical contribution: a modular MEL architecture that combines LLM-based description enrichment with sparse mixture-of-experts feature selection, evaluated on three standard benchmarks with code release and extensive ablations. The architecture is clearly described and the engineering effort is substantial. However, the headline claim that the MMoE architecture itself achieves state-of-the-art performance is not yet separable from the effect of the DME module, which appends a candidate WikiData description selected by an LLM. Because the candidate list at test time contains the gold entity's description, the reported gains may be attributable to the LLM's selection rather than to the mixture-of-experts design. The evidence needed to resolve this attribution is missing, so the significance of the architectural contribution is currently conditional.
major comments (3)
- [Section 3.2, Eq. (2)] The load-bearing assumption of the DME module is unsupported. The candidate set M_w includes the gold entity's WikiData description at test time, and GPT-3.5 returns the description judged 'most pertinent' given the mention context; that description is then concatenated into m_t^* and fed to the encoders. If GPT-3.5 frequently selects the gold description, the model's input contains the answer text before any SMoE matching occurs, so the reported gains cannot be attributed to the MMoE architecture. The paper reports neither the LLM selection accuracy nor any control condition (random description, gold description, no description, or a selector that cannot see the gold description). The only LLM ablation, Appendix C Table 7, shows that replacing GPT-3.5 with LLaMA2-7B drops WikiDiverse MRR from 84.23 to 74.34, which is consistent with the LLM choice dominating performance. I request these diagnostics before the headline SOTA claim is accepted.
- [Section 4.2, Table 3] The statement that 'regardless of whether the DME module is incorporated, MMoE consistently achieves the best performance on all datasets' is internally contradicted by the table: without DME, MMoE has WikiMEL MRR 92.53, below OT-MEL's 92.59. The no-DME condition is the only evidence for the MoE architecture's intrinsic value, so this discrepancy should be corrected and the claim weakened or supported with significance testing. The reported best-baseline margins are also small (e.g., WikiDiverse MRR +0.70, RichpediaMEL MRR +0.82), which makes the consistency claim especially delicate.
- [Table 3, baseline protocol] Several strong baselines, including OT-MEL, MELOV, and FissFuse, are marked with ♦ to indicate that their numbers are taken from the original papers rather than rerun under the same protocol. No error bars, seeds, or significance tests are reported for any condition. Given that the headline improvements over the strongest baselines are as small as 0.70 MRR on WikiDiverse and 0.82 MRR on RichpediaMEL, the SOTA claim requires either rerunning those baselines under the same protocol or reporting variance and significance. This is particularly important because the DME ablation for these baselines was not possible (no code), leaving the comparison asymmetric.
minor comments (4)
- [Section 3.5] There are several typos in this section: 'alliviate' should be 'alleviate' and 'acerage' should be 'average'.
- [Section 4.1] In the implementation details, 'CLIP-ViT-Base-Pathch32' contains a typo: 'Pathch' should be 'Patch'.
- [Section 4.4 and Figure 3] The notation 'E4T3' and similar is used to describe expert/top-expert configurations, but the notation is not explicitly defined; also, the axis labels in Figure 3 are garbled in the provided manuscript and should be redrawn for readability.
- [Table 3] The 'improvement (%)' row is ambiguous and includes a negative value for RichpediaMEL H@5 (-0.03) while the text emphasizes outstanding performance; the row should be defined precisely and negative entries should be discussed.
Circularity Check
No significant circularity: MMoE's MoE scoring chain is self-contained; the DME module raises answer-leakage concerns that are validity issues, not circular reductions.
full rationale
The derivation chain is self-contained: Equations (3)-(8) define SMoE routing, intra/inter-level matching, and a contrastive loss from the CLIP features of Eq. (2), with no parameter fitted to the headline metrics and no equation defined in terms of the target entity. The reported gains are empirical comparisons against external baselines (Table 3), and the self-citations (e.g., M3EL [17]) are used only as a baseline and as a source of low-resource results, not as an authority that forces the MMoE construction. The DME module (Section 3.2, Eq. (2)) does raise a legitimate answer-leakage concern: GPT-3.5 selects a WikiData description from a candidate list that contains the gold entity's description, the selected text is concatenated into the mention, and the paper does not report selection accuracy or random-/gold-description controls; Appendix C shows the LLM choice strongly affects results. However, this is an experimental-validity risk, not circularity: the LLM selection is an external, non-fitted component, and the matching equations do not reduce to that selection by construction. The Section 4.2 claim that MMoE without DME 'consistently achieves the best performance' is contradicted by Table 3 on WikiMEL MRR (92.53 vs OT-MEL 92.59), but that is a correctness/consistency issue rather than a circular step.
Assumptions & free parameters
free parameters (5)
- Number of experts K =
Not reported; searched over {2,4,6,8,10}, optimal 4 on all datasets per Figure 3(a)
- Number of top experts k =
Not reported; optimal 3 for WikiMEL, 4 for RichpediaMEL, 2 for WikiDiverse per Figure 3(a)
- Embedding dimension d =
96 optimal on all three datasets per Figure 3(c)
- Max text length =
50 optimal on all three datasets per Figure 3(d)
- Learning rate =
1e-5 optimal in sensitivity analysis; final value not reported
assumptions (5)
- domain assumption The candidate entity set E contains the correct entity for every mention.
- domain assumption GPT-3.5 reliably selects the correct WikiData description from the candidate description list.
- domain assumption WikiData descriptions are available for candidate entities and can be concatenated to mention contexts without leaking the answer.
- domain assumption Pre-trained CLIP textual and visual encoders provide adequate representations for MEL.
- standard math In-batch negatives approximate the full candidate distribution in the contrastive loss.
Cite this review
Pith. "Pith review of Multi-level Mixture of Experts for Multimodal Entity Linking." pith.science (2026). https://pith.science/paper/S3MCPN4U
@misc{pith2026250707108,
author = {Pith},
title = {Pith review of: Multi-level Mixture of Experts for Multimodal Entity Linking},
year = {2026},
howpublished = {\url{https://pith.science/paper/S3MCPN4U}},
note = {Machine review of arXiv:2507.07108}
}
read the original abstract
Multimodal Entity Linking (MEL) aims to link ambiguous mentions within multimodal contexts to associated entities in a multimodal knowledge base. Existing approaches to MEL introduce multimodal interaction and fusion mechanisms to bridge the modality gap and enable multi-grained semantic matching. However, they do not address two important problems: (i) mention ambiguity, i.e., the lack of semantic content caused by the brevity and omission of key information in the mention's textual context; (ii) dynamic selection of modal content, i.e., to dynamically distinguish the importance of different parts of modal information. To mitigate these issues, we propose a Multi-level Mixture of Experts (MMoE) model for MEL. MMoE has four components: (i) the description-aware mention enhancement module leverages large language models to identify the WikiData descriptions that best match a mention, considering the mention's textual context; (ii) the multimodal feature extraction module adopts multimodal feature encoders to obtain textual and visual embeddings for both mentions and entities; (iii)-(iv) the intra-level mixture of experts and inter-level mixture of experts modules apply a switch mixture of experts mechanism to dynamically and adaptively select features from relevant regions of information. Extensive experiments demonstrate the outstanding performance of MMoE compared to the state-of-the-art. MMoE's code is available at: https://github.com/zhiweihu1103/MEL-MMoE.
Figures
Reference graph
Works this paper leans on
-
[1]
Omar Adjali, Romaric Besançon, Olivier Ferret, Hervé Le Borgne, and Brigitte Grau. 2020. Multimodal Entity Linking for Tweets. In ECIR. 463–478
work page 2020
-
[2]
Ali Ahmadvand, Harshita Sahijwani, Jason Ingyu Choi, and Eugene Agichtein
-
[3]
Meta AI. 2024. Introducing meta llama 3: The most capable openly available LLM to date
work page 2024
-
[4]
Ujjwala Anantheswaran, Himanshu Gupta, Kevin Scaria, Shreyas Verma, Chitta Baral, and Swaroop Mishra. 2024. Investigating the Robustness of LLMs on Math Word Problems. arXiv (2024)
work page 2024
-
[5]
Tom Ayoola, Shubhi Tyagi, Joseph Fisher, Christos Christodoulopoulos, and Andrea Pierleoni. 2022. ReFinED: An Efficient Zero-shot-capable Approach to End-to-End Entity Linking. In NAACL. 209–220
work page 2022
-
[6]
Ilaria Bordino, Yelena Mejova, and Mounia Lalmas. 2013. Penguins in sweaters, or serendipitous entity search on user-generated content. In CIKM. 109–118. Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li, and Jeff Z. Pan
work page 2013
-
[7]
Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. 2021. Au- toregressive Entity Retrieval. In ICLR. 1–20
work page 2021
-
[8]
Nicola De Cao, Ledell Wu, Kashyap Popat, Mikel Artetxe, Naman Goyal, Mikhail Plekhanov, Luke Zettlemoyer, Nicola Cancedda, Sebastian Riedel, and Fabio Petroni. 2022. Multilingual Autoregressive Entity Linking. In TACL. 274–290
work page 2022
Show all 66 references
-
[9]
Yixin Cao, Lei Hou, Juanzi Li, and Zhiyuan Liu. 2018. Neural Collective Entity Linking. In COLING. 675–686
2018
-
[10]
Pan, Ian Horrocks, and Huajun Chen
Jiaoyan Chen, Freddy Lecue, Jeff Z. Pan, Ian Horrocks, and Huajun Chen. 2018. Knowledge-based Transfer Learning Explanation. In Proc. of the International Conference on Principles of Knowledge Representation and Reasoning (KR2018) . 349–358
2018
-
[11]
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. In EMNLP. 1724–1734
2014
-
[12]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL. 4171–4186
2019
-
[13]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recogn...
2021
-
[14]
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, Zicheng Liu, and Michael Zeng. 2022. An Empirical Study of Training End-to-End Vision-and- Language Transformers. In CVPR. 18145–18155
2022
-
[15]
Pan, Martin J
Biralatei Fawei, Jeff Z. Pan, Martin J. Kollingbaum, and Adam Z. Wyner. 2020. A Semi-automated Ontology Construction for Legal Question Answering. New Generation Computing (2020), 453–478
2020
-
[16]
Bowei He, Xu He, Yingxue Zhang, Ruiming Tang, and Chen Ma. 2023. Dynami- cally Expandable Graph Convolution for Streaming Recommendation. In WWW. 1457–1467
2023
-
[17]
Zhiwei Hu, Víctor Gutiérrez-Basulto, Ru Li, and Jeff Z. Pan. 2024. Multi-level Matching Network for Multimodal Entity Linking. arXiv (2024)
2024
-
[18]
Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li, and Jeff Z. Pan. 2024. Leveraging Intra-modal and Inter-modal Interaction for Multi-Modal Entity Alignment. arXiv (2024)
2024
-
[19]
Wenyu Huang, Guancheng Zhou, Mirella Lapata, Pavlos Vougiouklis, Sebastien Montella, and Jeff Z Pan. 2025. Prompting Large Language Models with Knowl- edge Graphs for Question Answering involving Long-tail Facts.Knowledge-Based Systems (KBS) (2025)
2025
-
[20]
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. In ICML. 5583–5594
2021
-
[21]
Phong Le and Ivan Titov. 2018. Improving Entity Linking by Modeling Latent Relations between Mentions. In ACL. 1595–1604
2018
-
[22]
Selvaraju, Akhilesh Gotmare, Shafiq R
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty, Caiming Xiong, and Steven Chu-Hong Hoi. 2021. Align before Fuse: Vision and Language Representation Learning with Momentum Distillation. In NeurIPS. 9694–9705
2021
-
[23]
Lizi Liao, Yunshan Ma, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2018. Knowledge-aware Multimodal Dialogue Systems. In MM. 801–809
2018
-
[24]
Qi Liu, Yongyi He, Tong Xu, Defu Lian, Che Liu, Zhi Zheng, and Enhong Chen
-
[25]
Xiao Liu, Shiyu Zhao, Kai Su, Yukuo Cen, Jiezhong Qiu, Mengdi Zhang, Wei Wu, Yuxiao Dong, and Jie Tang. 2022. Mask and Reason: Pre-Training Knowledge Graph Transformers for Complex Logical Queries. In KDD. 1120–1130
2022
-
[26]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv (2019)
2019
-
[27]
Xinwei Long, Jiali Zeng, Fandong Meng, Jie Zhou, and Bowen Zhou. 2024. Trust in Internal or External Knowledge? Generative Multi-Modal Entity Linking with Knowledge Retriever. In ACL. 7559–7569
2024
-
[28]
Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-Based Knowledge Conflicts in Question Answer- ing. In EMNLP. 7052–7063
2021
-
[29]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In ICLR. 1–11
2019
-
[30]
Pengfei Luo, Tong Xu, Che Liu, Suojuan Zhang, Linli Xu, Minglei Li, and Enhong Chen. 2024. Bridging Gaps in Content and Knowledge for Multimodal Entity Linking. In MM. 9311–9320
2024
-
[31]
Pengfei Luo, Tong Xu, Shiwei Wu, Chen Zhu, Linli Xu, and Enhong Chen. 2023. Multi-Grained Multimodal Interaction Network for Entity Linking. In KDD. 1583–1594
2023
-
[32]
Seungwhan Moon, Leonardo Neves, and Vitor Carvalho. 2018. Multimodal Named Entity Disambiguation for Noisy Social Media Posts. In ACL. 2000–2008
2018
-
[33]
OpenAI. 2023. Chatgpt
2023
-
[34]
J.Z. Pan, D. Calvanese, T. Eiter, I. Horrocks, M. Kifer, F. Lin, and Y. Zhao (Eds.)
-
[35]
Jeff Pan, Stuart Taylor, and Edward Thomas. 2009. Reducing Ambiguity in Tagging Systems with Folksonomy Search Expansion. In 6th Annual European Semantic Web Conference (ESWC2009). 669–683. http://data.semanticweb.org/ conference/eswc/2009/paper/150
2009
-
[36]
J.Z. Pan, G. Vetere, J.M. Gomez-Perez, and H. Wu (Eds.). 2017. Exploiting Linked Data and Knowledge Graphs for Large Organisations . Springer
2017
-
[37]
Pan, Mei Zhang, Kuldeep Singh, Frank Van Harmelen, Jinguang Gu, and Zhi Zhang
Jeff Z. Pan, Mei Zhang, Kuldeep Singh, Frank Van Harmelen, Jinguang Gu, and Zhi Zhang. 2019. Entity Enabled Relation Linking. In ISWC. 523–538
2019
-
[38]
Peters, Mark Neumann, Robert L
Matthew E. Peters, Mark Neumann, Robert L. Logan IV, Roy Schwartz, Vidur Joshi, Sameer Singh, and Noah A. Smith. 2019. Knowledge Enhanced Contextual Word Representations. In EMNLP. 43–54
2019
-
[39]
Flavio Petruzzellis, Alberto Testolin, and Alessandro Sperduti. 2024. Assessing the Emergent Symbolic Reasoning Abilities of Llama Large Language Models. arXiv (2024)
2024
-
[40]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In ICML. 8748–8763
2021
-
[41]
Chenwei Ran, Wei Shen, and Jianyong Wang. 2018. An Attention Factor Graph Model for Tweet Entity Linking. In WWW. 1135–1144
2018
-
[42]
Senbao Shi, Zhenran Xu, Baotian Hu, and Min Zhang. 2024. Generative Multi- modal Entity Linking. In COLING. 7654–7665
2024
-
[43]
Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Net- works for Large-Scale Image Recognition. In ICLR. 1–14
2015
-
[44]
Shezheng Song, Shan Zhao, Chengyu Wang, Tianwei Yan, Shasha Li, Xiaoguang Mao, and Meng Wang. 2024. A Dual-Way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking. In AAAI. 19008–19016
2024
-
[45]
Xuhui Sui, Ying Zhang, Yu Zhao, Kehui Song, Baohang Zhou, and Xiaojie Yuan
-
[46]
Pan, and Derek H
Edward Thomas, Jeff Z. Pan, and Derek H. Sleeman. 2007. ONTOSEARCH2: Searching Ontologies Semantically.. In OWLED
2007
-
[47]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lam- ple. 2023. LLaMA: Open and Efficient Foundation ...
2023
-
[48]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy F...
2023
-
[49]
MELOV: Multimodal Entity Linking with Optimized Visual Features in Latent Space. In ACL. 816–826
-
[50]
Denny Vrandecic and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Commun. ACM (2014), 78–85
2014
-
[51]
Pan, and Kam-Fai Wong
Hongru Wang, Wenyu Huang, Yufei Wang, Yuanhao Xi, Jianqiao Lu, Huan Zhang, Nan Hu, Zeming Liu, Jeff Z. Pan, and Kam-Fai Wong. 2025. Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges. InIn the Proceedings of the 63rd Annual Meeting of the Associati...
2025
-
[52]
Meng Wang, Haofen Wang, Guilin Qi, and Qiushuo Zheng. 2020. Richpedia: A Large-Scale, Comprehensive Multi-Modal Knowledge Graph. Big Data Res. (2020), 1–11
2020
-
[53]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In NeurIPS. 5998–6008
2017
-
[54]
Xuwu Wang, Junfeng Tian, Min Gui, Zhixu Li, Rui Wang, Ming Yan, Lihan Chen, and Yanghua Xiao. 2022. WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity Types. In ACL. 4785–4797
2022
-
[55]
Ledell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel, and Luke Zettle- moyer. 2020. Scalable Zero-shot Entity Linking with Dense Entity Retrieval. In EMNLP. 6397–6407
2020
-
[56]
Wenhan Xiong, Mo Yu, Shiyu Chang, Xiaoxiao Guo, and William Yang Wang
-
[57]
Peng Wang, Jiangheng Wu, and Xiaohang Chen. 2022. Multimodal Entity Linking with Gated Hierarchical Fusion and Contrastive Training. In SIGIR. 938–948
2022
-
[58]
Gongrui Zhang, Chenghuan Jiang, Zhongheng Guan, and Peng Wang. 2023. Multimodal Entity Linking with Mixed Fusion Mechanism. In DASFAA. 607– 622
2023
-
[59]
Li Zhang, Zhixu Li, and Qiang Yang. 2021. Attention-Based Multimodal Entity Linking with High-Quality Images. In DASFAA. 533–548
2021
-
[60]
Zefeng Zhang, Jiawei Sheng, Chuang Zhang, Liangyunzhi Liangyunzhi, Wenyuan Zhang, Siqi Wang, and Tingwen Liu. 2024. Optimal Transport Guided Correlation Assignment for Multimodal Entity Linking. In ACL. 4103–4117
2024
-
[61]
Improving Question Answering over Incomplete KBs with Knowledge- Aware Reader. In ACL. 4258–4264. Multi-level Mixture of Experts for Multimodal Entity Linking
-
[62]
Dongjie Zhang and Longtao Huang. 2022. Multimodal Knowledge Learning for Named Entity Disambiguation. In EMNLP. 3160–3169
2022
-
[66]
Person of the Year
Qiushuo Zheng, Hao Wen, Meng Wang, and Guilin Qi. 2022. Visual Entity Linking via Multi-modal Learning. Data Intell. (2022), 1–19. Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li, and Jeff Z. Pan Appendix A Baselines We compare MMoE with three types of baselines, in...
2022
-
[2017]
Springer
Reasoning Web: Logical Foundation of Knowledge Graph Construction and Querying Answering. Springer
-
[2019]
ConCET: Entity-Aware Topic Classification for Open-Domain Conversa- tional Agents. In CIKM. 1371–1380
-
[2024]
UniMEL: A Unified Framework for Multimodal Entity Linking with Large Language Models. In CIKM. 1909–1919
1909
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.