REVIEW 2 major objections 5 minor 1 cited by
REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims unified multimodal relation extraction is feasible, and that its REMOTE framework with multilevel optimal transport and mixture-of-experts outperforms all compared baselines on almost every metric across the new UMRE…
desk verdict The UMRE dataset is a genuine contribution; the model's optimal transport module is not reproducible as written, and the ablation that credits it most cannot be checked until the authors fix the math and release code. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the multilevel optimal transport (MOT) fusion module. It takes low-level features from encoder layers $0$ to $l-1$ as the source distribution $\mu$ and the high-level feature at layer $l$ as the target distribution $\nu$, solves a Sinkhorn-regularized transport problem to obtain a plan $\Pi^*$, and fuses the transported low-level features with the high-level features through a learned weight $\alpha$. Together with the multimodal mixture-of-experts (MMoE) router, which weight-combines textual-only, visual-only, and cross-modal hierarchical attention features per triplet, MOT is what the paper claims carries the improvement. The UMRE dataset itself is also machinery: it provides the three triplet types and the 28-relation label space that lets the method be trained and evaluated as a single task.
What would settle it
Run the released REMOTE code on UMRE with the MOT module replaced by a dimensionally correct transport plan, for example a Sinkhorn plan with feature rows as marginals, and compare against Table 3 (F1 69.17 with OT, 64.82 without). If the corrected OT plan does not reproduce the gain, the claimed mechanism is falsified. Also verify whether Eq. 6's constraints hold for the shapes in the implementation: with $a=2u$ and $b=lu$, $\Pi \mathbf{1}_b = \mu$ requires $\mu$ to be a vector of length $a$, not a $(b \times d)$ matrix.
Extended reading notes
Core claim
The central discovery claim is that unified multimodal relation extraction—predicting relations among textual entities, among visual objects, and across text and images in one model—is feasible and beats task-specific designs. The evidence is the UMRE dataset and the REMOTE framework: multilevel optimal transport preserves low-level visual and textual detail that single-layer fusion loses, and the multimodal mixture-of-experts router assigns each relational triplet the interaction features best suited to it, giving a 5.3-point F1 gain over the previous best on UMRE and consistent gains on MORE and MNRE. The paper further claims that this is the first unified formulation of the task, going beyond both MNRE-style text-entity extraction and MORE-style text-object extraction.
Load-bearing premise
The load-bearing premise is that the multilevel optimal transport module performs a genuine, well-defined feature alignment; as printed, the equations are dimensionally inconsistent, so if the implementation does not match the described transport problem, the attribution of the gains to OT is unsupported.
Editorial extensions
If this is right
- Unified MRE becomes a single benchmark problem, so future methods no longer need separate designs for MNRE-style and MORE-style extraction.
- MLLM-generated captions plus low-level features become a reusable recipe: the paper shows that stronger captioning MLLMs improve extraction without retraining the core model.
- Mixture-of-experts routing gives a per-triplet choice of modality evidence, and the weight visualizations suggest that spatial-temporal relations favor lower-layer features while role relations favor visual and bidirectional features.
- The UMRE dataset gives the community a large testbed with 55,021 triplets, 28 relations, and three triplet types, enabling direct comparison of unified approaches.
- The framework's state-of-the-art numbers on three datasets imply that task-specific modality filtering can be replaced by a single dynamic router without losing performance.
Reading between the lines
- The OT component may not be load-bearing: the ablation that removes MOT also removes multilevel feature depth, so the 4.35-point F1 drop on UMRE could be due to losing low-level features altogether rather than to transport itself.
- A dimensionally faithful reimplementation of Eqs. 5-8 would clarify whether the transport plan is computed over feature rows or over a per-example probability vector; the released code should settle this.
- The same router-plus-multilevel-fusion idea could transfer to other multimodal tasks such as entity normalization or visual question answering, where fine-grained object detail and text semantics interact.
- The dataset construction relies on MLLM candidates and human adjudication, and the reported Kappa of 0.7325 suggests the task is harder than typical relation extraction, so future work may need more refined annotation protocols.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes REMOTE, a unified multimodal relation extraction framework that jointly extracts intra-modal (text-text, object-object) and inter-modal (text-object) relational triplets. The method combines a Multilevel Optimal Transport (MOT) fusion module, intended to preserve low-level encoder features, with a Multimodal Mixture-of-Experts (MMoE) router that selects per-triplet interaction features. The authors also introduce the UMRE dataset, containing 55,021 triplets from 12,737 text-image pairs, built by extending MNRE and MORE with MLLM-assisted candidate extraction and manual verification. Experiments on UMRE, MORE, and MNRE report state-of-the-art or competitive results, with ablations attributing gains to the MOT and MMoE modules.
Significance. If the technical description is corrected and the results verified, the paper would make a useful contribution: it is the first to define and evaluate a unified task covering all three triplet types, it contributes a sizable human-verified benchmark, and it reports extensive comparisons with many recent baselines, including honest reproductions of FocalMRE on MNRE. The authors release their resources, which supports reproducibility. The main weakness is that the central MOT module is specified with dimensionally inconsistent equations, and the empirical claims on the more established datasets rest on small margins without variance reporting.
major comments (2)
- [Section 6.1 and Table 2] The optimal transport formulation is dimensionally inconsistent and therefore under-specified. The source is defined as μ∈R^{(l·2u)×d_v} and the target as ν∈R^{2u×d_v}; the text sets b=l·2u and a=2u. With these definitions, a plan transporting μ onto ν should be in R^{b×a}, with row-sum μ and column-sum ν. Instead, Eq. (6) states Π∈R^{a×b}_+ with constraints Π1_b=μ and Π^T1_a=ν. The left-hand sides are vectors of lengths a and b, while μ and ν are matrices with d_v columns, so the constraints are not well-formed. Equation (8), F_V^{l'}=Π^*·μ, is dimensionally valid only if Π^*∈R^{a×b} and μ∈R^{b×d_v}, which contradicts the stated a,b and the stated constraints. The cost matrix in Eq. (7) is also indexed as C_{ij} for i∈[a], j∈[b] but then described as C∈R^{m×n} without specifying m,n. Moreover, the paper does not explain how feature matrices are converted into probability distributions; an OT plan requires probability vectors, not raw feature rows. Because Table 3 attributes the largest ablation gain to MOT (78.47 vs 75.64 accuracy and 69.17 vs 64.82 F1 on UMRE), the reader cannot independently determine that the module performs optimal transport rather than a reweighted similarity operation. Please correct the plan dimensions, marginal constraints, cost indexing, and specify the normalization (or learnable weighting) used to obtain histograms from feature rows.
- [Section 6.4 and Table 2] The empirical claim that REMOTE outperforms all baselines on almost all metrics is not fully supported by the reported statistics because Section 6.1 only states that experiments are averaged over 3 runs, but no standard deviations or significance tests are provided. On the MORE dataset, the reported improvements over the best baseline are small (e.g., +0.14 accuracy and +1.16 F1), so without variance information these gains may be within run-to-run variability. The larger gains on the new UMRE dataset are more convincing, but the general claim of superiority across all three datasets would be stronger with error bars or significance tests.
minor comments (5)
- [Abstract and Section 5] The abstract contains a typo: 'performanc' should be 'performance'. In Section 5, 'as shown in in Fig. 3' duplicates the word 'in'.
- [Table 2] The formatting of the FocalMRE row (e.g., '88.85/86.96*' and '88.01/86.21*') and of the ΔSOTA row is confusing; please split reported and reproduced results into separate columns or clearly label them.
- [References] References [15] and [16] are duplicates of the same paper (Hu et al., ACL 2023); please merge them.
- [Section 4.2] The annotation section mentions a Weighted Cohen's Kappa of 0.7325, but it does not state the number of samples that were double-annotated or the exact scoring procedure; please clarify the protocol.
- [Eq. (13)] The notation 'H_{T,V}^{⟨s⟩}' and 'H_{V,T}^{o_i}' in Eq. (13) is not explicitly defined; please clarify which of the previously introduced interaction features are used for textual entities versus visual objects.
Circularity Check
No circularity found: REMOTE is an empirical systems paper whose claims are benchmark results, not derived quantities or fitted inputs renamed as predictions.
full rationale
I walked the paper's claimed derivation chain: the task definition (Section 3), dataset construction (Section 4), the proposed REMOTE architecture (Section 5), and the experimental evaluation (Section 6). No step reduces to its own input by construction. The UMRE dataset is assembled from MNRE/MORE data with MLLM-assisted extraction and human annotation, independently of the REMOTE model; the model's parameters are learned from training data and evaluated on held-out test splits. The MoE routing weights are learned behaviors assessed by ablation, not quantities forced by the problem definition. The MOT module's optimal-transport equations (Eqs. 5-8) are dimensionally inconsistent as written—the marginal constraints do not type-check against the matrix-valued distributions, and the cost matrix indices appear swapped—but this is a reproducibility/correctness defect, not circularity: the module's measured ablation gain is an empirical claim, not a quantity that is equivalent to the paper's inputs by definition. There are no load-bearing self-citations, no imported uniqueness theorems, and no fitted parameters renamed as predictions. The paper is self-contained against external benchmarks, so the honest circularity finding is a score of 0.
Assumptions & free parameters
free parameters (5)
- lambda (entropy regularization strength)
- l (number of lower layers for MOT source)
- alpha (learnable fusion weight)
- P (MoE mapping matrix)
- visual feature dimension 4096 =
4096
assumptions (4)
- standard math Sinkhorn-Knopp algorithm converges to the entropy-regularized optimal transport plan.
- domain assumption Visual objects closer to the image center, larger, and in the foreground are more likely to be primary objects.
- domain assumption MLLM-generated captions for visual objects accurately describe the objects.
- domain assumption The UMRE annotations are consistent enough (Kappa 0.7325) for reliable evaluation.
Cite this review
Pith. "Pith review of REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts." pith.science (2026). https://pith.science/paper/OYNS6Z55
@misc{pith2026250904844,
author = {Pith},
title = {Pith review of: REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts},
year = {2026},
howpublished = {\url{https://pith.science/paper/OYNS6Z55}},
note = {Machine review of arXiv:2509.04844}
}
read the original abstract
Multimodal relation extraction (MRE) is a crucial task in the fields of Knowledge Graph and Multimedia, playing a pivotal role in multimodal knowledge graph construction. However, existing methods are typically limited to extracting a single type of relational triplet, which restricts their ability to extract triplets beyond the specified types. Directly combining these methods fails to capture dynamic cross-modal interactions and introduces significant computational redundancy. Therefore, we propose a novel \textit{unified multimodal Relation Extraction framework with Multilevel Optimal Transport and mixture-of-Experts}, termed REMOTE, which can simultaneously extract intra-modal and inter-modal relations between textual entities and visual objects. To dynamically select optimal interaction features for different types of relational triplets, we introduce mixture-of-experts mechanism, ensuring the most relevant modality information is utilized. Additionally, considering that the inherent property of multilayer sequential encoding in existing encoders often leads to the loss of low-level information, we adopt a multilevel optimal transport fusion module to preserve low-level features while maintaining multilayer encoding, yielding more expressive representations. Correspondingly, we also create a Unified Multimodal Relation Extraction (UMRE) dataset to evaluate the effectiveness of our framework, encompassing diverse cases where the head and tail entities can originate from either text or image. Extensive experiments show that REMOTE effectively extracts various types of relational triplets and achieves state-of-the-art performanc on almost all metrics across two other public MRE datasets. We release our resources at https://github.com/Nikol-coder/REMOTE.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction
A multimodal relation extraction system that models text and image features as Gaussian distributions, denoises them with uncertainty-aware contrastive learning, and aligns the distributions with symmetric KL achieves...
Reference graph
Works this paper leans on
-
[1]
Christoph Alt, Aleksandra Gabryszak, and Leonhard Hennig. 2020. TACRED Revisited: A Thorough Evaluation of the TACRED Relation Extraction Task. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, ...
doi:10.18653/v1/ 2020
-
[2]
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Rep...
arXiv 2025
-
[3]
Yue Cao, Yangzhou Liu, Zhe Chen, Guangchen Shi, Wenhai Wang, Danhuai Zhao, and Tong Lu. 2024. MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding. arXiv:2410.11829 [cs.CV] https://arxiv.org/abs/2410.11829
arXiv 2024
-
[4]
Xiang Chen, Ningyu Zhang, Lei Li, Shumin Deng, Chuanqi Tan, Changliang Xu, Fei Huang, Luo Si, and Huajun Chen. 2022. Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph Completion. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval(Madrid, Spain)(SIGIR ’22). Association f...
arXiv 2022
-
[5]
Shiyao Cui, Jiangxia Cao, Xin Cong, Jiawei Sheng, Quangang Li, Tingwen Liu, and Jinqiao Shi. 2024. Enhancing Multimodal Entity and Relation Extraction With Variational Information Bottleneck.IEEE/ACM Trans. Audio, Speech and Lang. Proc.32 (Jan. 2024), 1274–1285. doi:10.1109/TASLP.2023.3345146
-
[6]
Marco Cuturi. 2013. Sinkhorn distances: lightspeed computation of optimal transport. InProceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2(Lake Tahoe, Nevada)(NIPS’13). Curran Associates Inc., Red Hook, NY, USA, 2292–2300
work page 2013
-
[7]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Jill Burstein, Christy...
2019
-
[8]
George Doddington, Alexis Mitchell, Mark Przybocki, Lance Ramshaw, Stephanie Strassel, and Ralph Weischedel. 2004. The Automatic Content Extraction (ACE) Program – Tasks, Data, and Evaluation. InProceedings of the Fourth International Conference on Language Resources and Evaluation (LREC’04), Maria Teresa Lino, Maria Francisca Xavier, Fátima Ferreira, Rut...
work page 2004
Show all 47 references
-
[9]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognit...
2021 arXiv
-
[10]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Schelten, et al. 2024. The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https://arxiv.org/abs/2407.21783
2024 arXiv
-
[11]
Tianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou
-
[12]
Liang He, Hongke Wang, Yongchang Cao, Zhen Wu, Jianbing Zhang, and Xinyu Dai. 2023. MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark Evaluation. InProceedings of the 31st ACM International Confer- ence on Multimedia(Ottawa ON, Canada)(MM ’23). Asso...
2023
-
[13]
Liang He, Hongke Wang, Zhen Wu, Jianbing Zhang, Xinyu Dai, and Jiajun Chen. 2024. Focus & Gating: A Multimodal Approach for Unveiling Relations in Noisy Social Media. InProceedings of the 32nd ACM International Conference on Multimedia(Melbourne VIC, Australia)(MM ’24). Associ...
2024
-
[14]
Xuming Hu, Junzhe Chen, Aiwei Liu, Shiao Meng, Lijie Wen, and Philip S. Yu
-
[15]
Xuming Hu, Zhijiang Guo, Zhiyang Teng, Irwin King, and Philip S. Yu. 2023. Multimodal Relation Extraction with Cross-Modal Retrieval and Synthesis. In Proceedings of the 61st Annual Meeting of the Association for Computational Lin- guistics (Volume 2: Short Papers), Anna Roger...
2023 doi
-
[16]
Xuming Hu, Zhijiang Guo, Zhiyang Teng, Irwin King, and Philip S. Yu. 2023. Multimodal Relation Extraction with Cross-Modal Retrieval and Synthesis
2023
-
[17]
Abdelwahed Khamis, Russell Tsuchida, Mohamed Tarek, Vivien Rolland, and Lars Petersson. 2024. Scalable Optimal Transport Methods in Machine Learning: A Contemporary Survey.IEEE Transactions on Pattern Analysis and Machine Intelligence(2024), 1–20. doi:10.1109/TPAMI.2024.3379571
2024
-
[18]
Jinyuan Li, Han Li, Zhuo Pan, Di Sun, Jiahao Wang, Wenkun Zhang, and Gang Pan. 2023. Prompting ChatGPT in MNER: Enhanced Multimodal Named Entity Recognition with Auxiliary Refined Knowledge. InFindings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor...
2023 doi
-
[19]
Jinyuan Li, Han Li, Di Sun, Jiahao Wang, Wenkun Zhang, Zan Wang, and Gang Pan. 2024. LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition. InFindings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikuma...
2024 doi
-
[20]
Jinyuan Li, Ziyan Li, Han Li, Jianfei Yu, Rui Xia, Di Sun, and Gang Pan. 2024. Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation. arXiv:2406.07268 [cs.MM] https: //arxiv.org/abs/2406.07268
2024 arXiv
-
[21]
Lei Li, Xiang Chen, Shuofei Qiao, Feiyu Xiong, Huajun Chen, and Ningyu Zhang
-
[22]
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang
-
[23]
Qian Li, Shu Guo, Cheng Ji, Xutan Peng, Shiyao Cui, Jianxin Li, and Lihong Wang. 2023. Dual-Gated Fusion with Prefix-Tuning for Multi-Modal Relation Extraction. InFindings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki O...
2023 doi
-
[24]
On analyzing the role of image for visual-enhanced relation extraction (student abstract). InProceedings of the Thirty-Seventh AAAI Conference on Arti- ficial Intelligence and Thirty-Fifth Conference on Innovative Applications of Arti- ficial Intelligence and Thirteenth Sympos...
1918 doi
-
[25]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual in- struction tuning. InProceedings of the 37th International Conference on Neural Information Processing Systems(New Orleans, LA, USA)(NIPS ’23). Curran Asso- ciates Inc., Red Hook, NY, USA, Article 1516, 25 pages
2023
-
[26]
arXiv:1908.03557 [cs.CV] https://arxiv.org/abs/1908.03557
VisualBERT: A Simple and Performant Baseline for Vision and Language. arXiv:1908.03557 [cs.CV] https://arxiv.org/abs/1908.03557
1908 arXiv
-
[27]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. arXiv:1711.05101 [cs.LG] https://arxiv.org/abs/1711.05101
2019 arXiv
-
[28]
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. 2024. LLaVA-NeXT: Improved reasoning, OCR, and world knowl- edge. https://llava-vl.github.io/blog/2024-01-30-llava-next/
2024
-
[29]
Sina Moradi. 2025. A Survey on Algorithmic Developments in Optimal Transport Problem with Applications. arXiv:2501.06247 [cs.DS] https://arxiv.org/abs/2501. 06247
2025 arXiv
-
[30]
Xiyang Liu, Chunming Hu, Richong Zhang, Kai Sun, Samuel Mensah, and Yongyi Mao. 2024. Multimodal Relation Extraction via a Mixture of Hierarchical Visual Context Learners. InProceedings of the ACM Web Conference 2024(Singapore, Singapore)(WWW ’24). Association for Computing Ma...
2024
-
[31]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models . In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Computer Society, Los Alamitos, CA, USA, 10...
2022
-
[32]
Wen Luo, Yu Xia, Shen Tianshu, and Sujian Li. 2024. Shapley Value-based Contrastive Alignment for Multimodal Information Extraction. InProceedings of the 32nd ACM International Conference on Multimedia(Melbourne VIC, Australia) (MM ’24). Association for Computing Machinery, Ne...
2024
-
[33]
Shengqiong Wu, Hao Fei, Yixin Cao, Lidong Bing, and Tat-Seng Chua. 2023. Infor- mation Screening whilst Exploiting! Multimodal Relation Extraction with Feature Denoising and Multimodal Topic Modeling. InProceedings of the 61st Annual Meet- ing of the Association for Computatio...
2023 doi
-
[34]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. InProceedings ...
2021
-
[35]
Xinjie Yang, Xiaocheng Gong, Binghao Tang, Yang Lei, Yayue Deng, Huan Ouyang, Gang Zhao, Lei Luo, Yunling Feng, Bin Duan, Si Li, and Yajing Xu
-
[36]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. Qwen2-VL: Enhancing Vision-Language Mode...
2024 arXiv
-
[37]
Zefan Zhang, Weiqi Zhang, Yanhui Li, and Tian Bai. 2024. Caption-Aware Multi- modal Relation Extraction with Mutual Information Maximization. InProceedings of the 32nd ACM International Conference on Multimedia(Melbourne VIC, Aus- tralia)(MM ’24). Association for Computing Mac...
2024
-
[38]
Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. 2024. Depth Anything V2. arXiv:2406.09414 [cs.CV] https://arxiv.org/abs/2406.09414
2024 arXiv
-
[39]
Changmeng Zheng, Junhao Feng, Yi Cai, Xiaoyong Wei, and Qing Li. 2023. Re- thinking Multimodal Entity and Relation Extraction from a Translation Point of View. InProceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), ...
2023 doi
-
[40]
Changmeng Zheng, Junhao Feng, Ze Fu, Yi Cai, Qing Li, and Tao Wang. 2021. Multimodal Relation Extraction with Efficient Graph Alignment. InProceedings of the 29th ACM International Conference on Multimedia(Virtual Event, China) (MM ’21). Association for Computing Machinery, Ne...
2021
-
[41]
Yaodong Yu, Tianzhe Chu, Shengbang Tong, Ziyang Wu, Druv Pai, Sam Buchanan, and Yi Ma. 2024. Emergence of Segmentation with Minimalistic White-Box Transformers. InConference on Parsimony and Learning (Proceedings of Machine Learning Research, Vol. 234), Yuejie Chi, Gintare Kar...
2024
-
[42]
Zihao Zheng, Tao He, Ming Liu, Zhongyuan Wang, Ruiji Fu, and Bing Qin. 2024. Relational Graph-Bridged Image-Text Interaction: A Novel Method for Multi- Modal Relation Extraction. InICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICA...
2024
-
[43]
Qihui Zhao, Tianhan Gao, and Nan Guo. 2023. TSVFN: Two-Stage Visual Fusion Network for multimodal relation extraction.Information Processing & Manage- ment60, 3 (2023), 103264. doi:10.1016/j.ipm.2023.103264
2023
-
[46]
Changmeng Zheng, Zhiwei Wu, Junhao Feng, Ze Fu, and Yi Cai. 2021. MNRE: A Challenge Multimodal Dataset for Neural Relation Extraction with Visual Evi- dence in Social Media Posts. In2021 IEEE International Conference on Multimedia and Expo (ICME). 1–6. doi:10.1109/ICME51207.20...
2021 arXiv
-
[2019]
InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Pro- cessing (EMNLP-IJCNLP)
FewRel 2.0: Towards More Challenging Few-Shot Relation Classification. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Pro- cessing (EMNLP-IJCNLP). Association for Computati...
2019 doi
-
[2023]
InProceedings of the 31st ACM International Confer- ence on Multimedia(Ottawa ON, Canada)(MM ’23)
Prompt Me Up: Unleashing the Power of Alignments for Multimodal Entity and Relation Extraction. InProceedings of the 31st ACM International Confer- ence on Multimedia(Ottawa ON, Canada)(MM ’23). Association for Computing Machinery, New York, NY, USA, 5185–5194. doi:10.1145/358...
-
[2024]
InProceedings of the 33rd ACM International Conference on Information and Knowledge Management(Boise, ID, USA)(CIKM ’24)
CAG: A Consistency-Adaptive Text-Image Alignment Generation for Joint Multimodal Entity-Relation Extraction. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management(Boise, ID, USA)(CIKM ’24). Association for Computing Machinery, New York,...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.