Pith. sign in

3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is designed to leverage insights from both fine-grained and coarse-grained levels by facilitating a nuanced correlation between token and entity representations, addressing the complexities inherent in form documents. Additionally, we introduce new inter-grained and cross-grained loss functions to further refine diverse multi-teacher knowledge distillation transfer process, presenting distribution gaps and a harmonised understanding of form documents. Through a comprehensive evaluation across publicly available form document understanding datasets, our proposed model consistently outperforms existing baselines, showcasing its efficacy in handling the intricate structures and content of visually complex form documents.

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding

cs.CV · 2025-06-02 · conditional · novelty 5.0

The VRD-IU competition results show that hierarchical decomposition and pretrained multimodal transformers are effective for form key-information extraction and localization, but the paper lacks baseline comparisons and contains internal inconsistencies.

citing papers explorer

Showing 1 of 1 citing paper.

  • VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding cs.CV · 2025-06-02 · conditional · none · ref 4 · internal anchor

    The VRD-IU competition results show that hierarchical decomposition and pretrained multimodal transformers are effective for form key-information extraction and localization, but the paper lacks baseline comparisons and contains internal inconsistencies.