UNITER-Based Situated Coreference Resolution with Rich Multimodal Input

Yichen Huang; Yik-Cheung Tam; Yuchen Wang

arxiv: 2112.03521 · v1 · pith:TAZZJGHFnew · submitted 2021-12-07 · 💻 cs.CL · cs.AI

UNITER-Based Situated Coreference Resolution with Rich Multimodal Input

Yichen Huang , Yuchen Wang , Yik-Cheung Tam This is my paper

classification 💻 cs.CL cs.AI

keywords multimodalobjectdialogcoreferencemodelresolutionrichcurrent

0 comments

read the original abstract

We present our work on the multimodal coreference resolution task of the Situated and Interactive Multimodal Conversation 2.0 (SIMMC 2.0) dataset as a part of the tenth Dialog System Technology Challenge (DSTC10). We propose a UNITER-based model utilizing rich multimodal context such as textual dialog history, object knowledge base and visual dialog scenes to determine whether each object in the current scene is mentioned in the current dialog turn. Results show that the proposed approach outperforms the official DSTC10 baseline substantially, with the object F1 score boosted from 36.6% to 77.3% on the development set, demonstrating the effectiveness of the proposed object representations from rich multimodal input. Our model ranks second in the official evaluation on the object coreference resolution task with an F1 score of 73.3% after model ensembling.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

LARE: Low-Attention Region Encoding for Text-Image Retrieval
cs.CV 2026-06 unverdicted novelty 5.0

LARE uses parallel encoding of full images and low-attention regions to improve text-image retrieval, shown on a new Dense-Set subset of COCO and Flickr30K with re-captioned overlooked areas.