pith. sign in

arxiv: 2112.03521 · v1 · pith:TAZZJGHFnew · submitted 2021-12-07 · 💻 cs.CL · cs.AI

UNITER-Based Situated Coreference Resolution with Rich Multimodal Input

classification 💻 cs.CL cs.AI
keywords multimodalobjectdialogcoreferencemodelresolutionrichcurrent
0
0 comments X
read the original abstract

We present our work on the multimodal coreference resolution task of the Situated and Interactive Multimodal Conversation 2.0 (SIMMC 2.0) dataset as a part of the tenth Dialog System Technology Challenge (DSTC10). We propose a UNITER-based model utilizing rich multimodal context such as textual dialog history, object knowledge base and visual dialog scenes to determine whether each object in the current scene is mentioned in the current dialog turn. Results show that the proposed approach outperforms the official DSTC10 baseline substantially, with the object F1 score boosted from 36.6% to 77.3% on the development set, demonstrating the effectiveness of the proposed object representations from rich multimodal input. Our model ranks second in the official evaluation on the object coreference resolution task with an F1 score of 73.3% after model ensembling.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LARE: Low-Attention Region Encoding for Text-Image Retrieval

    cs.CV 2026-06 unverdicted novelty 5.0

    LARE uses parallel encoding of full images and low-attention regions to improve text-image retrieval, shown on a new Dense-Set subset of COCO and Flickr30K with re-captioned overlooked areas.