Pith. sign in

REVIEW 1 cited by

Training Multimedia Event Extraction With Generated Images and Captions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.08966 v2 pith:Z5VHJ6XI submitted 2023-06-15 cs.MM cs.CV

classification cs.MMcs.CV
keywords trainingdatamultimediaeventcamelgeneratedimagemultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contemporary news reporting increasingly features multimedia content, motivating research on multimedia event extraction. However, the task lacks annotated multimodal training data and artificially generated training data suffer from distribution shift from real-world data. In this paper, we propose Cross-modality Augmented Multimedia Event Learning (CAMEL), which successfully utilizes artificially generated multimodal training data and achieves state-of-the-art performance. We start with two labeled unimodal datasets in text and image respectively, and generate the missing modality using off-the-shelf image generators like Stable Diffusion and image captioners like BLIP. After that, we train the network on the resultant multimodal datasets. In order to learn robust features that are effective across domains, we devise an iterative and gradual training strategy. Substantial experiments show that CAMEL surpasses state-of-the-art (SOTA) baselines on the M2E2 benchmark. On multimedia events in particular, we outperform the prior SOTA by 4.2% F1 on event mention identification and by 9.8% F1 on argument identification, which indicates that CAMEL learns synergistic representations from the two modalities. Our work demonstrates a recipe to unleash the power of synthetic training data in structured prediction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding

    cs.CV 2024-12 conditional novelty 5.0 of 10

    POBF paints new backgrounds around preserved objects to synthesize visual-grounding training data and filters those samples with teacher-model scores, improving accuracy by 5.83% over real-only training.

Pith tools