Pith. sign in

REVIEW 1 cited by

A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.09606 v1 pith:KBR33UIV submitted 2020-05-19 cs.CL

classification cs.CL
keywords instructionsrecipesdifferentdishsametextvideoacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many high-level procedural tasks can be decomposed into sequences of instructions that vary in their order and choice of tools. In the cooking domain, the web offers many partially-overlapping text and video recipes (i.e. procedures) that describe how to make the same dish (i.e. high-level task). Aligning instructions for the same dish across different sources can yield descriptive visual explanations that are far richer semantically than conventional textual instructions, providing commonsense insight into how real-world procedures are structured. Learning to align these different instruction sets is challenging because: a) different recipes vary in their order of instructions and use of ingredients; and b) video instructions can be noisy and tend to contain far more information than text instructions. To address these challenges, we first use an unsupervised alignment algorithm that learns pairwise alignments between instructions of different recipes for the same dish. We then use a graph algorithm to derive a joint alignment between multiple text and multiple video recipes for the same dish. We release the Microsoft Research Multimodal Aligned Recipe Corpus containing 150K pairwise alignments between recipes across 4,262 dishes with rich commonsense information.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Recipe-derived object status phrases added to vision-language matching improve recipe-step prediction by 20 to 26 accuracy points on instructional and real-world non-visual cooking videos.

Pith tools