Pith. sign in

REVIEW 1 cited by

Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.14712 v2 pith:4FPAS4KS submitted 2022-03-28 cs.CV

classification cs.CV
keywords actionassembly101datasetunderstandingactivityinvestigatemistakesmulti-view
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101 is the first multi-view action dataset, with simultaneous static (8) and egocentric (4) recordings. Sequences are annotated with more than 100K coarse and 1M fine-grained action segments, and 18M 3D hand poses. We benchmark on three action understanding tasks: recognition, anticipation and temporal segmentation. Additionally, we propose a novel task of detecting mistakes. The unique recording format and rich set of annotations allow us to investigate generalization to new toys, cross-view transfer, long-tailed distributions, and pose vs. appearance. We envision that Assembly101 will serve as a new challenge to investigate various activity understanding problems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

    cs.CV 2026-08 conditional novelty 6.0 of 10

    LVLM judges are near chance when asked to pick the correctly ordered version of an image sequence, and this temporal blindness persists after fine-tuning and at larger scale.

Pith tools