Pith. sign in

REVIEW 3 major objections 6 minor 8 cited by

HD-EPIC: A Highly-Detailed Egocentric Video Dataset

T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read HD-EPIC is a 41.3-hour unscripted egocentric kitchen-video dataset whose annotations—recipes, ingredient nutrition, fine-grained actions, audio, gaze, and 3D digital twins—reach lab-grade detail, and whose 26,650-question VQA benchmark…

desk verdict A dense, genuinely novel egocentric dataset that deserves serious review; the 'lab-matching' claim overreaches but the resource is solid. read the letter →

arxiv 2502.04144 v2 pith:UQDDULOW submitted 2025-02-06 cs.CV

classification cs.CV
keywords egocentricvideodatasetdigitaltwinquestionansweringactionrecognitiongazeprimingnutritiontrackingobjectaudioannotation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

HD-EPIC is a dataset of 41.3 hours of unscripted egocentric video recorded in nine home kitchens, annotated at a depth previously reserved for controlled labs. Over 263 annotations per minute link recipe steps, weighed ingredients with nutrition, fine-grained actions with how and why clauses, audio events, object movements, gaze, and a labelled 3D digital twin of each kitchen. The paper uses these labels to build a 26,650-question video QA benchmark on which the strongest tested model, Gemini Pro, scores 37.6% versus a 90.3% human baseline. The central claim is that this is the first in-the-wild dataset whose annotation detail matches lab environments.

What carries the argument

The load-bearing mechanism is the interconnected annotation hierarchy anchored by a digital twin of each kitchen. A custom-built 3D model of every fixture (cupboard, drawer, counter, appliance) is created from multi-day SLAM point clouds, and object masks, camera poses, and eye gaze are all expressed in that shared 3D frame. This lets the pipeline associate actions with exact fixtures, lift 2D object bounding boxes to 3D locations, and compute gaze-priming times for objects before pickup or put-down. The recipe prep/step structure and per-ingredient weighing-and-adding segments then tie fine-grained actions to long-horizon nutrition tracking, which is what makes the VQA benchmark's cross-scale questions possible.

What would settle it

Take a random sample of 50 filmed kitchen sessions, have independent annotators watch the video without the participant narrations and write their own step-by-step action lists and ingredient logs, then measure agreement with HD-EPIC's labels on action boundaries, recipe steps, and ingredient quantities; a large disagreement rate in either direction would show the self-report ground truth is not stable across observers.

Watch

Extended reading notes

Core claim

The paper's claim is that HD-EPIC closes the gap between unscripted egocentric video and lab-quality annotation density. Concretely, the dataset provides 69 recipes with temporally aligned steps and prep segments, 59,454 fine-grained actions parsed into verbs, nouns, hands, and how and why clauses, 50,968 audio events, 19,900 object movement tracks with 36,900 bounding boxes lifted to 3D using per-kitchen digital twins of 413 fixtures, and 7.7 million hand masks. Every annotation is interconnected through the digital twin and gaze, so questions can span from single frames to over seven hours of footage. On the resulting VQA benchmark of 30 question prototypes and 26,650 multiple-choice questions, the strongest long-context model reaches 37.6% accuracy while a sampled human baseline reaches 90.3%, which the authors present as evidence that the benchmark tests genuine video understanding.

Load-bearing premise

The whole annotation hierarchy hangs on participants' self-reported narrations, recipes, and nutrition logs, with no independent check that what they said matches what the video shows.

Editorial extensions

If this is right

  • Any vision-language model claiming long-video understanding can now be tested on questions requiring multi-hop object itineraries, ingredient-order reasoning, and nutrition changes, not just narration retrieval.
  • Action recognition models trained on existing egocentric kitchen datasets drop sharply on HD-EPIC, so the dataset can serve as a zero-shot validation set for generalization to unseen homes, recipes, and recording conditions.
  • Because all annotations share a 3D frame, benchmarks on 3D perception, gaze, and object-fixture interaction can run on the same videos as action, audio, and VQA benchmarks, enabling cross-task analysis.
  • The 90.3% human versus 37.6% model gap on the VQA benchmark sets a concrete accuracy target for future long-context video models to be measured against.
  • The 26,650-question benchmark is built from dense annotations that also support an estimated upper bound of 100,000 possible unique questions, so the evaluation space can be expanded without new video collection.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The gaze-priming statistics (94.8% of feasible pickups primed about four seconds before the action) suggest that gaze could serve as a self-supervised training signal for action anticipation, a direction the paper benchmarks but does not train on.
  • Because all annotations share a 3D frame, one could construct cross-day questions such as “where was this object yesterday” to test episodic memory beyond a single video, which the current VQA set does not fully explore.
  • The reliance on participant self-reports means the dataset inherits any systematic bias in what people choose to narrate or log; an independent verification study on a subset of videos would quantify that bias and could be a valuable follow-up contribution.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces HD-EPIC, a 41.3-hour egocentric video dataset recorded in nine home kitchens over multiple days, with digital twins of 413 fixtures, 69 annotated recipes, 59,454 fine-grained action segments, 50,968 audio events, 19.9K object movement tracks, and 37K object masks lifted to 3D. It also presents a 26,650-question VQA benchmark spanning recipe, ingredient, nutrition, fine-grained action, 3D perception, object motion, and gaze categories, plus action recognition, sound recognition, and long-term video object segmentation benchmarks. The central claim is that HD-EPIC is the first in-the-wild dataset with annotations matching those produced in controlled lab environments, supported by multi-tiered annotation, gaze-based object priming, and 3D grounding.

Significance. If the annotation-quality concerns are addressed, this is a substantial contribution to egocentric vision: the dataset is unusually dense, the annotations are interconnected (recipe steps, ingredients, nutrition, actions, gaze, 3D fixtures, and object tracks), and the release includes videos, digital twins, annotations, and benchmark protocols. The paper's evaluation is also appropriately cautious in several respects: the VQA models are external and not trained on HD-EPIC, language-only baselines are included, and a ground-truth narration ablation (Supp. Table A4) demonstrates that the questions are not answerable from text alone. The explicit quality checks that do exist, such as inter-annotator agreement for action boundaries, manual verification of object-fixture assignments, and OCR verification of scale readings, are concrete strengths. However, the load-bearing claim of equivalence with controlled-lab annotation quality currently rests on unverified participant self-reports for recipes, nutrition, and narration content, and the benchmark conclusions are reported without confidence intervals.

major comments (3)
  1. [Sec. 3; Supp. B.2-B.3; Sec. 4.4] Recipe composition, step ordering, ingredient quantities, and what/how/why narrations are collected from participant self-reports, with no independent verification beyond OCR of weighing-scale displays (Supp. C.1). No inter-annotator agreement or second-source check is reported for recipe, nutrition, or narration labels. Because these labels propagate into the 59,454 action segments, 69 recipes, and 26,650 VQA questions, the claim in Sec. 4.4 that HD-EPIC has 'detailed annotations matching those in controlled lab environments' is not yet evidenced. Please add either an independent verification study (e.g., third-party review of a random sample of recipe steps, ingredient-addition times, and narration content against the video) or revise the claim to distinguish verified annotations from self-reported ones, and report verification statistics.
  2. [Sec. 5.1; Tables 2 and A4; Fig. 10] VQA accuracies are reported as point estimates without confidence intervals, despite prototype sample sizes ranging from 50 to 1,000 questions. Some headline comparisons, such as blind Llama 3.2 at 26.5% versus blind Gemini Pro at 26.7%, are within plausible sampling variation, and per-prototype differences of a few points are not interpretable without intervals. Please report binomial or bootstrap confidence intervals for each prototype and for the aggregate, and state the exact number of questions behind each reported average.
  3. [Supp. D.3; Table 2] The 90.3% human baseline is computed from 600 questions, 20 per prototype, from 3 participants. This sample is too small to support per-prototype conclusions; for a 20-question prototype the 95% confidence interval is roughly ±13 percentage points. Please expand the human evaluation or report per-prototype human accuracy with confidence intervals, and clarify how the 600 questions were assigned to participants and whether each participant answered all prototypes.
minor comments (6)
  1. [Abstract vs. Sec. 5.1 and Table 2] The Gemini Pro accuracy is reported as 38.5% in the header abstract but 37.6% in the main abstract and Table 2; Table A4 shows 38.5% for the video+Q+A configuration. Please harmonize the number and state which configuration is being reported.
  2. [Abstract] The term 'HDEPIC' should be 'HD-EPIC' for consistency with the dataset name.
  3. [Supp. C.3, Fig. A10 caption] The word 'corectly' should be 'correctly'.
  4. [Supp. D.2] The phrase 'far from solves these categories' should read 'far from solving these categories'.
  5. [Table 1 and Table A1] The 'Fully annotated' column would benefit from a definition; as written, it is unclear whether it refers to all listed annotation types or specifically to 3D/pose coverage.
  6. [Sec. 5.2, Table 4] The statement that 'audio is not sufficiently robust to new scenes or devices' is stronger than the evidence, since HD-EPIC differs jointly in scene, device, and recording conditions; consider rephrasing to note that the cross-dataset drop is consistent with multiple confounds.

Circularity Check

1 steps flagged · score 4.0 of 10

Long-term VOS hand benchmark is partly circular: SAM2 is evaluated on hand masks that SAM2 itself generated; the central VQA benchmark is independent.

  1. self definitional [Sec. 4.3 (Hand Masks), Sec. 5.3 (Long-Term VOS), Supp. C.3]
    "Hand Masks. We annotate a handful of frames per video for both hands. ... In total, our dataset contains 7.7M hand masks: 3.9M right and 3.8M left of which 11K are manually annotated. ... For each video, we utilise SAM2 [64] to predict 2D hand masks for every frame. ... We construct a long-term video object segmentation benchmark using our segmentations and track associations (Sec. 4.3). ... We evaluate two models [12, 64] with a naive baseline."

    The VOS benchmark's hand ground truth is not independent of the model being evaluated. The dataset contains 7.7M hand masks, of which only roughly 20.5K (about 0.3%) are manually annotated or corrected; the rest are SAM2's own per-frame predictions. The benchmark then evaluates SAM2 against 'our segmentations', i.e., largely SAM2's own output. The reported result that SAM2 achieves 89.1 J&F for hands, and the conclusion that SAM2 surpasses Cutie for hands, therefore reduce partly to SAM2 agreeing with itself. Cutie is scored on labels produced by its competitor, and the manual-correction rate is too small to make the hand benchmark an independent test.

full rationale

The central VQA benchmark is not circular: it is evaluated on external models not trained on HD-EPIC, it includes blind language-only baselines (26.5/26.7%), and the input ablation (Supp. Table A4) shows that GT narrations alone reach at most 50.8% versus a 90.3% human baseline, so the questions are not answerable from the annotations used to create them. Self-citations to EPIC-KITCHENS [18] and EPIC-SOUNDS [31] supply methodology and cluster/class taxonomies; these are published artifacts used for cross-dataset evaluation, not load-bearing derivations of the paper's claims. The one genuinely circular step is in the Long-Term VOS benchmark for hands: the hand ground-truth masks are predominantly generated by SAM2 itself, and SAM2 is then evaluated against those same masks, making its reported hand-segmentation advantage over Cutie partly self-agreement by construction. The headline 'matching controlled lab environments' claim is an unsupported validity assertion, but that is a correctness gap rather than a circularity.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

As a dataset paper, the ledger records design thresholds and domain assumptions rather than fitted derivation constants. The listed free parameters are hand-set values that shape the benchmark difficulty and the gaze priming statistics; the axioms are the unverified premises about self-reported ground truth, SLAM accuracy, gaze attention, and audio class coverage on which the dataset's claimed label quality rests.

free parameters (5)
  • Gaze association window = 1 second
    Sec. 4.3: the fixture with the highest cumulative gaze intersection in the 1s before a narration is selected as the interacted fixture. This threshold, taken from prior literature [40], is applied uniformly across kitchens and actions.
  • Gaze priming exclusion time = 10 seconds
    Sec. 4.3: objects whose pick-up location is already close to gaze 10s earlier are excluded from priming. This hand-set cutoff filters 29.40% of start locations and shapes all priming statistics.
  • VQA weight negative multiplier range = (0.5, 5)
    Supp. D.1.2: negatives for ingredient weight and exact quantity are random multipliers of the correct value in the range (0.5, 5). This choice directly controls the difficulty of these prototypes.
  • Nutrition negative separation = 20%
    Supp. D.1.3: image nutrition estimation negatives require a nutritional value difference of at least 20% from the positive answer, defining the hardness of nutrition questions.
  • Object movement counting filters = 1.1 seconds and 20 cm
    Supp. D.1.6: consecutive object tracks are counted as distinct movements only if separated by at least 1.1 seconds and with spatial displacement of at most 20 cm. These ad hoc thresholds filter the ground truth and affect the object motion benchmark.
assumptions (4)
  • domain assumption Participant self-narrations and self-reported recipes and nutrition are accurate ground truth
    Sec. 3: all action descriptions, how/why clauses, recipe steps, prep segments, ingredient weights and adding times derive from participants reviewing their own videos and from MyFitnessPal logs. There is no independent verification of recipe content or quantities.
  • domain assumption The SLAM point cloud and camera poses provide a sufficiently accurate 3D scene for grounding annotations
    Sec. 4.3: object masks are lifted to 3D using MPS depth and sparse correspondences, and fixture associations depend on the reconstructed point cloud. Three videos failed SLAM and were aligned via COLMAP, and no metric accuracy of the digital twin against ground truth is reported.
  • domain assumption Gaze direction projected into 3D indicates visual attention and primes object interactions
    Sec. 4.3: the paper assumes that the wearer's gaze intersecting an object location marks priming before pick-up or place-down, with a 1 second fixation-to-interaction link from [40]. This assumption underpins all gaze priming statistics and the gaze-based VQA questions.
  • domain assumption The 44 audio classes from EPIC-Sounds are sufficient for HD-EPIC audio annotation
    Sec. 4.2 and Supp. C.2.4: sound annotations use the existing EPIC-Sounds class list, and the authors state no new classes were warranted. This assumes the kitchen sound vocabulary of EPIC-Sounds fully covers these new homes and devices.

how reviews work

0 comments
Cite this review

Pith. "Pith review of HD-EPIC: A Highly-Detailed Egocentric Video Dataset." pith.science (2026). https://pith.science/paper/UQDDULOW

@misc{pith2026250204144,
  author       = {Pith},
  title        = {Pith review of: HD-EPIC: A Highly-Detailed Egocentric Video Dataset},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UQDDULOW}},
  note         = {Machine review of arXiv:2502.04144}
}
read the original abstract

We present a validation dataset of newly-collected kitchen-based egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, all annotations are grounded in 3D through digital twinning of the scene, fixtures, object locations, and primed with gaze. Footage is collected from unscripted recordings in diverse home environments, making HDEPIC the first dataset collected in-the-wild but with detailed annotations matching those in controlled lab environments. We show the potential of our highly-detailed annotations through a challenging VQA benchmark of 26K questions assessing the capability to recognise recipes, ingredients, nutrition, fine-grained actions, 3D perception, object motion, and gaze direction. The powerful long-context Gemini Pro only achieves 38.5% on this benchmark, showcasing its difficulty and highlighting shortcomings in current VLMs. We additionally assess action recognition, sound recognition, and long-term video-object segmentation on HD-EPIC. HD-EPIC is 41 hours of video in 9 kitchens with digital twins of 413 kitchen fixtures, capturing 69 recipes, 59K fine-grained actions, 51K audio events, 20K object movements and 37K object masks lifted to 3D. On average, we have 263 annotations per minute of our unscripted videos.

Figures

Figures reproduced from arXiv: 2502.04144 by the authors.

Figure 1
Figure 1. Annotation Highlights. We capture multi-day recordings of unscripted activities. Centre-Top: Recipes are recorded with steps and their preparation temporally annotated, along with ingredient addition. Ingredients are weighed and nutrition recorded. Centre-Middle: Dense fine-grained narrations detailing what, how, and why are parsed and clustered. Audio events are also annotated. Centre-Bottom: Object movements are t… view at source ↗
Figure 2
Figure 2. Diversity in HD-EPIC, which is filmed over 3 days in-the-wild, resulting in many objects, activities and recipes. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Recipe modification in ingredients and steps. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: For the ‘Carbonara’ recipe, we visualise the [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Frequency of verb clusters (top) and noun clusters (bottom) in narrated sentences by category, shown on a logarithmic scale. [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 6
Figure 6. Figure 6: Nutrition is monitored throughout recipes as ingredients [PITH_FULL_IMAGE:figures/full_fig_p004_6.png]
Figure 8
Figure 8. Figure 8: Priming Object Interaction Through Gaze. Top: Camera position with projected eye-gaze and object positions in 3D. Middle: 2D gaze location. Bottom: Timeline for priming ob￾ject movement e.g. the glass is primed 8.3s before taking. thus offer full annotations of all obj…
Figure 9
Figure 9. Figure 9: (Top) Priming Statistics for both start and end locations [PITH_FULL_IMAGE:figures/full_fig_p006_9.png]
Figure 11
Figure 11. Figure 11: VQA Results per Question Prototype. Our benchmark contains many challenging questions for current models. Model Recipe Ingredient Nutrition Action 3D Motion Gaze Avg. Blind - Language Only Llama 3.2 33.5 25.0 36.7 23.3 22.3 25.5 19.5 26.5 Gemini Pro 38.0 26.8 30.0 22.…
Figure 12
Figure 12. Figure 12: Effect of Input Length. Models struggle with ques￾tions of all video input lengths. s=second, m=minute, h=hour. Model Modality Verb Noun Action Unseen EPIC-100 Action EPIC-KITCHENS-100 SOTA TIM [9] A+V 77.1 67.2 57.5 44.6 HD-EPIC Chance - 10.9 1.8 0.0 - SlowFast [24] …
Figure 13
Figure 13. Figure 13: VQA Qualitative Results. We mark GT answers with a green background, and predictions from different models, i.e., LLaMA 3.2, VideoLLaMA 2, LongVA, Gemini Pro with coloured dots. Note: Under Nutrition, [fat] values are not provided to the model. Model Modality Top-1 To…
Figure 14
Figure 14. Figure 14: shows the results. SAM2 [64] surpasses Cutie [12] for hands, but does worse on objects. Overall, objects have added challenge in diversity in perspective, lighting, loca￾Total Hands Objects Model J F J &F J F J &F J F J &F Static 8.0 10.3 9.2 14.6 14.4 14.5 4.8 8.4 6.…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding

    cs.CV 2025-05 conditional novelty 7.0 of 10

    EgoExOR is a new multimodal, multi-perspective OR dataset with 84,553 annotated frames, plus a benchmark showing that fusing egocentric and exocentric signals improves surgical scene graph generation.

  2. VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

    cs.RO 2026-08 conditional novelty 6.0 of 10

    From 204K egocentric human videos, the authors automatically extract visual, grasp, and trajectory affordances and train one vision-language model, VLAff, that predicts all three for robot manipulation.

  3. EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    EgoEverything is a new benchmark for long-context egocentric video understanding that uses human gaze-based attention signals to generate questions reflecting natural behavior.

  4. Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 3B agent trained with supervised tool-use traces and reinforcement learning answers week-long egocentric video questions by dynamically selecting hierarchical retrieval, video-LLM, and VLM tools.

  5. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AVHaystacks is a new 3100-question benchmark for audio-visual QA across 500 videos, and the MAGNET multi-agent pipeline beats current baselines on it.

  6. Enhancing Wearable Tap Water Audio Detection through Subclass Annotation in the HD-Epic Dataset

    cs.HC 2025-05 conditional novelty 6.0 of 10

    This paper adds 717 precisely timed tap water audio annotations to the HD-Epic dataset and reports that lightweight classifiers detect this subclass better than the broader water class when measured against a random baseline.

  7. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

  8. Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research

    cs.RO 2025-06 accept novelty 1.0 of 10

    A perspective article reviews the state of using foundation models for laboratory automation and proposes a roadmap for fully autonomous experiments.

Reference graph

Works this paper leans on

93 extracted references · 64 canonical work pages · cited by 8 Pith papers

  1. [1]

    https : / / facebookresearch

    Project aria machine perception services. https : / / facebookresearch . github . io / projectaria _ tools/docs/ARK/mps. 3, 2

  2. [2]

    https://www.blender.org/

    Blender. https://www.blender.org/. 5

  3. [3]

    https://www.myfitnesspal.com/

    My fitness app. https://www.myfitnesspal.com/. 3, 1

  4. [4]

    https://www.robots.ox.ac.uk/ vgg/software/lisa/

    VGG List Annotator (LISA) , 2022. https://www.robots.ox.ac.uk/ vgg/software/lisa/. 5

  5. [5]

    GPT-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ah- mad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. GPT-4 technical report. arXiv preprint arXiv:2303.08774 ,

  6. [6]

    Flamingo: a visual language model for few-shot learning

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems (NeurIPS), 2022. 2

  7. [7]

    HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos

    Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Fan Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, Jakob Ju- lian Engel, and Tomas Hodan. HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP...

  8. [8]

    TemporalBench: Towards fine-grained temporal understanding for multimodal video models

    Mu Cai, Reuben Tan, Jianrui Zhang, Bocheng Zou, Kai Zhang, Feng Yao, Fangrui Zhu, Jing Gu, Yiwu Zhong, Yuzhang Shang, Yao Dou, Jaden Park, Jianfeng Gao, Yong Jae Lee, and Jianwei Yang. TemporalBench: Towards fine-grained temporal understanding for multimodal video models. arXiv preprint arXiv:2410.10818, 2024. 2

Show all 93 references
  1. [9]

    TIM: A Time Interval Machine for Audio-Visual Action Recognition

    Jacob Chalk, Jaesung Huh, Evangelos Kazakos, Andrew Zis- serman, and Dima Damen. TIM: A Time Interval Machine for Audio-Visual Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 7, 8, 15

  2. [10]

    ReXTime: A benchmark suite for reasoning-across-time in videos

    Jr-Jen Chen, Yu-Chien Liao, Hsi-Che Lin, Yu-Chu Yu, Yen- Chun Chen, and Yu-Chiang Frank Wang. ReXTime: A benchmark suite for reasoning-across-time in videos. In Ad- vances in Neural Information Processing Systems (NeurIPS),

  3. [11]

    InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, et al. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF Conference on Computer Vi-...

  4. [12]

    Putting the object back into video object segmentation

    Ho Kei Cheng, Seoung Wug Oh, Brian Price, Joon-Young Lee, and Alexander Schwing. Putting the object back into video object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 8, 15

  5. [13]

    VideoLLaMA 2: Advancing Spatial- Temporal Modeling and Audio Understanding in Video- LLMs

    Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. VideoLLaMA 2: Advancing Spatial- Temporal Modeling and Audio Understanding in Video- LLMs. arXiv preprint arXiv:2406.07476, 2024. 2, 7, 14

  6. [14]

    Qwen-Audio: Advancing universal audio understanding via unified large-scale audio-language models

    Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shil- iang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou. Qwen-Audio: Advancing universal audio understanding via unified large-scale audio-language models. arXiv preprint arXiv:2311.07919, 2023. 3

  7. [15]

    TVBench: Re- designing video-language evaluation

    Daniel Cores, Michael Dorkenwald, Manuel Mucientes, Cees GM Snoek, and Yuki M Asano. TVBench: Re- designing video-language evaluation. arXiv preprint arXiv:2410.07752, 2024. 2

  8. [16]

    ScanNet: Richly-annotated 3D reconstructions of indoor scenes

    Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. ScanNet: Richly-annotated 3D reconstructions of indoor scenes. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2017. 2

  9. [17]

    Scaling egocentric vision: The EPIC-KITCHENS dataset

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The EPIC-KITCHENS dataset. In Proceedings of the European Conference on Comput...

  10. [18]

    Rescaling egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescaling egocentric vision: Collection, pipeline and challenges for EPIC-KITCHENS-100. Inter- nati...

  11. [19]

    EPIC-KITCHENS VISOR benchmark: Video segmentations and object relations

    Ahmad Darkhalil, Dandan Shan, Bin Zhu, Jian Ma, Am- lan Kar, Richard Higgins, Sanja Fidler, David Fouhey, and Dima Damen. EPIC-KITCHENS VISOR benchmark: Video segmentations and object relations. In Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Dat...

  12. [20]

    Guide to the Carnegie Mellon University Multimodal activity (CMU-MMAC) database

    Fernando De la Torre, Jessica Hodgins, Adam Bargteil, Xavier Martin, Justin Macey, Alex Collado, and Pep Bel- tran. Guide to the Carnegie Mellon University Multimodal activity (CMU-MMAC) database. 2009. 3

  13. [21]

    The Llama 3 Herd of Models

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783,

  14. [22]

    The VIA annota- tion software for images, audio and video

    Abhishek Dutta and Andrew Zisserman. The VIA annota- tion software for images, audio and video. In Proceedings of the 27th ACM International Conference on Multimedia (ACMMM), 2019. 5, 6

  15. [23]

    MMBench-Video: A long-form multi-shot benchmark for holistic video under- standing

    Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao, Yining Li, Dahua Lin, and Kai Chen. MMBench-Video: A long-form multi-shot benchmark for holistic video under- standing. In Advances in Neural Information Processing Sys- tems (NeurIPS), 2024. 2

  16. [24]

    SlowFast networks for video recognition

    Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. SlowFast networks for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019. 7, 15

  17. [25]

    Video-MME: The First-Ever Com- prehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

    Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-MME: The First-Ever Com- prehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis. In Proceedings of the IEEE/CVF Confer- ence...

  18. [26]

    Omnivore: A Sin- gle Model for Many Visual Modalities

    Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra. Omnivore: A Sin- gle Model for Many Visual Modalities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 7, 15

  19. [27]

    SSAST: Self-Supervised Audio Spectrogram Transformer

    Yuan Gong, Cheng-I Lai, Yu-An Chung, and James Glass. SSAST: Self-Supervised Audio Spectrogram Transformer. In Proceedings of the AAAI Conference on Artificial Intel- ligence, 2022. 8

  20. [28]

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jack- son Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Ku- mar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Me...

  21. [29]

    Jawahar, Richard Newcombe, Hyun Soo Park, James M

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zach Chavis, Joya Chen, Feng Cheng, Fu- Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong, Maria Es...

  22. [30]

    spaCy 2: Natural lan- guage understanding with Bloom embeddings, convolutional neural networks and incremental parsing

    Matthew Honnibal and Ines Montani. spaCy 2: Natural lan- guage understanding with Bloom embeddings, convolutional neural networks and incremental parsing. 2020. 5, 10

  23. [31]

    EPIC-SOUNDS: A Large- Scale Dataset of Actions that Sound

    Jaesung Huh, Jacob Chalk, Evangelos Kazakos, Dima Damen, and Andrew Zisserman. EPIC-SOUNDS: A Large- Scale Dataset of Actions that Sound. In IEEE Interna- tional Conference on Acoustics, Speech, & Signal Process- ing (ICASSP), 2023. 4, 5

  24. [32]

    Eye–Hand Coordination in Object Manipu- lation

    Roland Johansson, G ¨oran Westling, Anders B¨ackstr¨om, and John Flanagan. Eye–Hand Coordination in Object Manipu- lation. The Journal of neuroscience: the official journal of the Society for Neuroscience, 2001. 5

  25. [33]

    Toronto annotation suite

    Amlan Kar, Seung Wook Kim, Marko Boben, Jun Gao, Tianxing Li, Huan Ling, Zian Wang, and Sanja Fidler. Toronto annotation suite. https : / / aidemos . cs . toronto.edu/toras, 2021. 7

  26. [34]

    Slow-Fast Auditory Streams For Audio Recognition

    Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. Slow-Fast Auditory Streams For Audio Recognition. In IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP), 2021. 8

  27. [35]

    ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

    Ilker Kesen, Andrea Pedrotti, Mustafa Dogan, Michele Cafagna, Emre Can Acikgoz, Letitia Parcalabescu, Iacer Calixto, Anette Frank, Albert Gatt, Aykut Erdem, and Erkut Erdem. ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models. In Inter- ...

  28. [36]

    Segment any- thing

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision (ICCV), 2023. 5

  29. [37]

    Constituency parsing with a self-attentive encoder

    Nikita Kitaev and Dan Klein. Constituency parsing with a self-attentive encoder. In Proceedings of the Annual Meet- ing of the Association for Computational Linguistics (ACL),

  30. [38]

    Oleksii Kuchaiev, Jason Li, Huyen Nguyen, Oleksii Hrinchuk, Ryan Leary, Boris Ginsburg, Samuel Kriman, Stanislav Beliaev, Vitaly Lavrukhin, Jack Cook, Patrice Cas- tonguay, Mariya Popova, Jocelyn Huang, and Jonathan M. Cohen. NeMo: a toolkit for building AI applications using ...

  31. [39]

    LEGO: Learning EGOcen- tric Action Frame Generation via Visual Instruction Tuning

    Bolin Lai, Xiaoliang Dai, Lawrence Chen, Guan Pang, James M Rehg, and Miao Liu. LEGO: Learning EGOcen- tric Action Frame Generation via Visual Instruction Tuning. In Proceedings of the European Conference on Computer Vi- sion (ECCV), 2023. 3

  32. [40]

    The roles of vision and eye movements in the control of activities of daily living

    Michael Land, Neil Mennie, and Jennifer Rusted. The roles of vision and eye movements in the control of activities of daily living. Perception, 1999. 5

  33. [41]

    Dis- covering important people and objects for egocentric video summarization

    Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman. Dis- covering important people and objects for egocentric video summarization. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,

  34. [42]

    SEED-Bench: Benchmarking Multimodal Large Language Models

    Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. SEED-Bench: Benchmarking Multimodal Large Language Models. InPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2

  35. [43]

    MVBench: A Comprehensive Multi-modal Video Under- standing Benchmark

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. MVBench: A Comprehensive Multi-modal Video Under- standing Benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

  36. [44]

    VITATECS: A Diagnos- tic Dataset for Temporal Concept Understanding of Video- Language Models

    Shicheng Li, Lei Li, Shuhuai Ren, Yuanxin Liu, Yi Liu, Run- dong Gao, Xu Sun, and Lu Hou. VITATECS: A Diagnos- tic Dataset for Temporal Concept Understanding of Video- Language Models. In Proceedings of the European Confer- ence on Computer Vision (ECCV), 2024. 2

  37. [45]

    In the eye of the beholder: Gaze and actions in first person video

    Yin Li, Miao Liu, and James M Rehg. In the eye of the beholder: Gaze and actions in first person video. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2021. 3

  38. [46]

    HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction

    Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...

  39. [47]

    Aria Everyday Activities Dataset

    Zhaoyang Lv, Nickolas Charron, Pierre Moulon, Alexander Gamino, Cheng Peng, Chris Sweeney, Edward Miller, Huix- uan Tang, Jeff Meissner, Jing Dong, et al. Aria Everyday Activities Dataset. arXiv preprint arXiv:2402.13349, 2024. 6, 3

  40. [48]

    OpenEQA: Embodied Question Answering in the Era of Foundation Models

    Arjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta, Sriram Yenamandra, Mikael Henaff, Sneha Sil- wal, Paul Mcvay, Oleksandr Maksymets, Sergio Arnaud, Karmesh Yadav, Qiyang Li, Ben Newman, Mohit Sharma, Vincent Berges, Shiqi Zhang, Pulkit Agrawal, Yonatan Bisk, Dhruv Batr...

  41. [49]

    EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

    Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding. Advances in Neural Information Processing Systems (NeurIPS), 2023. 2

  42. [50]

    Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Percep- tion

    Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Pe- ters, Thomas Whelan, Chen Kong, Omkar Parkhi, Richard Newcombe, and Yuheng Carl Ren. Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Percep- tion. In Proceedings of the IEEE/CVF International Confer- ...

  43. [51]

    Asano, Is- han Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Jo ˜ao F

    Mandela Patrick, Dylan Campbell, Yuki M. Asano, Is- han Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Jo ˜ao F. Henriques. Keeping your eye on the ball: Trajectory attention in video transformers. In Advances in Neural Information Processing Systems (NeurIP...

  44. [52]

    CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities

    Rohith Peddi, Shivvrat Arya, Bharath Challa, Likhitha Pal- lapothula, Akshay Vyas, Bhavya Gouripeddi, Jikai Wang, Qifan Zhang, Vasundhara Komaragiri, Eric Ragan, Nicholas Ruozzi, Yu Xiang, and Vibhav Gogate. CaptainCook4D: A Dataset for Understanding Errors in Procedural Activ...

  45. [53]

    A benchmark dataset and evaluation methodology for video object segmentation

    Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (...

  46. [54]

    It’s just another day: Unique video captioning by discriminitave prompting

    Toby Perrett, Tengda Han, Dima Damen, and Andrew Zis- serman. It’s just another day: Unique video captioning by discriminitave prompting. In Proceedings of the Asian Con- ference on Computer Vision (ACCV), 2024. 2

  47. [55]

    Detecting activities of daily living in first-person camera views

    Hamed Pirsiavash and Deva Ramanan. Detecting activities of daily living in first-person camera views. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), 2012. 3

  48. [56]

    Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind

    Chiara Plizzari, Shubham Goel, Toby Perrett, Jacob Chalk, Angjoo Kanazawa, and Dima Damen. Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind. In Proceedings of the IEEE International Conference on 3D Vi- sion (3DV), 2025. 7

  49. [57]

    The 2017 davis challenge on video object segmentation

    Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Ar- bel´aez, Alexander Sorkine-Hornung, and Luc Van Gool. The 2017 davis challenge on video object segmentation. arXiv:1704.00675, 2017. 15

  50. [58]

    Perception test: A diagnostic benchmark for multimodal video models

    Viorica P ˘atr˘aucean, Lucas Smaira, Ankush Gupta, Adri`a Re- casens Continente, Larisa Markeeva, Dylan Banarse, Skanda Koppula, Joseph Heyward, Mateusz Malinowski, Yi Yang, Carl Doersch, Tatiana Matejovicova, Yury Sulsky, Antoine Miech, Alex Frechette, Hanna Klimczak, Raphael...

  51. [59]

    Learn- ing transferable visual models from natural language super- vision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International Conference on Machine Learning...

  52. [60]

    Robust speech recognition via large-scale weak supervision

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In Interna- tional Conference on Machine Learning (ICML), 2023. 3

  53. [61]

    EGO- CH: Dataset and fundamental tasks for visitors behavioral understanding using egocentric vision

    Francesco Ragusa, Antonino Furnari, Sebastiano Battiato, Giovanni Signorello, and Giovanni Maria Farinella. EGO- CH: Dataset and fundamental tasks for visitors behavioral understanding using egocentric vision. Pattern Recognition Letters, 2020. 3

  54. [62]

    The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like Domain

    Francesco Ragusa, Antonino Furnari, Salvatore Livatino, and Giovanni Maria Farinella. The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like Domain. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer ...

  55. [63]

    Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

    Santhosh K Ramakrishnan, Aaron Gokaslan, Erik Wijmans, Oleksandr Maksymets, Alex Clegg, John Turner, Eric Un- dersander, Wojciech Galuba, Andrew Westbury, Angel X Chang, et al. Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI. In Ad- vance...

  56. [64]

    SAM 2: Segment anything in images and videos

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, et al. SAM 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714, 2024. 5, 8, 7, 15

  57. [65]

    Structure-from-motion revisited

    Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 2

  58. [66]

    Pixelwise view selection for un- structured multi-view stereo

    Johannes Lutz Sch ¨onberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for un- structured multi-view stereo. InProceedings of the European Conference on Computer Vision (ECCV), 2016. 2

  59. [67]

    IndustReal: A Dataset for Procedure Step Recognition Handling Execution Errors in Egocentric Videos in an Industrial-Like Setting

    Tim J Schoonbeek, Tim Houben, Hans Onvlee, Fons van der Sommen, et al. IndustReal: A Dataset for Procedure Step Recognition Handling Execution Errors in Egocentric Videos in an Industrial-Like Setting. In Proceedings of the IEEE/CVF Winter Conference on Applications of Compute...

  60. [68]

    As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

    Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...

  61. [69]

    Actor and observer: Joint modeling of first and third-person videos

    Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. Actor and observer: Joint modeling of first and third-person videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 3

  62. [70]

    FLA V A: A foundational language and vision alignment model

    Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guil- laume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. FLA V A: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,

  63. [71]

    Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks

    Krishna Kumar Singh, Kayvon Fatahalian, and Alexei A Efros. Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks. In Pro- ceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), 2016. 3

  64. [72]

    Somasundaram, Jing Dong, Huixuan Tang, Ju- lian Straub, Mingfei Yan, Michael Goesele, Jakob J

    Kiran K. Somasundaram, Jing Dong, Huixuan Tang, Ju- lian Straub, Mingfei Yan, Michael Goesele, Jakob J. Engel, Renzo De Nardi, and Richard A. Newcombe. Project Aria: A New Tool for Egocentric Multi-Modal AI Research.arXiv preprint arXiv:2308.13561, 2023. 3, 1

  65. [73]

    Ego4D Goal-Step: Toward Hierarchical Understanding of Procedural Activi- ties

    Yale Song, Eugene Byrne, Tushar Nagarajan, Huiyu Wang, Miguel Martin, and Lorenzo Torresani. Ego4D Goal-Step: Toward Hierarchical Understanding of Procedural Activi- ties. Advances in Neural Information Processing Systems (NeurIPS), 2024. 3

  66. [74]

    EFM3D: A Bench- mark for Measuring Progress Towards 3D Egocentric Foun- dation Models

    Julian Straub, Daniel DeTone, Tianwei Shen, Nan Yang, Chris Sweeney, and Richard Newcombe. EFM3D: A Bench- mark for Measuring Progress Towards 3D Egocentric Foun- dation Models. arXiv preprint arXiv:2406.10224, 2024. 6, 3

  67. [75]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text. arXiv preprint arXiv:2403.05530, 2024. 7, 14

  68. [76]

    VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

    Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training. InAdvances in Neural Information Processing Systems (NeurIPS), 2022. 7, 15

  69. [77]

    EPIC Fields: Marrying 3D Geometry and Video Understanding

    Vadim Tschernezki, Ahmad Darkhalil, Zhifan Zhu, David Fouhey, Iro Larina, Diane Larlus, Dima Damen, and Andrea Vedaldi. EPIC Fields: Marrying 3D Geometry and Video Understanding. In Advances in Neural Information Process- ing Systems (NeurIPS), 2023. 2

  70. [78]

    SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge

    Andong Wang, Bo Wu, Sunli Chen, Zhenfang Chen, Hao- tian Guan, Wei-Ning Lee, Li Erran Li, and Chuang Gan. SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...

  71. [79]

    InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions

    Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiaowei Hu, Tong Lu, Lewei Lu, Hongsheng Li, et al. InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions. In Proceedings of the IEEE/CVF Conference on Computer Vision a...

  72. [80]

    HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World

    Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bu- gra Tekin, Felipe Vieira Frujeri, Neel Joshi, and Marc Polle- feys. HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World...

  73. [81]

    Florence-2: Advancing a unified representation for a variety of vision tasks

    Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan. Florence-2: Advancing a unified representation for a variety of vision tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2

  74. [82]

    Can I trust your answer? Visually grounded video question answering

    Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua. Can I trust your answer? Visually grounded video question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 7

  75. [83]

    YouTube-VOS: A large-scale video object segmentation benchmark

    Ning Xu, Linjie Yang, Yuchen Fan, Dingcheng Yue, Yuchen Liang, Jianchao Yang, and Thomas Huang. YouTube-VOS: A large-scale video object segmentation benchmark. arXiv preprint arXiv:1809.03327, 2018. 15

  76. [84]

    The 2nd large-scale video object segmentation challenge - video object segmen- tation track, 2019

    Linjie Yang, Yuchen Fan, and Ning Xu. The 2nd large-scale video object segmentation challenge - video object segmen- tation track, 2019. 15

  77. [85]

    Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

    Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 5

  78. [86]

    MM-Ego: Towards Building Egocentric Multimodal LLMs

    Hanrong Ye, Haotian Zhang, Erik Daxberger, Lin Chen, Zongyu Lin, Yanghao Li, Bowen Zhang, Haoxuan You, Dan Xu, Zhe Gan, et al. MM-Ego: Towards Building Egocentric Multimodal LLMs. In International Conference on Learning Representations (ICLR), 2025. 2

  79. [87]

    Es- timating body and hand motion in an ego-sensed world

    Brent Yi, Vickie Ye, Maya Zheng, Lea M ¨uller, Georgios Pavlakos, Yi Ma, Jitendra Malik, and Angjoo Kanazawa. Es- timating body and hand motion in an ego-sensed world. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2025. 5

  80. [88]

    Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

    Hang Zhang, Xin Li, and Lidong Bing. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2023. 2

  81. [89]

    Long context transfer from language to vision

    Peiyuan Zhang, Kaichen Zhang, Bo Li, Guangtao Zeng, Jingkang Yang, Yuanhan Zhang, Ziyue Wang, Haoran Tan, Chunyuan Li, and Ziwei Liu. Long context transfer from language to vision. arXiv preprint arXiv:2406.16852, 2024. 7, 14

  82. [90]

    Video instruction tuning with synthetic data

    Yuanhan Zhang, Jinming Wu, Wei Li, Bo Li, Zejun Ma, Zi- wei Liu, and Chunyuan Li. Video instruction tuning with synthetic data. arXiv preprint arXiv:2410.02713, 2024. 7, 14

  83. [91]

    Instance tracking in 3D scenes from egocentric videos

    Yunhan Zhao, Haoyu Ma, Shu Kong, and Charless Fowlkes. Instance tracking in 3D scenes from egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), 2024. 2, 3

  84. [92]

    Needle In A Video Haystack: A Scalable Synthetic Framework for Benchmarking Video MLLMs

    Zijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du, Tongtian Yue, Longteng Guo, Bingning Wang, Weipeng Chen, and Jing Liu. Needle In A Video Haystack: A Scalable Synthetic Framework for Benchmarking Video MLLMs. In Inter- national Conference on Learning Representations (ICLR) ,

  85. [93]

    use something

    Luowei Zhou, Chenliang Xu, and Jason Corso. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence, 2018. 2 HD-EPIC: A Highly-Detailed Egocentric Video Dataset Supplementary Material Contents A ....

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.