Pith. sign in

REVIEW 4 major objections 4 minor 4 cited by

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Vinci claims to be the first always-on wearable assistant built on an egocentric vision-language model, answering spoken questions about both the current scene and past events in real time.

desk verdict A well-scoped system-demo paper with a real code release; the integration is new, but "real-time" is asserted, not measured, and the memory module's exact timestamps need explaining. read the letter →

arxiv 2412.21080 v1 pith:RIT5NFRV submitted 2024-12-30 cs.CV

classification cs.CV
keywords egocentricvisionvision-languagemodelreal-timeassistantstreamingvideounderstandingtemporalgroundinggenerationmemorymodulewearableAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Vinci is a proposed system that tries to establish that an egocentric vision-language model can run as an always-on assistant on portable devices, watching a continuous video stream and answering spoken questions about both the present scene and past events. The authors argue that, by coupling a video encoder with a large language model and adding a memory module that periodically writes timestamped text descriptions, the system can perform temporal grounding, video summarization, and future planning from first-person video. It also combines two forms of visual guidance: a generation module that produces short synthesized how-to clips and a retrieval module that pulls relevant third-person instructional videos. The paper reports qualitative demonstrations of each capability, positioning Vinci as a first step toward practical real-time egocentric AI assistants.

What carries the argument

The load-bearing component is EgoVideo-VL, an egocentric vision-language model formed by attaching the EgoVideo encoder to the InternLM-7B large language model and instruction-tuning the pair on egocentric video-text data. The memory module is the second central mechanism: it continuously writes short timestamped text descriptions of observed actions, so that queries about the past can be answered without storing the full video stream. The generation module, based on SEINE, turns the current frame plus a user prompt into a two-second synthesized clip, and the retrieval module matches query text against cached features from HowTo100M to return third-person demonstrations. These modules run concurrently through a backend hub that connects a camera, a web frontend, and text-to-speech audio output.

What would settle it

Measure end-to-end latency on the deployed smartphone setup from wake-word detection to the start of the spoken answer, and measure memory-module processing rate against the incoming frame rate; if the latency exceeds a few seconds or the memory backlog grows without bound over an hour-long stream, the real-time claim is refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that a single system, called Vinci, can deliver real-time embodied assistance by fine-tuning an egocentric video foundation model into a vision-language model, EgoVideo-VL, and then integrating it with a memory module, a video-generation module, and a retrieval module. EgoVideo-VL connects the EgoVideo encoder to a fixed InternLM-7B language model and is instruction-tuned on a curated dataset built from Ego4D, EgoExoLearn, and Ego4D-Goalstep, giving it free-form conversational ability in egocentric settings. The memory module periodically captures video, writes detailed textual descriptions with timestamps, and feeds relevant history into the model on query, enabling answers about past actions. The generation module, a fine-tuned SEINE model, outputs two-second egocentric video demonstrations of requested actions, while the retrieval module searches HowTo100M for third-person how-to videos. The paper shows qualitative examples of current scene understanding, temporal grounding, video summarization, future planning, action prediction, and cross-view retrieval.

Load-bearing premise

The central real-time claim depends on an unmeasured latency and resource budget: if EgoVideo-VL's per-query response time is slower than a natural conversational exchange, or if the memory module cannot keep up with the always-on video stream, the system would not actually be a real-time assistant.

Editorial extensions

If this is right

  • A user wearing a camera can ask "When did I add sugar?" and receive a timestamped answer grounded in the memory log.
  • The system can summarize long multi-step activities from first-person video and propose next steps based on current state.
  • For tasks requiring manipulation, Vinci can generate a short synthesized clip of the next action or retrieve an existing how-to video.
  • The open-sourced implementation (model weights plus frontend and backend code) provides a complete blueprint for other researchers to deploy similar wearable assistants.
  • The instruction-tuning dataset constructed from Ego4D, EgoExoLearn, and Ego4D-Goalstep is claimed to equip the model with egocentric conversational ability and procedural reasoning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the real-time claim is taken literally, the system's usefulness depends on end-to-end latency; the paper does not report numbers, so a natural next step is measuring wake-word-to-answer delay on a deployed device.
  • The memory module's periodic snapshots could be extended to support long-horizon episodic queries over days, but would require a compression or summarization strategy to avoid unbounded storage.
  • The video generation module is egocentric-specific and likely fails on third-person views; a testable extension is measuring generation quality as a function of camera perspective.
  • Beyond assistance, the same architecture could serve as a data-collection device for egocentric activity understanding, automatically generating timestamped narrations that could train future models.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces Vinci, an embodied smart assistant built on an egocentric vision-language model called EgoVideo-VL, together with an input processing module, a periodic memory module, a video generation module, and a retrieval module. The system is claimed to run in real time on portable devices, support always-on observation, answer questions about current and past events, provide future planning, and generate or retrieve short instructional video clips. The authors release the implementation and a demo web platform. The evaluation section consists entirely of qualitative examples from uploaded videos and a deployed demo; no quantitative metrics, latency measurements, ablations, or baseline comparisons are reported.

Significance. If the performance claims were substantiated, Vinci would represent a useful integration of an egocentric vision-language model with streaming memory, retrieval, and video generation, and the open release of the full system would be a practical contribution to the community. The paper also demonstrates a plausible, unified architecture for wearable embodied assistants. However, the current evidence is insufficient to validate the central claims of real-time operation and accurate temporal grounding, because these rest on unquantified system behavior and hand-picked qualitative demonstrations rather than measurable evaluation.

major comments (4)
  1. [Sections 3.6 and 4] The central claim that Vinci operates in real time is not supported by any measurement. The paper never reports end-to-end latency from user speech to audio response, video frame processing rate, GPU or CPU utilization, memory footprint, or network delay. Figure 3 shows a backend-server architecture, so the real-time behavior depends on server scheduling and network conditions that are not quantified. Without these numbers, the repeated use of “real-time” in the abstract, introduction, and Section 3.6 is an unsupported assertion rather than an established property of the system.
  2. [Section 3.3] The memory module is described as “periodically capturing short video snapshots” and storing descriptions with timestamps, but the period is never stated. Figures 9 and 10 claim exact second-level temporal grounding (e.g., “the sugar was added at 58 seconds” and “you washed the bell pepper at 136 seconds”). If the snapshot period is coarse, such precision cannot be derived from stored memory entries and would have to be hallucinated or interpolated, which is not explained. If the period is fine enough to support such precision, the always-on computational cost of running EgoVideo-VL on every snapshot must be quantified, and it is not.
  3. [Section 4] The experimental evaluation is exclusively qualitative, consisting of a small set of hand-picked examples from the Gradio demo. There are no quantitative results on standard egocentric benchmarks, no ablations of the memory period, the two-stage fine-tuning, or the generation module, and no comparisons with the streaming-video systems discussed in Section 2.3 (e.g., VideoLLM-Online, Flash-VStream, MMDuet). Moreover, the examples appear to come from the same types of egocentric datasets used to train the model, so the demonstrations do not establish generalization, reliability, or the marginal contribution of any individual module. This evidential basis is too thin to support the paper’s capability claims.
  4. [Section 3.4] The generation module is a fine-tuned SEINE diffusion model, and Section 4.7 reports that it outputs 2-second video clips. The paper does not report the inference latency of this module, which is relevant because video diffusion models are typically computationally expensive and often far from real time on portable hardware. Since this module is invoked during a user conversation and the user is waiting for the visual demonstration, the paper must state at least the generation time and the hardware on which it runs; otherwise the “real-time assistant” framing is incomplete for this interaction path.
minor comments (4)
  1. [Section 4.2] The heading “Current scene undersanding” contains a typo and should read “Current scene understanding.”
  2. [Figure 1] The figure contains typographical errors in the displayed text, including “you have doned the following” and “peper”; these should be corrected to “done” and “pepper.”
  3. [Section 3.4] The optical-flow threshold and the verb-frequency criterion for filtering the generation training dataset are described only as “a defined threshold” and “reasonable frequency”; these values should be stated to make the training procedure reproducible.
  4. [Section 3.1] The external services (Baidu ASR and the wake-up keyword API) are mentioned by name without references or version information; a citation or URL would help readers reproduce the deployed system.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: Vinci is a system-integration report with qualitative demonstrations, and the self-citations are component citations rather than load-bearing derivations.

full rationale

The paper does not derive a quantitative prediction from a fitted parameter, and no equation or definitional equivalence reduces a reported output to the paper's own inputs. The self-citations to EgoVideo [65], EgoInstructor [84], and SEINE [15] are used to identify components that Vinci builds upon or integrates; none of these citations is invoked as an external theorem that forces the paper's conclusions. The central 'real-time' and temporal-grounding claims are indeed unquantified: no end-to-end latency, memory sampling period, or device load is reported, and the evaluation is entirely qualitative. These are validity and reproducibility weaknesses, not circularity. The instruction-tuning dataset is assembled from public ego datasets, and the model is then shown on demo examples; this is standard fine-tune-and-demonstrate practice, and concerns about cherry-picking or overfitting belong to correctness risk, not circular derivation. Accordingly, the circularity score is 0.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central system claim depends on unverified domain assumptions about the underlying egocentric foundation model, the quality of generated memory descriptions, and the sufficiency of the curated instruction data, plus several hand-chosen hyperparameters. No new scientific entities are introduced.

free parameters (4)
  • memory snapshot period = unspecified
    The memory module periodically captures video snapshots, but the exact interval is not given; this affects temporal grounding recall and real-time performance.
  • optical flow threshold = a defined threshold
    Used to filter training videos for the generation module to select smooth global motion; the threshold value is not specified.
  • verb frequency threshold = reasonable frequency
    Used to filter training data for video generation to avoid rare verbs; the threshold is described qualitatively.
  • retrieval top-K = K=3
    Number of instructional videos retrieved by the retrieval module; chosen by hand.
assumptions (5)
  • domain assumption EgoVideo provides suitable visual features for instruction tuning and live streaming understanding.
    EgoVideo-VL is built directly on EgoVideo without comparing to alternative backbones; Section 3.2.
  • domain assumption Narrations from Ego4D and EgoExoLearn, with LLM-refined prompt templates, are sufficient instruction data for egocentric alignment.
    No scaling or data quality analysis is given; Section 3.2.
  • domain assumption The memory module's periodic descriptions are accurate enough for temporal grounding and summarization.
    Memory stores model-generated descriptions without verification; Section 3.3.
  • domain assumption Fine-tuned SEINE can produce actionable two-second egocentric demonstrations from a single frame and a user prompt.
    The paper states generation works for egocentric scenarios but is less effective for third-person video, and there is no user study; Sections 3.4 and 4.7.
  • domain assumption The selected qualitative examples are representative of overall system performance.
    All conclusions are drawn from hand-picked figures; Section 4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model." pith.science (2026). https://pith.science/paper/RIT5NFRV

@misc{pith2026241221080,
  author       = {Pith},
  title        = {Pith review of: Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RIT5NFRV}},
  note         = {Machine review of arXiv:2412.21080}
}
read the original abstract

We introduce Vinci, a real-time embodied smart assistant built upon an egocentric vision-language model. Designed for deployment on portable devices such as smartphones and wearable cameras, Vinci operates in an "always on" mode, continuously observing the environment to deliver seamless interaction and assistance. Users can wake up the system and engage in natural conversations to ask questions or seek assistance, with responses delivered through audio for hands-free convenience. With its ability to process long video streams in real-time, Vinci can answer user queries about current observations and historical context while also providing task planning based on past interactions. To further enhance usability, Vinci integrates a video generation module that creates step-by-step visual demonstrations for tasks that require detailed guidance. We hope that Vinci can establish a robust framework for portable, real-time egocentric AI systems, empowering users with contextual and actionable insights. We release the complete implementation for the development of the device in conjunction with a demo web platform to test uploaded videos at https://github.com/OpenGVLab/vinci.

Figures

Figures reproduced from arXiv: 2412.21080 by the authors.

Figure 1
Figure 1. Overview of Vinci’s capabilities demonstrated through a streaming video timeline. At different timestamps, Vinci showcases [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The overall structure of the EgoVideo-VL model. The visual encoder leverages the egocentric video foundation model, EgoVideo, [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Overview of the Vinci system. The system integrates four components: the camera, frontend, backend, and models. The frontend [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Real-world deployment of Vinci. (a) The deployed system on a OnePlus smartphone mounted on the user’s head. (b) The [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Example of Vinci’s ability to analyze the current video state and accurately respond to user queries. In this scenario, at 35.2 [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Example of Vinci’s ability to generalize to embodied navigation scenarios without additional tuning. In this case, Vinci accurately [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 7
Figure 7. Figure 7: Example of Vinci’s general video scene understanding capabilities. In this scenario, Vinci successfully identifies the filming [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Example of Vinci’s strong OCR capability. In this scenario, Vinci correctly recognizes the Chinese characters on the box when [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]
Figure 9
Figure 9. Figure 9: Vinci can perform temporal grounding, helping the user to locate the timestamp of specific previous actions. [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Vinci can correctly locate previous actions even when they are queried out of chronological order or with a significant time gap [PITH_FULL_IMAGE:figures/full_fig_p009_10.png]
Figure 11
Figure 11. Figure 11: Vinci can make the stepwise summarization of long-horizon videos. [PITH_FULL_IMAGE:figures/full_fig_p009_11.png]
Figure 12
Figure 12. Figure 12: Vinci can provide future planning based on the current video state. In the bottom example, Vinci demonstrates its ability to create [PITH_FULL_IMAGE:figures/full_fig_p010_12.png]
Figure 13
Figure 13. Figure 13: Examples of retrieval results from Vinci’s retrieval module. The system identifies relevant how-to demonstration videos from a [PITH_FULL_IMAGE:figures/full_fig_p010_13.png]
Figure 14
Figure 14. Figure 14: Examples of visual demonstrations generated by Vinci’s generation module. Each 2-second video clip is synthesized based on the [PITH_FULL_IMAGE:figures/full_fig_p011_14.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

    cs.CV 2026-07 conditional novelty 6.5 of 10

    EgoMemo uses multi-scale temporal summaries, a knowledge graph, and visual archives to decide whether and when to intervene proactively on continuous egocentric video, setting baselines on the new EgoServe benchmark o...

  2. EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?

    cs.CV 2025-06 conditional novelty 6.0 of 10

    EOC-Bench evaluates MLLMs on egocentric object cognition across past, present, and future temporal dimensions, finding large gaps versus humans, especially in absolute time perception.

  3. Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EAIL localizes a person in a 3D point cloud from head-mounted IMU signals by aligning short action segments with scene locations using vision-language training guidance.

  4. Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision

    cs.CV 2025-06 accept novelty 3.0 of 10

    A comprehensive review of cross-view video understanding that uses both first-person and third-person cameras, organized into a three-direction taxonomy with a dataset catalog and future research gaps.

Reference graph

Works this paper leans on

97 extracted references · 58 canonical work pages · cited by 4 Pith papers

  1. [1]

    Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P

    Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, An- drew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michae...

  2. [2]

    Swim- master: a wearable assistant for swimmer

    Marc B¨achlin, Kilian F ¨orster, and Gerhard Tr ¨oster. Swim- master: a wearable assistant for swimmer. In Proceedings of the 11th international conference on Ubiquitous computing, pages 215–224, 2009. 1

  3. [3]

    Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023. 1, 3

  4. [4]

    Analysis of the hands in egocentric vision: A survey

    Andrea Bandini and Jos ´e Zariffa. Analysis of the hands in egocentric vision: A survey. IEEE transactions on pattern analysis and machine intelligence, 45(6):6846–6866, 2020. 1

  5. [5]

    Siddhant Bansal, Chetan Arora, and C.V . Jawahar. My view is the best view: Procedure learning from egocentric videos. In Proceedings of the European Conference on Computer Vision (ECCV), 2022. 2

  6. [6]

    Internlm2 technical report, 2024

    Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiaoyi Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang, Peng- long Jiao, Zhenjiang Jin, Zhikai Lei, Jiaxing Li, Jingwen Li, Linyang Li,...

  7. [7]

    Internvideo-ego4d: A pack of champion solu- tions to ego4d challenges

    Guo Chen, Sen Xing, Zhe Chen, Yi Wang, Kunchang Li, Yizhuo Li, Yi Liu, Jiahao Wang, Yin-Dong Zheng, Bingkun Huang, et al. Internvideo-ego4d: A pack of champion solu- tions to ego4d challenges. arXiv preprint arXiv:2211.09529,

  8. [8]

    Dcan: improving temporal action detection via dual context aggregation

    Guo Chen, Yin-Dong Zheng, Limin Wang, and Tong Lu. Dcan: improving temporal action detection via dual context aggregation. In Proceedings of the AAAI conference on artifi- cial intelligence, pages 248–257, 2022. 3

Show all 97 references
  1. [9]

    Elan: Enhancing temporal action detection with location awareness

    Guo Chen, Yin-Dong Zheng, Zhe Chen, Jiahao Wang, and Tong Lu. Elan: Enhancing temporal action detection with location awareness. In 2023 IEEE International Conference on Multimedia and Expo (ICME), pages 1020–1025. IEEE, 2023

  2. [10]

    Videollm: Modeling video sequence with large language models

    Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, et al. Videollm: Modeling video sequence with large language models. arXiv preprint arXiv:2305.13292, 2023

  3. [11]

    Video mamba suite: State space model as a ver- satile alternative for video understanding

    Guo Chen, Yifei Huang, Jilan Xu, Baoqi Pei, Zhe Chen, Zhiqi Li, Jiahao Wang, Kunchang Li, Tong Lu, and Limin Wang. Video mamba suite: State space model as a ver- satile alternative for video understanding. arXiv preprint arXiv:2403.09626, 2024

  4. [12]

    Cg-bench: Clue-grounded question answering benchmark for long video understanding

    Guo Chen, Yicheng Liu, Yifei Huang, Yuping He, Baoqi Pei, Jilan Xu, Yali Wang, Tong Lu, and Limin Wang. Cg-bench: Clue-grounded question answering benchmark for long video understanding. arXiv preprint arXiv:2412.12075, 2024. 3

  5. [13]

    Gatehub: Gated history unit with background suppression for online action detection

    Junwen Chen, Gaurav Mittal, Ye Yu, Yu Kong, and Mei Chen. Gatehub: Gated history unit with background suppression for online action detection. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 19893–19902, 2022. 3

  6. [14]

    Videollm-online: Online video large language model for streaming video

    Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, and Mike Zheng Shou. Videollm-online: Online video large language model for streaming video. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  7. [15]

    Seine: Short-to-long video diffusion model for generative transition and prediction

    Xinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang, Xin Ma, Jiashuo Yu, Yali Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Seine: Short-to-long video diffusion model for generative transition and prediction. In ICLR, 2023. 5

  8. [16]

    gSDF: Geometry-Driven signed distance functions for 3D hand-object reconstruction

    Zerui Chen, Shizhe Chen, Cordelia Schmid, and Ivan Laptev. gSDF: Geometry-Driven signed distance functions for 3D hand-object reconstruction. In CVPR, 2023. 2

  9. [17]

    Expanding performance boundaries of open-source multimodal models with model, data, and test- time scaling

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhang- wei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test- time scaling. arXiv preprint arXiv:2412.05271 , 2024. 1, 3

  10. [18]

    Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

    Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In 2024 IEE...

  11. [19]

    Scaling egocentric vision: The epic-kitchens dataset

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In Pro- ceedings of the European Conference on Comp...

  12. [20]

    Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, An- tonino Furnari, Jian Ma, Evangelos Kazakos, Davide Molti- santi, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. Interna...

  13. [21]

    Wearable reasoner: towards enhanced human rational- ity through a wearable device with an explainable ai assistant

    Valdemar Danry, Pat Pataranutaporn, Yaoli Mao, and Pattie Maes. Wearable reasoner: towards enhanced human rational- ity through a wearable device with an explainable ai assistant. In Proceedings of the Augmented Humans International Con- ference, pages 1–12, 2020. 1

  14. [22]

    Summarization of egocentric videos: A com- prehensive survey

    Ana Garcia Del Molino, Cheston Tan, Joo-Hwee Lim, and Ah-Hwee Tan. Summarization of egocentric videos: A com- prehensive survey. IEEE Transactions on Human-Machine Systems, 47(1):65–76, 2016. 2

  15. [23]

    Un- locking exocentric video-language data for egocentric video representation learning

    Zi-Yi Dou, Xitong Yang, Tushar Nagarajan, Huiyu Wang, Jing Huang, Nanyun Peng, Kris Kitani, and Fu-Jen Chu. Un- locking exocentric video-language data for egocentric video representation learning. arXiv preprint arXiv:2408.03567,

  16. [24]

    Learning to recognize objects in egocentric activities

    Alireza Fathi, Xiaofeng Ren, and James M Rehg. Learning to recognize objects in egocentric activities. In CVPR, pages 3281–3288, 2011. 1

  17. [25]

    What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention

    Antonino Furnari and Giovanni Farinella. What would you expect? anticipating egocentric actions with rolling-unrolling lstms and modality attention. In 2019 IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), pages 6251– 6260, 2019. 3

  18. [26]

    What would you expect? anticipating egocentric actions with rolling- unrolling lstms and modality attention

    Antonino Furnari and Giovanni Maria Farinella. What would you expect? anticipating egocentric actions with rolling- unrolling lstms and modality attention. In Proceedings of the International Conference on Computer Vision (ICCV) ,

  19. [27]

    Unsupervised video summarization via relation- aware assignment learning

    Junyu Gao, Xiaoshan Yang, Yingying Zhang, and Chang- sheng Xu. Unsupervised video summarization via relation- aware assignment learning. IEEE Transactions on Multime- dia, 23:3203–3214, 2020. 3

  20. [28]

    Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos

    Reza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, and Be- hzad Dariush. Weakly-supervised action segmentation and unseen error detection in anomalous instructional videos. 2023 IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 10094–10104, 2023. 3

  21. [29]

    Ego4d: Around the World in 3,000 Hours of Egocentric Video

    Kristen Grauman, Andrew Westbury, and Eugene Byrne et al. Ego4d: Around the World in 3,000 Hours of Egocentric Video. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 3, 5, 6

  22. [30]

    Ego-exo4d: Un- derstanding skilled human activity from first-and third-person perspectives

    Kristen Grauman, Andrew Westbury, et al. Ego-exo4d: Un- derstanding skilled human activity from first-and third-person perspectives. arXiv preprint arXiv:2311.18259, 2023. 1

  23. [31]

    Weiming Hu, Qiang Wang, Li Zhang, Luca Bertinetto, and Philip H.S. Torr. Siammask: A framework for fast online object tracking and segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(3):3072–3089,

  24. [32]

    Pre- dicting gaze in egocentric video by learning task-dependent attention transition

    Yifei Huang, Minjie Cai, Zhenqiang Li, and Yoichi Sato. Pre- dicting gaze in egocentric video by learning task-dependent attention transition. In Proceedings of the European Confer- ence on Computer Vision (ECCV), 2018. 1

  25. [33]

    Mutual context network for jointly estimating egocentric gaze and action

    Yifei Huang, Minjie Cai, Zhenqiang Li, Feng Lu, and Yoichi Sato. Mutual context network for jointly estimating egocentric gaze and action. IEEE Transactions on Image Processing, 29: 7795–7806, 2020. 2

  26. [34]

    Improving action segmentation via graph-based temporal reasoning

    Yifei Huang, Yusuke Sugano, and Yoichi Sato. Improving action segmentation via graph-based temporal reasoning. In Proceedings of the Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 14024–14034, 2020. 1

  27. [35]

    Compound proto- type matching for few-shot action recognition

    Yifei Huang, Lijin Yang, and Yoichi Sato. Compound proto- type matching for few-shot action recognition. InProceedings of the European Conference on Computer Vision (ECCV) ,

  28. [36]

    Weakly supervised temporal sentence grounding with uncertainty-guided self- training

    Yifei Huang, Lijin Yang, and Yoichi Sato. Weakly supervised temporal sentence grounding with uncertainty-guided self- training. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3

  29. [37]

    Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world

    Yifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang, Lijin Yang, Baoqi Pei, Hongjie Zhang, Lu Dong, Yali Wang, Limin Wang, et al. Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world. In Proceedings of the Conference on...

  30. [38]

    VBench: Com- prehensive benchmark suite for video generative models

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. VBench: Com- prehensive benchmark suite for video generative models....

  31. [39]

    Vbench++: Comprehensive and versatile benchmark suite for video generative models

    Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying-Cong Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench++: Comprehensive and versatile benchmark suite for v...

  32. [40]

    Towards intelligent wearable assistants

    Nuwan Janaka. Towards intelligent wearable assistants. In Companion of the 2024 on ACM International Joint Confer- ence on Pervasive and Ubiquitous Computing, pages 618–621,

  33. [41]

    Demonstrating tom: A de- velopment platform for wearable intelligent assistants

    Nuwan Janaka, Shengdong Zhao, David Hsu, Sherisse Tan Jing Wen, and Chun Keat Koh. Demonstrating tom: A de- velopment platform for wearable intelligent assistants. In Companion of the 2024 on ACM International Joint Confer- ence on Pervasive and Ubiquitous Computing, pages 214–219,

  34. [42]

    Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities

    Baoxiong Jia, Yixin Chen, Siyuan Huang, Yixin Zhu, and Song-Chun Zhu. Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2

  35. [43]

    Epic-fusion: Audio-visual temporal binding for egocentric action recognition

    Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. Epic-fusion: Audio-visual temporal binding for egocentric action recognition. In Proceedings of the In- ternational Conference on Computer Vision (ICCV) , 2019. 1

  36. [44]

    Time- conditioned action anticipation in one shot

    Qiuhong Ke, Mario Fritz, and Bernt Schiele. Time- conditioned action anticipation in one shot. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9917–9926, 2019. 3

  37. [45]

    Learning to discriminate information for online action detection: Anal- ysis and application

    Sumin Lee, Hyunjun Eun, Jinyoung Moon, Seokeon Choi, Yoonhyung Kim, Chanho Jung, and Changick Kim. Learning to discriminate information for online action detection: Anal- ysis and application. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(5):5918–5934, 2023. 3

  38. [46]

    Ego-body pose es- timation via ego-head pose estimation

    Jiaman Li, Karen Liu, and Jiajun Wu. Ego-body pose es- timation via ego-head pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17142–17151, 2023. 2

  39. [47]

    Videochat: Chat-centric video understanding

    KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023. 1

  40. [48]

    Jointly localizing and describing events for dense video captioning

    Yehao Li, Ting Yao, Yingwei Pan, Hongyang Chao, and Tao Mei. Jointly localizing and describing events for dense video captioning. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), pages 7492–7500,

  41. [49]

    Egocentric video-language pretraining

    Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z XU, Difei Gao, Rong-Cheng Tu, Wen- zhe Zhao, Weijie Kong, et al. Egocentric video-language pretraining. In Proceedings of the Advances in Neural Infor- mation Processing Systems (NeurIPS), 2022. 3

  42. [50]

    Univtg: Towards unified video-language temporal grounding

    Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick, Difei Gao, Alex Jinpeng Wang, Rui Yan, and Mike Zheng Shou. Univtg: Towards unified video-language temporal grounding. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 2794–2804,

  43. [51]

    Visual instruction tuning, 2023

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning, 2023. 3

  44. [52]

    Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

    Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reason- ing, ocr, and world knowledge, 2024. 3

  45. [53]

    Video summariza- tion through reinforcement learning with a 3d spatio-temporal u-net

    Tianrui Liu, Qingjie Meng, Jun-Jie Huang, Athanasios Vlont- zos, Daniel Rueckert, and Bernhard Kainz. Video summariza- tion through reinforcement learning with a 3d spatio-temporal u-net. IEEE Transactions on Image Processing , 31:1573– 1586, 2022. 3

  46. [54]

    Latte: Latent diffusion transformer for video generation

    Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048, 2024. 5

  47. [55]

    Online episodic memory visual query localization with egocentric streaming object memory

    Zaira Manigrasso, Matteo Dunnhofer, Antonino Furnari, Moritz Nottebaum, Antonio Finocchiaro, Davide Marana, Giovanni Maria Farinella, and Christian Micheloni. Online episodic memory visual query localization with egocentric streaming object memory. arXiv preprint arXiv:2411.16934,

  48. [56]

    HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

    Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. In Proceedings of the International Conference on Computer Vision (ICCV), 2019. 6, 7

  49. [57]

    Integrating human gaze into attention for egocentric activity recognition

    Kyle Min and Jason J Corso. Integrating human gaze into attention for egocentric activity recognition. In Proceedings of the Winter Conference on Applications of Computer Vision (WACV), pages 1069–1078, 2021. 2

  50. [58]

    Learning affordance landscapes for interaction exploration in 3d environments

    Tushar Nagarajan and Kristen Grauman. Learning affordance landscapes for interaction exploration in 3d environments. In NeurIPS, 2020. 2

  51. [59]

    Egoenv: Human- centric environment representations from egocentric video

    Tushar Nagarajan, Santhosh Kumar Ramakrishnan, Ruta De- sai, James Hillis, and Kristen Grauman. Egoenv: Human- centric environment representations from egocentric video. In NeurIPS, 2023. 2

  52. [60]

    Gpt-4v(ision) system card

    OpenAI. Gpt-4v(ision) system card. 2023. 1

  53. [61]

    Hello gpt-4o

    OpenAI. Hello gpt-4o. 2024. 1

  54. [62]

    Actionvos: Actions as prompts for video object segmentation

    Liangyang Ouyang, Ruicong Liu, Yifei Huang, Ryosuke Fu- ruta, and Yoichi Sato. Actionvos: Actions as prompts for video object segmentation. In European Conference on Com- puter Vision, pages 216–235, 2024. 2

  55. [63]

    Wear- able augmented reality system using gaze interaction

    Hyung Min Park, Seok Han Lee, and Jong Soo Choi. Wear- able augmented reality system using gaze interaction. In Proceedings of the IEEE/ACM International Symposium on Mixed and Augmented Reality, 2008. 1

  56. [64]

    Deep learning-based smart task assistance in wearable augmented reality

    Kyeong-Beom Park, Minseok Kim, Sung Ho Choi, and Jae Yeol Lee. Deep learning-based smart task assistance in wearable augmented reality. Robotics and Computer- Integrated Manufacturing, 63:101887, 2020. 1

  57. [65]

    Egovideo: Exploring egocentric foun- dation model and downstream adaptation

    Baoqi Pei, Guo Chen, Jilan Xu, Yuping He, Yicheng Liu, Kanghua Pan, Yifei Huang, Yali Wang, Tong Lu, Limin Wang, and Yu Qiao. Egovideo: Exploring egocentric foun- dation model and downstream adaptation. arXiv preprint arXiv:2406.18070, 2024. 1, 4

  58. [66]

    An outlook into the future of egocentric vision

    Chiara Plizzari, Gabriele Goletto, Antonino Furnari, Siddhant Bansal, Francesco Ragusa, Giovanni Maria Farinella, Dima Damen, and Tatiana Tommasi. An outlook into the future of egocentric vision. arXiv preprint arXiv:2308.07123, 2023. 1

  59. [67]

    Spatial cognition from egocentric video: Out of sight, not out of mind.arXiv preprint arXiv:2404.05072, 2024

    Chiara Plizzari, Shubham Goel, Toby Perrett, Jacob Chalk, Angjoo Kanazawa, and Dima Damen. Spatial cognition from egocentric video: Out of sight, not out of mind.arXiv preprint arXiv:2404.05072, 2024. 2

  60. [68]

    Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

    Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang. Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. In Proceedings of the International Conference on Computer Vision (ICCV) ,

  61. [69]

    Streaming long video understanding with large language models, 2024

    Rui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Shuangrui Ding, Dahua Lin, and Jiaqi Wang. Streaming long video understanding with large language models, 2024. 3

  62. [70]

    Sener, D

    F. Sener, D. Chatterjee, D. Shelepov, K. He, D. Singhania, R. Wang, and A. Yao. Assembly101: A large-scale multi- view video dataset for understanding procedural activities. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2

  63. [71]

    Understanding human hands in contact at internet scale

    Dandan Shan, Jiaqi Geng, Michelle Shu, and David Fouhey. Understanding human hands in contact at internet scale. In Proceedings of the Conference on Computer Vision and Pat- tern Recognition (CVPR), 2020. 2

  64. [72]

    Charades-ego: A large-scale dataset of paired third and first person videos

    Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. Charades-ego: A large-scale dataset of paired third and first person videos. arXiv preprint arXiv:1804.09626, 2018. 2

  65. [73]

    Towards diverse paragraph captioning for untrimmed videos

    Yuqing Song, Shizhe Chen, and Qin Jin. Towards diverse paragraph captioning for untrimmed videos. InProceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), pages 11245–11254, 2021. 3

  66. [74]

    Ego4d goal-step: Toward hierarchical understanding of procedural activities

    Yale Song, Eugene Byrne, Tushar Nagarajan, Huiyu Wang, Miguel Martin, and Lorenzo Torresani. Ego4d goal-step: Toward hierarchical understanding of procedural activities. Proceedings of the Advances in Neural Information Process- ing Systems (NeurIPS), 2024. 5

  67. [75]

    H+ o: Unified egocentric recognition of 3d hand-object poses and interactions

    Bugra Tekin, Federica Bogo, and Marc Pollefeys. H+ o: Unified egocentric recognition of 3d hand-object poses and interactions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4511–4520,

  68. [76]

    Memory-and-anticipation transformer for online action understanding

    Jiahao Wang, Guo Chen, Yifei Huang, Limin Wang, and Tong Lu. Memory-and-anticipation transformer for online action understanding. In Proceedings of the International Conference on Computer Vision (ICCV), 2023. 2

  69. [77]

    Scene-aware ego- centric 3d human pose estimation

    Jian Wang, Diogo Luvizon, Weipeng Xu, Lingjie Liu, Kri- pasindhu Sarkar, and Christian Theobalt. Scene-aware ego- centric 3d human pose estimation. In Proceedings of the Con- ference on Computer Vision and Pattern Recognition (CVPR),

  70. [78]

    Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language model’s pe...

  71. [79]

    Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip H.S. Torr. Fast online object tracking and segmentation: A unifying approach. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 1328–1338, 2019. 3

  72. [80]

    Lavie: High-quality video generation with cascaded latent diffusion models

    Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. Lavie: High-quality video generation with cascaded latent diffusion models. IJCV, 2024. 5

  73. [81]

    In- ternvideo2: Scaling foundation models for multimodal video understanding

    Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Zun Wang, Yansong Shi, Tianxiang Jiang, Songze Li, Jilan Xu, Hongjie Zhang, Yifei Huang, Yu Qiao, Yali Wang, and Limin Wang. In- ternvideo2: Scaling foundation models for multimodal vide...

  74. [82]

    Vide- ollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format, 2024

    Yueqian Wang, Xiaojun Meng, Yuxuan Wang, Jianxin Liang, Jiansheng Wei, Huishuai Zhang, and Dongyan Zhao. Vide- ollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format, 2024. 3

  75. [83]

    Gaze-enabled egocentric video summarization via constrained submodular maximiza- tion

    Jia Xu, Lopamudra Mukherjee, Yin Li, Jamieson Warner, James M Rehg, and Vikas Singh. Gaze-enabled egocentric video summarization via constrained submodular maximiza- tion. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 2

  76. [84]

    Retrieval-augmented egocentric video captioning

    Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, and Weidi Xie. Retrieval-augmented egocentric video captioning. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages 13525–13536, 2024. 1, 6

  77. [85]

    Finebio: A fine-grained video dataset of biological experiments with hierarchical annotations

    Takuma Yagi, Misaki Ohashi, Yifei Huang, Ryosuke Furuta, Shungo Adachi, Toutai Mitsuyama, and Yoichi Sato. Finebio: A fine-grained video dataset of biological experiments with hierarchical annotations. In Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Reco...

  78. [86]

    Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition

    Lijin Yang, Yifei Huang, Yusuke Sugano, and Yoichi Sato. Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1

  79. [87]

    Deco: Decomposition and reconstruction for compositional temporal grounding via coarse-to-fine contrastive ranking

    Lijin Yang, Quan Kong, Hsuan-Kung Yang, Wadim Kehl, Yoichi Sato, and Norimasa Kobori. Deco: Decomposition and reconstruction for compositional temporal grounding via coarse-to-fine contrastive ranking. In Proceedings of the Con- ference on Computer Vision and Pattern Recogniti...

  80. [88]

    Basictad: an astounding rgb-only baseline for temporal action detection

    Min Yang, Guo Chen, Yin-Dong Zheng, Tong Lu, and Limin Wang. Basictad: an astounding rgb-only baseline for temporal action detection. Computer Vision and Image Understanding, 232:103692, 2023. 3

  81. [89]

    Flash-vstream: Memory- based real-time understanding for long video streams, 2024

    Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Jiashi Feng, Jifeng Dai, and Xiaojie Jin. Flash-vstream: Memory- based real-time understanding for long video streams, 2024. 3

  82. [90]

    Fine-grained egocentric hand-object segmentation: Dataset, model, and applications

    Lingzhi Zhang, Shenghao Zhou, Simon Stent, and Jianbo Shi. Fine-grained egocentric hand-object segmentation: Dataset, model, and applications. In Proceedings of the European Conference on Computer Vision (ECCV), 2022. 2

  83. [91]

    Masked video and body-worn imu autoencoder for egocentric action recognition

    Mingfang Zhang, Yifei Huang, Ruicong Liu, and Yoichi Sato. Masked video and body-worn imu autoencoder for egocentric action recognition. In European Conference on Computer Vision, pages 312–330, 2024. 2

  84. [92]

    Internlm-xcomposer2.5-omnilive: A comprehensive multimodal system for long-term streaming video and audio interactions, 2024

    Pan Zhang, Xiaoyi Dong, Yuhang Cao, Yuhang Zang, Rui Qian, Xilin Wei, Lin Chen, Yifei Li, Junbo Niu, Shuangrui Ding, Qipeng Guo, Haodong Duan, Xin Chen, Han Lv, Zheng Nie, Min Zhang, Bin Wang, Wenwei Zhang, Xinyue Zhang, Jiaye Ge, Wei Li, Jingwen Li, Zhongying Tu, Conghui He, ...

  85. [93]

    Ego- body: Human body shape and motion of interacting people from head-mounted devices

    Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. Ego- body: Human body shape and motion of interacting people from head-mounted devices. In Proceedings of the European Conference on Computer Vision (ECCV), 2022. 2

  86. [94]

    Training a large video model on a single machine in a day

    Yue Zhao and Philipp Kr ¨ahenb¨uhl. Training a large video model on a single machine in a day. arXiv preprint arXiv:2309.16669, 2023. 1

  87. [95]

    Learning video representations from large language mod- els

    Yue Zhao, Ishan Misra, Philipp Kr ¨ahenb¨uhl, and Rohit Gird- har. Learning video representations from large language mod- els. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3

  88. [96]

    Learning video representations from large language models

    Yue Zhao, Ishan Misra, Philipp Kr ¨ahenb¨uhl, and Rohit Gird- har. Learning video representations from large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6586–6597,

  89. [97]

    Mrsn: Multi-relation support network for video action detec- tion

    Yin-Dong Zheng, Guo Chen, Minglei Yuan, and Tong Lu. Mrsn: Multi-relation support network for video action detec- tion. In 2023 IEEE International Conference on Multimedia and Expo (ICME), pages 1026–1031. IEEE, 2023. 3

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.