Pith. sign in

REVIEW 4 major objections 4 minor 2 cited by

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The view that best predicts a view-agnostic narration of an activity is the most informative view, and a selector trained on this language proxy needs no best-view labels.

desk verdict A clean weakly-supervised view selection idea with genuine novelty, but the independent human validation is thinner than the claims; deserves a serious review with requests for code, variance, and an unseen captioner. read the letter →

arxiv 2411.08753 v4 pith:OKAJ6JF7 submitted 2024-11-13 cs.CV

classification cs.CV
keywords viewselectionmulti-viewvideoweaksupervisioncaptioningpseudo-labelinginstructionalvideosrelativecameraposeegocentric
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tackles a practical problem: of the several cameras recording the same instructional activity, which view should a viewer watch at each moment. Its central claim is that the camera view whose per-view caption most closely matches a view-agnostic narration of the activity is the view a human would find most informative. To turn that claim into a method, LangView first finetunes captioners on the target datasets, ranks every view by how well its predicted caption scores against the narration with CIDEr, and aggregates ranks across captioners into best-view pseudo-labels. A view selector is then trained on those pseudo-labels with an auxiliary relative-camera-pose prediction task, so that at inference it needs only the multi-view video. The paper reports that this weakly supervised model outperforms heuristic baselines and prior view-selection methods on Ego-Exo4D and LEMMA, both in automatic metrics and in pairwise human preference.

What carries the argument

The load-bearing object is the best-view pseudo-labeler. For each training clip it produces one captioned version per view using three off-the-shelf video captioners, scores every caption against the view-agnostic narration using the CIDEr metric, ranks the views per captioner, and aggregates the top-ranked views into a (possibly multi-element) pseudo-label set. The second piece is the view selector: a shared TimeSformer visual encoder feeding a view classification head and a relative camera pose prediction head; pose prediction is posed as classification over discretized relative-pose angle bins for every view pair, and acts as a regularizer that prevents the encoder from collapsing all views into narration-predictive but viewpoint-insensitive representations. The training loss is the min over pseudo-labels of the cross-entropy of the predicted view, plus a weighted pose classification loss.

What would settle it

Take a held-out set of multi-view clips with narrations, compute the pseudo-labeler's per-view CIDEr scores, and collect human pairwise preferences among all views, not just the extremes. If the top-scored view loses to a lower-scored view at or above chance, or if the CIDEr ranking agrees with human preference no better than the Hand-object or Body-area heuristics, the core proxy is not doing the work. A cheaper quantitative check is to compare the pseudo-labeler's top view against a human-labeled best view on a dataset with such labels; chance-level agreement would refute the claim.

Watch

Extended reading notes

Core claim

The discovery the paper aims to establish is that language can isolate the informative viewpoint in a multi-view instructional video: the more accurately an individual view predicts a view-agnostic text summary of the activity, the more informative that view is. The paper operationalizes this by using an ensemble of video captioners to caption each view independently, scoring each caption against the ground-truth narration with CIDEr, and taking the consensus top-ranked views as pseudo-labels. It then shows that a view classifier can be trained on these pseudo-labels alone, provided it is regularized by an auxiliary relative camera pose predictor that keeps the visual features sensitive to viewpoint. The result is a selector that, at test time, consumes only the multi-view video and returns a best view per clip, with reported gains over heuristics, caption-length scoring, and prior cinematographic baselines on both evaluation datasets.

Load-bearing premise

The entire method assumes that the CIDEr score between a view's predicted caption and the view-agnostic narration reliably tracks how informative that view would be for a human viewer; if captioners describe views in ways that do not reflect what is visible, the pseudo-labels and the selection model trained on them are wrong.

Editorial extensions

If this is right

  • Best-view labels can be replaced by view-agnostic narrations wherever such narrations already exist, removing the most expensive annotation bottleneck in view selection.
  • The auxiliary relative-pose task measurably improves selection: in the paper's ablations, dropping it hurts performance on most metrics, while the rank aggregator mainly protects noun and noun-chunk overlap.
  • The approach transfers to both a multi-exo setup (Ego-Exo4D) and a single-exo setup (LEMMA), so it is not tied to a particular camera topology.
  • At inference the system is text-free and pose-free: language and camera geometry are used only to manufacture training signal.
  • Pseudo-label quality depends on finetuned captioners; with frozen off-the-shelf captioners the paper reports near-zero CIDEr pseudo-labels, so captioner finetuning on the target domain is a required step.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: a natural next test is replacing CIDEr with a learned or human-calibrated caption-to-narration similarity, since captioning metrics reward lexical overlap and might rank views differently when captioners are verbose or paraphrastic.
  • Beyond the paper: the same 'predictive accuracy against view-agnostic text' recipe should carry over to other multi-view settings with descriptions, such as broadcast sports, surveillance with operator logs, or user-generated event coverage.
  • Beyond the paper: because the pseudo-labeler is an ensemble of captioners, the whole pipeline can inherit future captioner improvements without architectural change, so the selector should improve as captioners get better at naming visible entities and interactions.
  • Beyond the paper: the min-over-pseudo-labels loss implicitly assumes the selector should pick the easiest-to-predict among the plausible best views; a ranking loss over the captioners' numeric view scores might make better use of the pseudo-labeler's output.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes LangView, a weakly supervised approach to best-view selection in multi-view instructional videos. Given a multi-view clip and a view-agnostic narration, a pseudo-labeler uses off-the-shelf video captioners (Video-Llama with two LLM decoders and VideoChat2), finetuned on the target datasets, to caption each view; the views are ranked by CIDEr similarity between the predicted caption and the ground-truth narration, and a rank aggregator produces a best-view pseudo-label set. A view selector, built on an EgoVLPv2/TimeSformer encoder with a view classification head and an auxiliary relative-camera-pose classification head, is trained with a min-over-pseudo-labels cross-entropy loss plus a weighted pose loss. At inference, only the multi-view video is used, without narrations or camera poses. Experiments on Ego-Exo4D and LEMMA report automatic captioning metrics (CIDEr, METEOR, V/N/NC-IoU) and a pairwise human study, claiming consistent improvement over baselines.

Significance. If the central proxy holds, the paper makes a useful contribution: it replaces expensive best-view labels with widely available view-agnostic narrations and shows a concrete way to mine those narrations for view-selection supervision. The inclusion of a human evaluation is a genuine strength, as it provides a signal that is not directly aligned with the pseudo-labeling objective. The relative-pose auxiliary task is a sensible mechanism for increasing view sensitivity, and the supplementary material is thorough, including a 3-fold evaluation, pseudo-labeling cost analysis, and additional ablations. The main weakness is that the independent evidence for the core hypothesis is currently thin: the human study is small and tests only extreme pairs, while the automatic metrics share the same captioner families, ground-truth narrations, and metric family as the training signal.

major comments (4)
  1. [Sec. 3.2 and Sec. 4.1 (evaluation metrics)] The automatic evaluation is partially aligned with the training signal by construction. The pseudo-labeler ranks views using CIDEr between per-view captions and the ground-truth narration (Sec. 3.2), and the automatic metrics (CIDEr, METEOR, V-IoU, N-IoU, NC-IoU) are computed by captioning the selected view with the same captioner families (Video-Llama, VideoChat2) and comparing to the same ground-truth narrations. A selector that optimizes the pseudo-labeling objective will therefore tend to score higher on these metrics even if it does not improve human informativeness. The authors should provide an automatic evaluation that is more independent of the training signal, for example by using a captioner that was not used in pseudo-labeling, or by reporting metrics that do not involve caption-narration agreement at all.
  2. [Table 2 (human evaluation)] The human evaluation results are more modest than the text claims. For pseudo-label quality, the best view wins only 53.3% and 46.7% on Ego-Exo4D and LEMMA, with loss rates of 28.9% and 38.9%. For view prediction on LEMMA, the win rate over Hand-object is 43.3% versus a 41.1% loss rate, a difference of only 2.2 percentage points. These numbers do not support the unqualified claim that 'our selected views are preferred significantly more than the two top baselines.' The paper should report confidence intervals or per-participant analysis, and should temper the claim for the LEMMA Hand-object comparison.
  3. [Supp. Table 4 and Sec. 6.2] The ablation without captioner finetuning yields near-collapse (CIDEr 0.4, METEOR 12.2, V-IoU 1.4), showing that the pseudo-labeling pipeline is entirely dependent on finetuning the captioners on target-dataset narrations. This dependency is load-bearing for the claimed 'weak supervision from language': the narrations are used not only as a ranking signal but also to finetune the captioners themselves. The paper should explicitly acknowledge this and discuss the extent to which the method's success relies on having a large set of in-domain clip-narration pairs for captioner finetuning.
  4. [Table 1 (main results)] Gains over the strongest baselines are small on several metrics: on Ego-Exo4D, CIDEr is 13.5 vs. 12.9 for Body-area and METEOR is 48.4 vs. 48.2; on LEMMA, CIDEr is 42.7 vs. 42.1 and METEOR is 74.4 vs. 73.8. No error bars, standard deviations, or number of runs are reported, so it is unclear whether these differences are within noise. The authors should report variance (e.g., bootstrapped confidence intervals or multiple training runs) for the main table, or at least for the comparison against the strongest baseline.
minor comments (4)
  1. [Throughout] There are several typos: 'c.f.' should be 'cf.' (Sec. 3.4 and elsewhere), 'we against observe' appears in Supp. Sec. 6.2, and 'and and' appears in Sec. 4.2 in the qualitative examples paragraph.
  2. [Table 2 caption] The caption states 'Significance, p ≤ 0.05' but does not name the statistical test or how ties were handled in the significance computation; please specify the test and the tie-handling procedure.
  3. [Abstract and Sec. 4.2] The phrase 'state-of-the-art baselines' overstates the baseline set, which includes simple heuristics (Ego-only, Random, Random-exo). Suggest rephrasing to 'a range of baselines, including heuristics and prior methods.'
  4. [Supp. Sec. 10.2.1] The captioner training uses 'a maximum of 1.6 million iterations' for both datasets; please clarify whether this is the same budget for LEMMA and Ego-Exo4D given their very different sizes, and whether early stopping was applied.

Circularity Check

1 steps flagged · score 4.0 of 10

Automatic metrics are coupled to the pseudo-label objective; the key proxy claim rests mainly on the small, extreme-pair human study.

  1. fitted input called prediction [Sec. 3.2 (Best view pseudo-labeler L) and Sec. 4.1 (Evaluation metrics)]
    "we predict the narrations for the N views separately using each captioner ... scores its narrations by comparing them to the ground-truth narration N∗ using a standard captioning metric [7,84,108] ... we use a state-of-the-art video captioner [64,122] to predict the narrations given our chosen views, and then compare the predictions with the view-agnostic ground-truth narrations through standard captioning metrics: CIDEr [108] and METEOR [7]."

    The pseudo-labeler defines the training target as CIDEr agreement between a per-view caption and the view-agnostic narration; the automatic evaluation measures the same CIDEr/METEOR agreement, using the same captioner families (Video-Llama [122], VideoChat2 [64]), for the view the selector outputs. Thus the automatic metrics largely test whether the selector generalizes the pseudo-label objective to new clips, not whether that objective matches human informativeness. The paper's own evaluation sentence ('the more the view deemed as "best" by our model predicts things consistent with the comprehensive view-agnostic ground truth, the better it is') restates the pseudo-labeler's scoring rule.

full rationale

The framework is not derivationally circular at the level of the selector: the pseudo-labeler and the selector are separate modules, and at inference the selector takes only multi-view video and never sees narrations or CIDEr scores. The human study is also independent of narrations and provides some external evidence. However, the automatic evaluation is coupled to the training signal by construction: the pseudo-labels are generated by comparing per-view captions with view-agnostic narrations via captioning metrics, and the automatic metrics reuse the same captioner families and the same caption-narration comparison on the selected view. This means the quantitative results in Table 1 mostly confirm that the selector learns to optimize the pseudo-label objective, rather than independently validating that the objective identifies human-preferred views. The paper further finetunes the captioners on the target datasets before scoring (Sec. 4.1), and the supplementary ablation shows that without this finetuning the pseudo-labels collapse (Table 4), which strengthens the coupling. The human study is the only non-circular validation of the core proxy claim, but it tests only extreme pairs (best vs. worst pseudo-label views) and shows modest win rates (53.3% and 46.7%), so it provides limited independent support. No load-bearing self-citation or imported uniqueness theorem appears; citations to the authors' prior work are for clip pairing, pretraining, and complementary research only. Overall, the central claim retains independent content through the human evaluation, but the automatic quantitative headline overstates independent confirmation, yielding a moderate circularity score of 4.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The method introduces no new physical or ontological entities. Its burdens are concentrated in the pseudo-labeling hypothesis and the dependence on finetuned captioners, which are captured as axioms and free parameters rather than invented entities.

free parameters (4)
  • pose loss weight w = 0.5
    Set on a disjoint validation set to balance view classification and pose prediction losses (Sec. 4.1 Implementation).
  • relative pose bin size = 30 degrees
    Chosen on validation set for discretizing relative camera pose direction labels (Sec. 4.1 Implementation).
  • number of captioners K = 3
    Design choice; additional ablations show that more captioners generally improves performance (Supp. Table 7).
  • captioner finetuning budget = up to 1.6 million iterations
    The pseudo-label quality depends heavily on finetuning the captioners on the target datasets; without finetuning, performance collapses to CIDEr 0.4 (Supp. Table 4). This is a fitted, data-dependent component rather than a fixed constant.
assumptions (4)
  • domain assumption Narrations in Ego-Exo4D and LEMMA are view-agnostic and provide reliable ground truth for the activity content.
    The entire pseudo-labeling pipeline compares per-view captions to these narrations (Sec. 3.1 and 3.2). If narrations are biased toward any particular view, the pseudo-labels would be biased.
  • ad hoc to paper The view whose predicted caption best matches the narration is the most informative view for a human observer.
    This is the paper's core hypothesis (abstract and Sec. 3.2). The human evaluation in Table 2 partially validates it, but only on a small sample of 70 pairs per dataset.
  • domain assumption Off-the-shelf video captioners, after finetuning on the target dataset, produce captions sensitive enough to viewpoint changes to rank views reliably.
    Supp. Table 4 shows that frozen captioners yield nearly useless pseudo-labels (CIDEr 0.4), so the success of the method depends on this finetuning assumption being true.
  • domain assumption EgoVLPv2 features pretrained on Ego-Exo4D transfer to both datasets and support view discrimination.
    EgoVLPv2 is used as the visual encoder for the view selector (Sec. 10.2.2). No ablation of this choice is provided, yet the model's ability to distinguish views depends on these features.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos." pith.science (2026). https://pith.science/paper/OKAJ6JF7

@misc{pith2026241108753,
  author       = {Pith},
  title        = {Pith review of: Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OKAJ6JF7}},
  note         = {Machine review of arXiv:2411.08753}
}
read the original abstract

Given a multi-view video, which viewpoint is most informative for a human observer? Existing methods rely on heuristics or expensive "best-view" supervision to answer this question, limiting their applicability. We propose a weakly supervised approach that leverages language accompanying an instructional multi-view video as a means to recover its most informative viewpoint(s). Our key hypothesis is that the more accurately an individual view can predict a view-agnostic text summary, the more informative it is. To put this into action, we propose LangView, a framework that uses the relative accuracy of view-dependent caption predictions as a proxy for best view pseudo-labels. Then, those pseudo-labels are used to train a view selector, together with an auxiliary camera pose predictor that enhances view-sensitivity. During inference, our model takes as input only a multi-view video--no language or camera poses--and returns the best viewpoint to watch at each timestep. On two challenging datasets comprised of diverse multi-camera setups and how-to activities, our model consistently outperforms state-of-the-art baselines, both with quantitative metrics and human evaluation. Project page: https://vision.cs.utexas.edu/projects/which-view-shows-it-best.

Figures

Figures reproduced from arXiv: 2411.08753 by the authors.

Figure 1
Figure 1. LANGVIEW idea: given multi-view instructional videos, we aim to learn a view selection model that can identify the best view for seeing how to perform the activity shown in the videos, in the absence of best view labels. To achieve this, we compare each estimated view-dependent caption to the view-agnostic ground￾truth video narration of the human activity, and use their respective accuracies as a proxy for view qua… view at source ↗
Figure 2
Figure 2. (a) Our model uses language guidance to train a view-selector for multi-view instructional videos, such that the chosen views help best understand the shown activity. To do so, we first generate best view pseudo-labels during training by leveraging clip narrations, where each narration is a view-agnostic and detailed description of the activity. Specifically, given a training clip, we use off-the-shelf video caption… view at source ↗
Figure 3
Figure 3. Left: sample successful predictions by our view selector. For each clip, our model chooses the view that shows the action, and the objects and body parts involved in it, most clearly, and hence, is most informative. Right: Sample failure cases for our model, where there are multiple high-quality views that differ only in certain nuances, which are discernible by a human but not our model trained through narration gu… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: t-SNE [106] plots of exo visual features of sample Ego-Exo4D [37] videos from basketball, bike repair, dance and cooking scenarios. Our model, when trained with the relative camera pose predictor, produces visual features that form neater clusters when grouped on the b…
Figure 5
Figure 5. Figure 5: Our model’s attention heatmaps on two best view clips from Ego-Exo4D [37]. Yellow patches indicate highest attention. exo views with almost equal likelihood. However, for LEMMA, our model tends to prefer the ego view much more than the exo view, re-emphasizing the prev…
Figure 6
Figure 6. Figure 6: Examples of predicted narrations, and the ranks and scores of the views, per our pseudo-labeler L, shown alongside ground-truth narrations, in addition to what is provided in Sec. 3.2 in main. Best (0.84) Worst (0.16) 1. Best (0.42) Worst (0.04) 2. Best (0.67) Worst (0…
Figure 7
Figure 7. Figure 7: Additional examples of best and worst views, and their scores, per our pseudo-labeler L. (‘Dataset’ in Sec. 4.1 in main). In addition to the ones provided in Fig. 2b in main, we show more pseudo-labeler outputs, comprising view ranks and predicted narrations, alongside…
Figure 8
Figure 8. Figure 8: Test CIDEr difference between our model and the Body-area [57] baseline vs. verb-noun pair frequency in train narrations, sorted in decreasing order Model CIDEr METEOR V-IoU N-IoU NC-IoU Body-area 10.5 46.6 30.0 35.2 30.4 Ours 11.4 46.9 31.2 37.0 31.9 [PITH_FULL_IMAGE…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

    cs.CV 2024-12 conditional novelty 6.0 of 10

    The authors train a view-switch predictor on pseudo-labeled web videos and show it transfers to selecting ego/exo views in new multi-view videos with limited labels.

  2. Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision

    cs.CV 2025-06 accept novelty 3.0 of 10

    A comprehensive review of cross-view video understanding that uses both first-person and third-person cameras, organized into a three-direction taxonomy with a dataset catalog and future research gaps.

Reference graph

Works this paper leans on

132 extracted references · 51 canonical work pages · cited by 2 Pith papers

  1. [1]

    Deep learning using rectified linear units (relu), 2018

    Abien Fred Agarap. Deep learning using rectified linear units (relu), 2018. cite arxiv:1803.08375Comment: 7 pages, 11 figures, 9 tables. 19

  2. [2]

    McCrae, Kenton Murray, Maria Nadejde, Satoshi Nakamura, Matteo Negri, Ha Nguyen, Jan Niehues, Xing Niu, Atul Kr

    Milind Agarwal, Sweta Agrawal, Antonios Anastasopou- los, Luisa Bentivogli, Ondˇrej Bojar, Claudia Borg, Marine Carpuat, Roldano Cattoni, Mauro Cettolo, Mingda Chen, William Chen, Khalid Choukri, Alexandra Chronopoulou, Anna Currey, Thierry Declerck, Qianqian Dong, Kevin Duh, Yannick Estève, Marcello Federico, Souhir Gahbiche, Barry Haddow, Benjamin Hsu, ...

  3. [3]

    A dataset for develop- ing and benchmarking active vision

    Phil Ammirato, Patrick Poirson, Eunbyung Park, Jana Košecká, and Alexander C Berg. A dataset for develop- ing and benchmarking active vision. In 2017 IEEE Inter- national Conference on Robotics and Automation (ICRA), pages 1378–1385. IEEE, 2017. 3

  4. [4]

    Automatic editing of footage from multi- ple social cameras

    Ido Arev, Hyun Soo Park, Yaser Sheikh, Jessica Hodgins, and Ariel Shamir. Automatic editing of footage from multi- ple social cameras. ACM Trans. Graph., 33(4), 2014. 2

  5. [5]

    R. Bajcsy. Active perception. Proceedings of the IEEE, 76 (8):966–1005, 1988. 2

  6. [6]

    Weaqa: Weak supervision via captions for visual question answering

    Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, and Chitta Baral. Weaqa: Weak supervision via captions for visual question answering. arXiv preprint arXiv:2012.02356, 2020. 3

  7. [7]

    METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments

    Satanjeev Banerjee and Alon Lavie. METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Work- shop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan, 2005. Association for Computational Linguistics. 4, 6, 7, 8, 15, 19

  8. [8]

    Is space-time attention all you need for video understanding? In ICML, page 4, 2021

    Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, page 4, 2021. 5, 18

Show all 132 references
  1. [9]

    High- lightme: Detecting highlights from human-centric videos

    Uttaran Bhattacharya, Gang Wu, Stefano Petrangeli, Viswanathan Swaminathan, and Dinesh Manocha. High- lightme: Detecting highlights from human-centric videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8157–8167, 2021. 2

  2. [10]

    Extreme rotation estimation using dense correlation volumes

    Ruojin Cai, Bharath Hariharan, Noah Snavely, and Hadar Averbuch-Elor. Extreme rotation estimation using dense correlation volumes. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 14566–14575, 2021. 5

  3. [11]

    Davis, and Lei Zhang

    Sijia Cai, Wangmeng Zuo, Larry S. Davis, and Lei Zhang. Weakly-supervised video summarization using variational encoder-decoder and web prior. In Computer Vision – ECCV 2018 - 15th European Conference, 2018, Proceedings, pages 193–210. Springer-Verlag, 2018. 15th European Conf...

  4. [12]

    Enhanced interactive 360° viewing via automatic guidance

    Seunghoon Cha, Jungjin Lee, Seunghwa Jeong, Younghui Kim, and Junyong Noh. Enhanced interactive 360° viewing via automatic guidance. ACM Trans. Graph., 39(5), 2020. 6, 7, 18, 19, 20

  5. [13]

    Learn- ing sports camera selection from internet videos

    Jianhui Chen, Keyu Lu, Sijia Tian, and Jim Little. Learn- ing sports camera selection from internet videos. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1682–1691. IEEE, 2019. 2

  6. [14]

    Wide- baseline relative camera pose estimation with directional learning

    Kefan Chen, Noah Snavely, and Ameesh Makadia. Wide- baseline relative camera pose estimation with directional learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3258–3268,

  7. [15]

    Geometry-aware recurrent neural networks for active visual recognition

    Ricson Cheng, Ziyan Wang, and Katerina Fragkiadaki. Geometry-aware recurrent neural networks for active visual recognition. Advances in Neural Information Processing Systems, 31, 2018. 3

  8. [16]

    Towards a richer 2d understanding of hands at scale

    Tianyi Cheng, Dandan Shan, Ayda Sultan Hassen, Richard Ely Locke Higgins, and David Fouhey. Towards a richer 2d understanding of hands at scale. In Thirty-seventh Confer- ence on Neural Information Processing Systems, 2023. 6, 7, 19 9

  9. [17]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 6

  10. [18]

    Self-view ground- ing given a narrated 360 {\deg} video

    Shih-Han Chou, Yi-Chun Chen, Kuo-Hao Zeng, Hou- Ning Hu, Jianlong Fu, and Min Sun. Self-view ground- ing given a narrated 360 {\deg} video. arXiv preprint arXiv:1711.08664, 2017. 2

  11. [19]

    Video co-summarization: Video summarization by visual co- occurrence

    Wen-Sheng Chu, Yale Song, and Alejandro Jaimes. Video co-summarization: Video summarization by visual co- occurrence. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3584–3592, 2015. 2

  12. [20]

    elochoice

    Andrew P. Clark, Kate L. Howard, Andy T. Woods, Ian S. Penton-V oak, and Christof Neumann. Why rate when you could compare? using the “elochoice” package to assess pairwise comparisons of perceived physical strength. PLOS ONE, 13(1):1–16, 2018. 6

  13. [21]

    Scaling egocentric vision: The epic- kitchens dataset

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Da- vide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Scaling egocentric vision: The epic- kitchens dataset. In European Conference on Computer Vi...

  14. [22]

    Flashattention-2: Faster attention with bet- ter parallelism and work partitioning

    Tri Dao. Flashattention-2: Faster attention with bet- ter parallelism and work partitioning. arXiv preprint arXiv:2307.08691, 2023. 18

  15. [23]

    Flashattention: Fast and memory-efficient exact attention with io-awareness

    Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christo- pher Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in Neural Informa- tion Processing Systems, 35:16344–16359, 2022. 18

  16. [24]

    Velastin

    Hannah Mary Dee and Sergio A. Velastin. How close are we to solving the problem of automated visual surveillance? a review of real-world surveillance, scientific progress and evaluative mechanisms. Machine Vision and Applications, 19(5-6):329–343, 2008. Dee, H. M.; Velastin, S...

  17. [25]

    Virtex: Learning visual representations from textual annotations

    Karan Desai and Justin Johnson. Virtex: Learning visual representations from textual annotations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11162–11173, 2021. 3

  18. [26]

    An image is worth 16x16 words: Trans- formers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint a...

  19. [27]

    Dense and aligned captions (dac) promote compositional reasoning in vl models

    Sivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig, Donghyun Kim, Paola Cascante-Bonilla, Amit Alfassy, Rameswar Panda, Raja Giryes, Rogerio Feris, et al. Dense and aligned captions (dac) promote compositional reasoning in vl models. Advances in Neural Information Processin...

  20. [28]

    Multi-view active fine- grained visual recognition

    Ruoyi Du, Wenqing Yu, Heqing Wang, Ting-En Lin, Dongliang Chang, and Zhanyu Ma. Multi-view active fine- grained visual recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1568–1578, 2023. 3

  21. [29]

    Multi- stream dynamic video summarization

    Mohamed Elfeki, Liqiang Wang, and Ali Borji. Multi- stream dynamic video summarization. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 339–349, 2022. 2

  22. [30]

    Elson and Mark O

    David K. Elson and Mark O. Riedl. A lightweight intelligent virtual cinematography system for machinima production. In Proceedings of the Third AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment , page 8–13. AAAI Press, 2007. 2

  23. [31]

    Foote and D

    J. Foote and D. Kimber. Flycam: practical panoramic video and automatic camera control. In 2000 IEEE International Conference on Multimedia and Expo. ICME2000. Proceed- ings. Latest Advances in the Fast Changing World of Multi- media (Cat. No.00TH8532), pages 1419–1422 vol.3, 2000. 2

  24. [32]

    Multi-view video summa- rization

    Yanwei Fu, Yanwen Guo, Yanshu Zhu, Feng Liu, Chuan- ming Song, and Zhi-Hua Zhou. Multi-view video summa- rization. IEEE Transactions on Multimedia, 12(7):717–729,

  25. [33]

    Gleicher, Rachel M

    Michael L. Gleicher, Rachel M. Heck, and Michael N. Wal- lick. A framework for virtual videography. In Proceedings of the 2nd International Symposium on Smart Graphics , page 9–16, New York, NY , USA, 2002. Association for Computing Machinery. 2

  26. [34]

    Peavs: Perceptual evaluation of audio-visual syn- chrony grounded in viewers’ opinion scores

    Lucas Goncalves, Prashant Mathur, Chandrashekhar La- vania, Metehan Cekic, Marcello Federico, and Kyu J Han. Peavs: Perceptual evaluation of audio-visual syn- chrony grounded in viewers’ opinion scores. arXiv preprint arXiv:2404.07336, 2024. 6

  27. [35]

    Diverse sequential subset selection for supervised video summarization

    Boqing Gong, Wei-Lun Chao, Kristen Grauman, and Fei Sha. Diverse sequential subset selection for supervised video summarization. In Advances in Neural Information Process- ing Systems. Curran Associates, Inc., 2014. 2

  28. [36]

    Ego4d: Around the world in 3,000 hours of egocentric video

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. InPro- ceedings of the IEEE/CVF Conference on Computer Vision...

  29. [37]

    Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. arXiv preprint...

  30. [38]

    Temporal difference varia- tional auto-encoder

    Karol Gregor, George Papamakarios, Frederic Besse, Lars Buesing, and Theophane Weber. Temporal difference varia- tional auto-encoder. arXiv preprint arXiv:1806.03107, 2018. 5

  31. [39]

    From images to textual prompts: Zero-shot vqa with frozen large language models

    Jiaxian Guo, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Boyang Li, Dacheng Tao, and Steven CH Hoi. From images to textual prompts: Zero-shot vqa with frozen large language models. arXiv preprint arXiv:2212.10846, 2022. 3 10

  32. [40]

    Using closed captions as supervision for video activity recognition

    Sonal Gupta and Raymond Mooney. Using closed captions as supervision for video activity recognition. Proceedings of the AAAI Conference on Artificial Intelligence , 24(1): 1083–1088, 2010. 3

  33. [41]

    Creating summaries from user videos

    Michael Gygli, Helmut Grabner, Hayko Riemenschneider, and Luc Van Gool. Creating summaries from user videos. In Computer Vision – ECCV 2014, pages 505–520, Cham,

  34. [42]

    Video summarization by learning submodular mixtures of objec- tives

    Michael Gygli, Helmut Grabner, and Luc Van Gool. Video summarization by learning submodular mixtures of objec- tives. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3090–3098, 2015. 2

  35. [43]

    Align and attend: Multimodal summarization with dual contrastive losses

    Bo He, Jun Wang, Jielin Qiu, Trung Bui, Abhinav Shrivas- tava, and Zhaowen Wang. Align and attend: Multimodal summarization with dual contrastive losses. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14867–14878, 2023. 2

  36. [44]

    Cohen, and David H

    Li-wei He, Michael F. Cohen, and David H. Salesin. The virtual cinematographer: a paradigm for automatic real-time camera control and directing. In Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques, page 217–224, New York, NY , USA, 1996...

  37. [45]

    Cohen, and David H

    Li-wei He, Michael F. Cohen, and David H. Salesin. The Virtual Cinematographer: A Paradigm for Automatic Real- Time Camera Control and Directing. Association for Com- puting Machinery, New York, NY , USA, 1 edition, 2023. 2

  38. [46]

    Vir- tual videography

    Rachel Heck, Michael Wallick, and Michael Gleicher. Vir- tual videography. In Proceedings of the 14th ACM Interna- tional Conference on Multimedia, page 961–962, New York, NY , USA, 2006. Association for Computing Machinery. 2

  39. [47]

    Lora: Low-rank adaptation of large language models

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021. 18

  40. [48]

    Deep 360 pilot: Learning a deep agent for piloting through 360deg sports videos

    Hou-Ning Hu, Yen-Chen Lin, Ming-Yu Liu, Hsien-Tzu Cheng, Yung-Ju Chang, and Min Sun. Deep 360 pilot: Learning a deep agent for piloting through 360deg sports videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3451–3460, 2017. 2

  41. [49]

    Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activi- ties in real world

    Yifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang, Li- jin Yang, Baoqi Pei, Hongjie Zhang, Lu Dong, Yali Wang, Limin Wang, et al. Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activi- ties in real world. arXiv preprint arXiv:2403.16182, ...

  42. [50]

    Batch normalization: accelerating deep network training by reducing internal co- variate shift

    Sergey Ioffe and Christian Szegedy. Batch normalization: accelerating deep network training by reducing internal co- variate shift. In Proceedings of the 32nd International Con- ference on International Conference on Machine Learning - Volume 37, page 448–456. JMLR.org, 2015. 19

  43. [51]

    Look-ahead be- fore you leap: end-to-end active recognition by forecasting the effect of motion

    Dinesh Jayaraman and Kristen Grauman. Look-ahead be- fore you leap: end-to-end active recognition by forecasting the effect of motion. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, Octo- ber 11-14, 2016, Proceedings, Part V 14 , pages 489–...

  44. [52]

    Learning to look around: Intelligently exploring unseen environments for unknown tasks

    Dinesh Jayaraman and Kristen Grauman. Learning to look around: Intelligently exploring unseen environments for unknown tasks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1238–1247,

  45. [53]

    End-to-end policy learning for active visual categorization

    Dinesh Jayaraman and Kristen Grauman. End-to-end policy learning for active visual categorization. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(7):1601– 1614, 2019. 3

  46. [54]

    Time-agnostic prediction: Predicting pre- dictable video frames

    Dinesh Jayaraman, Frederik Ebert, Alexei A Efros, and Sergey Levine. Time-agnostic prediction: Predicting pre- dictable video frames. arXiv preprint arXiv:1808.07784,

  47. [55]

    Simglim: Simplifying glimpse based active visual reconstruction

    Abhishek Jha, Soroush Seifi, and Tinne Tuytelaars. Simglim: Simplifying glimpse based active visual reconstruction. In Proceedings of the IEEE/CVF Winter Conference on Appli- cations of Computer Vision (WACV), pages 269–278, 2023. 3

  48. [56]

    Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities

    Baoxiong Jia, Yixin Chen, Siyuan Huang, Yixin Zhu, and Song-chun Zhu. Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities. In European Conference on Computer Vision, pages 767–786. Springer, 2020. 2, 4, 5, 7, 16, 18, 20

  49. [57]

    Rtmpose: Real-time multi-person pose estimation based on mmpose

    Tao Jiang, Peng Lu, Li Zhang, Ning Ma, Rui Han, Chengqi Lyu, Yining Li, and Kai Chen. Rtmpose: Real-time multi-person pose estimation based on mmpose. ArXiv, abs/2303.07399, 2023. 6, 18, 19

  50. [58]

    Large-scale video summarization using web-image priors

    Aditya Khosla, Raffay Hamid, Chih-Jen Lin, and Neel Sun- daresan. Large-scale video summarization using web-image priors. In 2013 IEEE Conference on Computer Vision and Pattern Recognition, pages 2698–2705, 2013. 2

  51. [59]

    Gunhee Kim and Eric P. Xing. Reconstructing storyline graphs for image recommendation from web community photos. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, pages 3882–3889, 2014. 2

  52. [60]

    Segment any- thing

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment any- thing. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 4015–4026, 2023. 6, 20

  53. [61]

    Hyperbolic learning with synthetic captions for open-world detection

    Fanjie Kong, Yanbei Chen, Jiarui Cai, and Davide Modolo. Hyperbolic learning with synthetic captions for open-world detection. arXiv preprint arXiv:2404.05016, 2024. 3

  54. [62]

    A memory network approach for story-based temporal summarization of 360° videos

    Sangho Lee, Jinyoung Sung, Youngjae Yu, and Gunhee Kim. A memory network approach for story-based temporal summarization of 360° videos. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1410– 1419, 2018. 2

  55. [63]

    Predicting important objects for egocentric video summarization

    Yong Jae Lee and Kristen Grauman. Predicting important objects for egocentric video summarization. International Journal of Computer Vision, 114(1):38–55, 2015. 2

  56. [64]

    Mvbench: A comprehensive multi-modal video understand- ing benchmark

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understand- ing benchmark. arXiv preprint arXiv:2311.17005, 2023. 2, 3, 4, 5, 6, 15, 18 11

  57. [65]

    Grounded language-image pre-training

    Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jian- wei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, and Jianfeng Gao. Grounded language-image pre-training. In 2022 IEEE/CVF Conference on Computer Vision and Pat- tern Rec...

  58. [66]

    How local is the local diversity? reinforcing sequen- tial determinantal point processes with dynamic ground sets for supervised video summarization

    Yandong Li, Liqiang Wang, Tianbao Yang, and Boqing Gong. How local is the local diversity? reinforcing sequen- tial determinantal point processes with dynamic ground sets for supervised video summarization. In Proceedings of the European Conference on Computer Vision (ECCV), p...

  59. [67]

    Egocentric video-language pretraining

    Kevin Qinghong Lin, Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Z Xu, Difei Gao, Rong-Cheng Tu, Wenzhe Zhao, Weijie Kong, et al. Egocentric video-language pretraining. Advances in Neural Information Processing Systems, 35:7575–7586, 2022. 18

  60. [68]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023. 3, 6, 20

  61. [69]

    Sgdr: Stochastic gradient descent with warm restarts

    Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016. 18

  62. [70]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 18, 20

  63. [71]

    Story-driven summariza- tion for egocentric video

    Zheng Lu and Kristen Grauman. Story-driven summariza- tion for egocentric video. In Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition, page 2714–2721, USA, 2013. IEEE Computer Society. 2

  64. [72]

    Switch-a-view: Few-shot view selection learned from edited videos

    Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, and Kristen Grauman. Switch-a-view: Few-shot view selection learned from edited videos. arXiv preprint arXiv:2412.18386, 2024. 2

  65. [73]

    Video summarization via multi- view representative selection

    Jingjing Meng, Suchen Wang, Hongxing Wang, Yap-Peng Tan, and Junsong Yuan. Video summarization via multi- view representative selection. In 2017 IEEE International Conference on Computer Vision Workshops (ICCVW), pages 1189–1198, 2017. 2

  66. [74]

    Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

    Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF international conference on computer vision, pages ...

  67. [75]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, 2020. 1

  68. [76]

    Automatized summarization of multi- player games

    Peter Mindek, Ladislav ˇCmolík, Ivan Viola, Eduard Gröller, and Stefan Bruckner. Automatized summarization of multi- player games. In Proceedings of the 31st Spring Conference on Computer Graphics, page 73–80, New York, NY , USA,

  69. [77]

    Egoenv: Human- centric environment representations from egocentric video

    Tushar Nagarajan, Santhosh Kumar Ramakrishnan, Ruta Desai, James Hillis, and Kristen Grauman. Egoenv: Human- centric environment representations from egocentric video. Advances in Neural Information Processing Systems , 36: 60130–60143, 2023. 5

  70. [78]

    Tl; dw? summarizing instructional videos with task relevance and cross-modal saliency

    Medhini Narasimhan, Arsha Nagrani, Chen Sun, Michael Rubinstein, Trevor Darrell, Anna Rohrbach, and Cordelia Schmid. Tl; dw? summarizing instructional videos with task relevance and cross-modal saliency. InEuropean Conference on Computer Vision, pages 540–557. Springer, 2022. 2

  71. [79]

    Adaptive skip intervals: Temporal abstraction for recurrent dynamical models

    Alexander Neitz, Giambattista Parascandolo, Stefan Bauer, and Bernhard Schölkopf. Adaptive skip intervals: Temporal abstraction for recurrent dynamical models. Advances in Neural Information Processing Systems, 31, 2018. 5

  72. [80]

    Au- tomatic video summarization by graph modeling

    Chong-Wah Ngo, Yu-Fei Ma, and Hong-Jiang Zhang. Au- tomatic video summarization by graph modeling. In Pro- ceedings Ninth IEEE International Conference on Computer Vision, pages 104–109 vol.1, 2003. 2

  73. [81]

    Collabora- tive summarization of topic-related videos

    Rameswar Panda and Amit K Roy-Chowdhury. Collabora- tive summarization of topic-related videos. In Proceedings of the IEEE Conference on computer vision and pattern recognition, pages 7083–7092, 2017. 2

  74. [82]

    Multi-view surveillance video summarization via joint embedding and sparse optimization

    Rameswar Panda and Amit K Roy-Chowdhury. Multi-view surveillance video summarization via joint embedding and sparse optimization. IEEE Transactions on Multimedia, 19 (9):2010–2021, 2017. 2

  75. [83]

    Roy-Chowdhury

    Rameswar Panda, Abir Das, Ziyan Wu, Jan Ernst, and Amit K. Roy-Chowdhury. Weakly supervised summariza- tion of web videos. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 3677–3686, 2017. 2

  76. [84]

    Bleu: a method for automatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311– 318, Philadelphia, Pennsylvania, USA, 2002. Assoc...

  77. [85]

    Sumgraph: Video summarization via recursive graph modeling

    Jungin Park, Jiyoung Lee, Ig-Jae Kim, and Kwanghoon Sohn. Sumgraph: Video summarization via recursive graph modeling. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceed- ings, Part XXV 16, pages 647–663. Springer, 2020. 2

  78. [86]

    Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

    Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang. Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. In Proceedings of the IEEE/CVF International Conference on Computer Vis...

  79. [87]

    Vloc- net++: Deep multitask learning for semantic visual localiza- tion and odometry

    Noha Radwan, Abhinav Valada, and Wolfram Burgard. Vloc- net++: Deep multitask learning for semantic visual localiza- tion and odometry. IEEE Robotics and Automation Letters, 3(4):4407–4414, 2018. 5

  80. [88]

    Sidekick policy learning for active visual exploration

    Santhosh K Ramakrishnan and Kristen Grauman. Sidekick policy learning for active visual exploration. In Proceedings of the European conference on computer vision (ECCV) , pages 413–430, 2018. 3

  81. [89]

    Emergence of exploratory look-around behaviors through active observation completion

    Santhosh K Ramakrishnan, Dinesh Jayaraman, and Kristen Grauman. Emergence of exploratory look-around behaviors through active observation completion. Science Robotics, 4 (30):eaaw6326, 2019. 3 12

  82. [90]

    Naq: Leveraging narrations as queries to super- vise episodic memory

    Santhosh Kumar Ramakrishnan, Ziad Al-Halah, and Kristen Grauman. Naq: Leveraging narrations as queries to super- vise episodic memory. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6694–6703, 2023. 18

  83. [91]

    Video summarization by learning from unpaired data

    Mrigank Rochan and Yang Wang. Video summarization by learning from unpaired data. In Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition, pages 7902–7911, 2019. 2

  84. [92]

    Adaptive video highlight detection by learning from user history

    Mrigank Rochan, Mahesh Kumar Krishna Reddy, Linwei Ye, and Yang Wang. Adaptive video highlight detection by learning from user history. In Computer Vision – ECCV 2020, pages 261–278, Cham, 2020. Springer International Publishing. 2

  85. [93]

    Chowdhury

    Abhimanyu Sahu and Ananda S. Chowdhury. Shot level egocentric video co-summarization. In 2018 24th Interna- tional Conference on Pattern Recognition (ICPR) , pages 2887–2892, 2018. 2

  86. [94]

    Attend and segment: Attention guided active semantic segmentation

    Soroush Seifi and Tinne Tuytelaars. Attend and segment: Attention guided active semantic segmentation. In European Conference on Computer Vision, pages 305–321. Springer,

  87. [95]

    Glimpse- attend-and-explore: Self-attention for active visual explo- ration

    Soroush Seifi, Abhishek Jha, and Tinne Tuytelaars. Glimpse- attend-and-explore: Self-attention for active visual explo- ration. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 16137–16146, 2021. 3

  88. [96]

    Actor and observer: Joint modeling of first and third-person videos

    Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. Actor and observer: Joint modeling of first and third-person videos. In proceedings of the IEEE conference on computer vision and pattern recognition, pages 7396–7404, 2018. 1

  89. [97]

    Tvsum: Summarizing web videos using titles

    Yale Song, Jordi Vallmitjana, Amanda Stent, and Alejandro Jaimes. Tvsum: Summarizing web videos using titles. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5179–5187, 2015. 2

  90. [98]

    Making 360 ° video watchable in 2d: Learning videography for click free view- ing

    Yu-Chuan Su and Kristen Grauman. Making 360 ° video watchable in 2d: Learning videography for click free view- ing. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1368–1376, 2017. 2

  91. [99]

    Pano2vid: Automatic cinematography for watching 360 videos

    Yu-Chuan Su, Dinesh Jayaraman, and Kristen Grauman. Pano2vid: Automatic cinematography for watching 360 videos. In Asian Conference on Computer Vision , pages 154–171. Springer, 2016. 2

  92. [100]

    Automatic con- cept discovery from parallel text and visual corpora

    Chen Sun, Chuang Gan, and Ram Nevatia. Automatic con- cept discovery from parallel text and visual corpora. In Proceedings of the IEEE international conference on com- puter vision, pages 2596–2604, 2015. 3

  93. [101]

    Foote, D

    Xinding Sun, J. Foote, D. Kimber, and B. S. Manjunath. Re- gion of interest extraction and virtual camera control based on panoramic video capturing. Trans. Multi., 7(5):981–990,

  94. [102]

    Multi-channel attention selection gan with cas- caded semantic guidance for cross-view image translation

    Hao Tang, Dan Xu, Nicu Sebe, Yanzhi Wang, Jason J Corso, and Yan Yan. Multi-channel attention selection gan with cas- caded semantic guidance for cross-view image translation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2417–2426...

  95. [103]

    Coin: A large-scale dataset for comprehensive instructional video analysis

    Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216,

  96. [104]

    Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models.arXiv preprint arXiv:2307.09288, 2023. 4, 6

  97. [105]

    Image captioners are scalable vision learners too

    Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiao- hua Zhai, Neil Houlsby, and Lucas Beyer. Image captioners are scalable vision learners too. Advances in Neural Infor- mation Processing Systems, 36, 2024. 3

  98. [106]

    Visualizing data using t-sne

    Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9 (86):2579–2605, 2008. 16

  99. [107]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 19

  100. [108]

    Lawrence Zitnick, and Devi Parikh

    Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description eval- uation. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4566–4575, 2014. 4, 6, 7, 15, 19

  101. [109]

    Mask- free ovis: Open-vocabulary instance segmentation without manual mask annotations

    Vibashan VS, Ning Yu, Chen Xing, Can Qin, Mingfei Gao, Juan Carlos Niebles, Vishal M Patel, and Ran Xu. Mask- free ovis: Open-vocabulary instance segmentation without manual mask annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,...

  102. [110]

    Mvsgcn: A novel graph convolutional network for multi-video summa- rization

    Jiaxin Wu, Sheng-Hua Zhong, and Yan Liu. Mvsgcn: A novel graph convolutional network for multi-video summa- rization. In Proceedings of the 27th ACM International Conference on Multimedia, page 827–835, New York, NY , USA, 2019. Association for Computing Machinery. 2

  103. [111]

    Be- trayed by captions: Joint caption grounding and generation for open vocabulary instance segmentation

    Jianzong Wu, Xiangtai Li, Henghui Ding, Xia Li, Guan- gliang Cheng, Yunhai Tong, and Chen Change Loy. Be- trayed by captions: Joint caption grounding and generation for open vocabulary instance segmentation. In Proceedings of the IEEE/CVF International Conference on Computer V...

  104. [112]

    Snap angle prediction for 360° panoramas

    Bo Xiong and Kristen Grauman. Snap angle prediction for 360° panoramas. In Proceedings of the European Confer- ence on Computer Vision (ECCV) , 2018. 2, 6, 7, 18, 19, 20

  105. [113]

    Open-vocabulary panop- tic segmentation with text-to-image diffusion models

    Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiao- long Wang, and Shalini De Mello. Open-vocabulary panop- tic segmentation with text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 2955–2966, 2023. 3

  106. [114]

    Learning to answer visual questions from web videos

    Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid. Learning to answer visual questions from web videos. arXiv preprint arXiv:2205.05019, 2022. 3

  107. [115]

    Unsupervised extraction of video highlights via robust recurrent auto-encoders

    Huan Yang, Baoyuan Wang, Stephen Lin, David Wipf, Minyi Guo, and Baining Guo. Unsupervised extraction of video highlights via robust recurrent auto-encoders. In Pro- 13 ceedings of the IEEE international conference on computer vision, pages 4633–4641, 2015. 2

  108. [116]

    Alip: Adaptive language-image pre-training with synthetic caption

    Kaicheng Yang, Jiankang Deng, Xiang An, Jiawei Li, Ziy- ong Feng, Jia Guo, Jing Yang, and Tongliang Liu. Alip: Adaptive language-image pre-training with synthetic caption. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2922–2931, 2023. 3

  109. [117]

    Cap2det: Learning to amplify weak caption supervision for object detection

    Keren Ye, Mingda Zhang, Adriana Kovashka, Wei Li, Dan- feng Qin, and Jesse Berent. Cap2det: Learning to amplify weak caption supervision for object detection. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision (ICCV), 2019. 3

  110. [118]

    Temporal cue guided video highlight detection with low-rank audio-visual fusion

    Qinghao Ye, Xiyue Shen, Yuan Gao, Zirui Wang, Qi Bi, Ping Li, and Guang Yang. Temporal cue guided video highlight detection with low-rank audio-visual fusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 7950–7959, 2021. 2

  111. [119]

    Coca: Contrastive captioners are image-text foundation models

    Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mo- jtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022. 3

  112. [120]

    Open-vocabulary object detection using captions

    Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih- Fu Chang. Open-vocabulary object detection using captions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14393–14402, 2021. 3

  113. [121]

    An automated end-to-end lecture capture and broadcasting sys- tem

    Cha Zhang, Yong Rui, Jim Crawford, and Li-Wei He. An automated end-to-end lecture capture and broadcasting sys- tem. ACM Trans. Multimedia Comput. Commun. Appl., 4 (1), 2008. 2

  114. [122]

    Video-llama: An instruction-tuned audio-visual language model for video understanding

    Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023. 2, 3, 4, 5, 6, 15, 18

  115. [123]

    Multi-view video synopsis via simultaneous object-shifting and view- switching optimization

    Zhensong Zhang, Yongwei Nie, Hanqiu Sun, Qing Zhang, Qiuxia Lai, Guiqing Li, and Mingyu Xiao. Multi-view video synopsis via simultaneous object-shifting and view- switching optimization. IEEE Transactions on Image Pro- cessing, 29:971–985, 2020. 2

  116. [124]

    Towards automatic learning of procedures from web instructional videos

    Luowei Zhou, Chenliang Xu, and Jason Corso. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence, 2018. 4

  117. [125]

    Cross- task weakly supervised learning from instructional videos

    Dimitri Zhukov, Jean-Baptiste Alayrac, Ramazan Gokberk Cinbis, David Fouhey, Ivan Laptev, and Josef Sivic. Cross- task weakly supervised learning from instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3537–3545, 2...

  118. [128]

    6.1), as referenced in ‘Qualitative examples’ in Sec

    Supplementary material In this supplementary material we provide additional details about: • Video (with audio) for qualitative illustration of our task and qualitative assessment of our view predictions (Sec. 6.1), as referenced in ‘Qualitative examples’ in Sec. 4.2 in main •...

  119. [129]

    5, we provide examples of our model’s attention heatmaps on Ego-Exo4D [ 37]

    Attention heatmaps of our view selector In Fig. 5, we provide examples of our model’s attention heatmaps on Ego-Exo4D [ 37]. Our model tends to focus on the salient objects for an activity, even if they are dynamic, indicating its strong activity understanding ability. 7.1. An...

  120. [130]

    Our model significantly outperforms Body-Area, the best baseline

    3-fold evaluation on Ego-Exo4D In Table 9, we report the results from 3-fold evaluation with Ego- Exo4D [37]. Our model significantly outperforms Body-Area, the best baseline. This shows that our model’s improvement over the baselines sustains across multiple test datasets

  121. [131]

    3.2 in main)

    Pseudo-labeling cost We use 8 NVIDIA V100 GPUs for training and performing in- ference with the captioners in our pseudo-labeler (Sec. 3.2 in main). When pseudo-labeling Ego-Exo4D [37], it takes ∼2.5 days with VideoLlama captioners, and 3 hours with VideoChat2. For LEMMA [56],...

  122. [132]

    contextual variable length clip pairing strategy

    Model performance vs. distribution of con- cepts in ground-truth train narrations Fig. 8 plots our test gains over Body-area [57], the strongest base- line, versus the frequency (most to least) of occurrence of different concepts in the ground-truth train narrations. The lack ...

  123. [2014]

    Springer International Publishing. 2

  124. [2015]

    Association for Computing Machinery. 2

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.