REVIEW 3 major objections 6 minor 85 references
NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding
T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read NoteIt turns instructional videos into structured, interactive notes that preserve the original hierarchy and multimodal cues.
desk verdict A well-built HCI system whose core DAG accuracy claim is never measured; deserves review, with direct evaluation of parallel/sequential edges required. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the hierarchical DAG of the video: nodes are chapters or steps, directed edges encode sequence, and multiple successors from the same predecessor encode parallel or alternative relations. The DAG is produced by a Chain-of-Thought prompt that first identifies chapter and step elements from GPT-4o-generated differential captions and Whisper transcripts, then classifies relations and assembles the graph. Around this DAG, the pipeline layers redundant-free keyframe selection (CLIP semantics plus DINO visual distinctiveness), a GPT-4o agentic workflow for static visual key frames, a scene-based detector for dynamic camera changes, BLIP-2 thumbnail retrieval, and an HTML/CSS/
What would settle it
Use a set of instructional videos whose steps have known parallel/alternative orders (e.g., recipes where ingredients can be swapped or repairs with optional sub-tasks), run NoteIt, and compare the generated DAG's horizontal edges to human annotations of order flexibility; if the horizontal relations match poorly while chapter segmentation remains high, the central structural fidelity claim fails.
Extended reading notes
Core claim
NoteIt's central claim is that a structured representation—a directed acyclic graph of chapters and steps, with sequential edges for vertically ordered content and multiple successors for parallel or alternative content—can be extracted from an instructional video by prompting GPT-4o with differential frame captions and speech transcripts. When this representation is combined with frames that carry visual key information (text overlays, graphic and diagram annotations, special marks, and camera-perspective changes), and rendered as a note scheme with step summaries, thumbnails, and GIFs, the resulting notes maintain structural consistency with the source video and comprehensively convey both
Load-bearing premise
The load-bearing premise is that GPT-4o correctly classifies every chapter and step relation as sequential or parallel/alternative from transcripts and captions—a step that is never verified against ground truth, so the horizontal structure could be wrong in ways the evaluation does not detect.
Editorial extensions
If this is right
- Users can review a how-to video as structured notes, skipping or reordering sections without rewatching.
- The DAG visualization makes parallel or alternative steps visible, helping users understand task flexibility for recipes, repairs, and fitness routines.
- The zero-shot pipeline generalizes across video categories without handcrafted rules, so new instructional content can be processed on demand.
- Customization (modality, detail, interaction mode) lets the same note serve novices and experts, and printable or interactive contexts.
- The note scheme could be exported to other tools or formats beyond the built-in UI.
Reading between the lines
- Because the paper only measures chapter-level segmentation, not the accuracy of the horizontal (parallel/alternative) relations, the strongest untested part is whether the DAG's branching matches user perception; a direct test would compare the generated DAG to human annotations of step order flexibility.
- The pipeline's reliance on GPT-4o for structure classification suggests an inexpensive improvement: a verification step that asks the model to justify each parallel edge with evidence from the transcript could catch misclassifications the current prompt does not catch.
- If the fidelity claim holds beyond the selected 32 videos, the same scheme could turn other procedural media—PDF manuals, wikiHow-style articles—into interactive notes by swapping the parsing layer while keeping the DAG note scheme.
- The paper's own limitation notes suggest that multi-window or slide-like video layouts break the keyframe similarity assumptions; that is a concrete boundary condition to test before wide deployment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. NoteIt is a web-based system that automatically converts instructional videos into interactive, customizable notes. The pipeline consists of video parsing (CLIP/DINO keyframe filtering plus Whisper transcription), hierarchical structure extraction (GPT-4o differential captions, chapter/step clustering, and construction of a DAG with sequential vs. parallel/alternative edges), visual key-information extraction (an agentic GPT-4o workflow for static overlays, plus PySceneDetect and perceptual metrics for camera perspective changes), note creation, and an interactive user interface. The technical evaluation uses 32 instructional videos and reports strong static keyframe extraction (F1 = 91.63%), weaker dynamic keyframe extraction (F1 = 70.86%), and an average chapter-level MRA of 75.31%. A user study (N = 36, counterbalanced within-subject, compared with NoteGPT) finds significantly higher ratings for consistency, informativeness, adaptability, and satisfaction, with a mean SUS of 78.1.
Significance. If the results hold, NoteIt is a useful contribution to mixed-media tutorials and automatic note generation: it provides a concrete design space, a complete MLLM-based pipeline, and a fully implemented interface with three customization axes. The user study is well-formed: it uses G*Power for sample-size estimation, a counterbalanced within-subject design, and appropriate Wilcoxon signed-rank tests. The static keyframe metrics are strong, the chapter-level MRA is a standard segmentation measure, and the system is demonstrated on a diverse set of categories. The main caveat is that the paper's most novel representation—the horizontal/parallel DAG of chapter and step relations—is never directly evaluated; the significance of the claimed contribution therefore depends on an added evaluation of that representation.
major comments (3)
- [Section 4.2; Section 5.2; Section 5.3] The core novel representation—the DAG with sequential vs. parallel/alternative edges—is never directly evaluated. Section 4.2's final step prompts GPT-4o to classify every chapter and step relation and assemble a DAG, directly implementing design goal D1 and the Section 2.2 claim that NoteIt 'faithfully captures chapter- and step-level hierarchies and displays their relationships.' However, Section 5.2 states that evaluation 'focused primarily on chapter-level segmentation,' Section 5.3 reports only MRA and keyframe metrics, and step-level relations are not measured at all. The user-study item Q3 is a subjective rating, not a validation of edge correctness. If GPT-4o mislabels parallel/sequential edges, the note structure misrepresents the video workflow and D1 fails. I request either a direct evaluation of the DAG edges against independently annotated ground truth (e.g., edge-level prec
- [Section 5.3; Table 1] The technical evaluation has no baseline comparison. The paper motivates NoteIt by arguing that direct MLLM summarization produces flat, structure-less notes (Section 2.1, Section 4), and the user study compares against NoteGPT, but the objective metrics in Table 1 and the MRA result are reported without any comparison condition. A chapter MRA of 75.31% and a dynamic keyframe recall of 67.94% are hard to interpret in isolation. I recommend adding at least one baseline to the technical evaluation—e.g., direct GPT-4o summarization with the same transcripts for chapter segmentation, or a standard scene-detection baseline for dynamic keyframes—so that the measured performance can be attributed to the proposed pipeline rather than to the underlying MLLM.
- [Section 5.1] The evaluation videos are selected from the original 80 using criteria that require 'clearly shows the hierarchical structure, with both vertical and horizontal structures' and 'at least two of the visual key information' types. The technical evaluation therefore measures performance on a favorable subset, while Section 5.1 claims the selection 'maintains the same level of diversity' without giving quantitative support. This weakens the generalizability claim (D4). I suggest either reporting results on all 80 videos or a random sample, or providing per-category selection statistics and an analysis of why the excluded videos are not representative of common instructional-video conditions.
minor comments (6)
- [Section 4.2] Typo: 'GTP-4o' should be 'GPT-4o'. Also, the DAG definition is formally imprecise: 'for every sequence of edge ... it must hold that v1 != vk' does not exclude all cycles (e.g., a cycle v2->v3->v2 does not return to v1), and 'parallel or alternative relation as multiple successors from the same predecessors' conflates two distinct notions. Please use standard DAG terminology and separate parallel from alternative edges.
- [Section 4.3.1] The text states that the agent's prompts were 'further refined through human knowledge.' Please report how many refinement iterations were performed, whether the examples were drawn from the same 32 evaluation videos, and whether the prompt tuning was completed before the technical evaluation. This is needed to assess potential experimenter bias.
- [Section 5.2] The ground-truth annotations by 10 raters are described, but no inter-rater reliability statistic (e.g., Cohen's kappa or Fleiss' kappa) is reported. Since chapter boundaries and visual keyframes are subjective, please provide agreement measures to support the reliability of the ground truth.
- [Section 3.2] The text contains a duplicated sentence: 'In contrast, other notes adopt a concise style, distilling steps to essential actions, such as: ...' appears twice in the Content verbosity paragraph. Please remove the duplicate.
- [Section 4.5; Section 4.6; Section 5.3] Minor typos: 'botton' should be 'button' in Section 4.5; 'pre-trainied' should be 'pre-trained' in Section 4.6; 'infomation' should be 'information' in Section 5.3. Also, Section 5.1 has 'The purposefully chosen maintain the same level of diversity'—missing a word.
- [General] No code or dataset is released; the project website is a demo but not a reproducibility artifact. I recommend adding a reproducibility statement (e.g., GitHub repository, annotation files, prompts) or at least a clear description of what is available.
Circularity Check
No significant circularity: the NoteIt pipeline is not constructed from its evaluation targets, and no load-bearing claim reduces to a fit or a self-citation.
full rationale
The central derivation chain is: analyze 80 instructional videos to derive a design space (Section 3), translate that into design goals D1–D4, build a pipeline around GPT-4o and pretrained encoders (Section 4), then evaluate the pipeline on 32 of the same videos with human-annotated ground truth (Section 5). This chain contains no step where an output is defined in terms of the quantity it is said to predict. The hierarchical structure extraction (Section 4.2) prompts GPT-4o to cluster chapters, extract steps, and emit a DAG; no parameter of this module is fitted to the ground-truth chapter boundaries or to any ground-truth DAG edges. The visual key information extraction (Section 4.3) is also prompt-based and zero-shot, not trained on the raters' annotations. The evaluation metrics (MRA, precision, recall, F1) compare system outputs to independently annotated labels, so the measurements are genuine. The fact that the evaluation categories (text overlays, graphic/diagram annotations, special marks, camera perspective manipulation) were defined in the authors' own design-space analysis is a methodological limitation—the evaluation is in-sample relative to the design analysis—but it is not circular: the system could have failed to detect these categories even though they were named in the prompts. The paper's self-citations (e.g., [24], [25], [26]) are not load-bearing; they support minor methodological references, not the central claims. The noted gap that step-level and horizontal DAG relations are never directly evaluated is a serious validity concern, not a circularity: the paper simply does not measure that part of its claim. Therefore no circular step can be exhibited with the required specificity, and the appropriate score is 0.
Assumptions & free parameters
free parameters (3)
- Keyframe filtering similarity thresholds (CLIP, DINO) =
not disclosed
- Dynamic keyframe detection criteria (SSIM, color histogram, ORB, CLIP, MiDaS depth) =
not disclosed
- GPT-4o prompt templates (chapter/step clustering, CoT DAG construction, keyframe agents, summarization) =
not disclosed
assumptions (4)
- domain assumption GPT-4o follows the designed prompts accurately enough that captions, chapter/step clusters, DAG relations, and summaries contain no task-breaking errors.
- domain assumption The purposive sample of 80 YouTube videos across 8 categories, 10 Notion templates, and 30 tutorials is representative enough to yield generalizable design goals.
- domain assumption In-house rater annotations (2 annotators plus 2 checkers per video) are valid ground truth.
- domain assumption Chapter-level structural segmentation (MRA) is a sufficient proxy for the structural fidelity of the notes.
Cite this review
Pith. "Pith review of NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding." pith.science (2026). https://pith.science/paper/E52JLEVW
@misc{pith2026250814395,
author = {Pith},
title = {Pith review of: NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding},
year = {2026},
howpublished = {\url{https://pith.science/paper/E52JLEVW}},
note = {Machine review of arXiv:2508.14395}
}
read the original abstract
Users often take notes for instructional videos to access key knowledge later without revisiting long videos. Automated note generation tools enable users to obtain informative notes efficiently. However, notes generated by existing research or off-the-shelf tools fail to preserve the information conveyed in the original videos comprehensively, nor can they satisfy users' expectations for diverse presentation formats and interactive features when using notes digitally. In this work, we present NoteIt, a system, which automatically converts instructional videos to interactable notes using a novel pipeline that faithfully extracts hierarchical structure and multimodal key information from videos. With NoteIt's interface, users can interact with the system to further customize the content and presentation formats of the notes according to their preferences. We conducted both a technical evaluation and a comparison user study (N=36). The solid performance in objective metrics and the positive user feedback demonstrated the effectiveness of the pipeline and the overall usability of NoteIt. Project website: https://zhaorunning.github.io/NoteIt/
Reference graph
Works this paper leans on
-
[1]
Aaron Bangor, Philip T Kortum, and James T Miller. 2008. An empirical evaluation of the system usability scale. Intl. Journal of Human–Computer Interaction 24, 6 (2008), 574–594
2008
-
[2]
Peter Brandl, Christoph Richter, and Michael Haller. 2010. NiCEBook: supporting natural note taking. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Atlanta, Georgia, USA) (CHI ’10). Association for Computing Machinery, New York, NY, USA, 599–608. doi:10.1145/1753326.1753417
arXiv 2010
-
[3]
ByteDance. 2024. Seed-TTS: A Family of High-Quality Versatile Speech Genera- tion Models. arXiv:2406.02430 [eess.AS] https://arxiv.org/abs/2406.02430
arXiv 2024
-
[4]
Yining Cao, Hariharan Subramonyam, and Eytan Adar. 2022. VideoSticker: A Tool for Active Viewing and Visual Note-taking from Videos. In Proceedings of the 27th International Conference on Intelligent User Interfaces (Helsinki, Finland) (IUI ’22). Association for Computing Machinery, New York, NY, USA, 672–690. doi:10.1145/3490099.3511132
arXiv 2022
-
[5]
Brandon Castellano. [n. d.]. PySceneDetect: Video Cut Detection and Analysis Tool. https://github.com/Breakthrough/PySceneDetect
-
[6]
Guillain, Hyeungshik Jung, Vivian M
Minsuk Chang, Leonore V. Guillain, Hyeungshik Jung, Vivian M. Hare, Juho Kim, and Maneesh Agrawala. 2018. RecipeScape: An Interactive Tool for Analyzing Cooking Instructions at Scale. InProceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–12....
arXiv 2018
-
[7]
Jun Chen, Deyao Zhu, Kilichbek Haydarov, Xiang Li, and Mohamed Elhoseiny
-
[8]
Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Zhenyu Tang, Li Yuan, et al. 2024. Sharegpt4video: Im- proving video understanding and generation with better captions. Advances in Neural Information Processing Systems 37 (2024), 19472–19495
2024
Show all 85 references
-
[9]
Yuexi Chen, Vlad I Morariu, Anh Truong, and Zhicheng Liu. 2024. TutoAI: a cross- domain framework for AI-assisted mixed-media tutorial creation on physical tasks. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Ass...
2024
-
[10]
Yi-Ting Chen, Chi-Hsuan Hsu, Chih-Han Chung, Yu-Shuen Wang, and Sabarish V. Babu. 2019. iVRNote: Design, Creation and Evaluation of an Interactive Note- Taking Interface for Study and Reflection in VR Learning Environments. In 2019 IEEE Conference on Virtual Reality and 3D Use...
2019
-
[11]
Peggy Chi, Nathan Frey, Katrina Panovich, and Irfan Essa. 2021. Automatic Instructional Video Creation from a Markdown-Formatted Tutorial. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21). Association for Computing Mach...
2021
-
[12]
Pei-Yu Chi, Sally Ahn, Amanda Ren, Mira Dontcheva, Wilmot Li, and Björn Hartmann. 2012. MixT: automatic generation of step-by-step mixed media tutorials. In Proceedings of the 25th Annual ACM Symposium on User Interface Software and Technology (Cambridge, Massachusetts, USA) (...
2012
-
[13]
Pei-Yu Chi, Joyce Liu, Jason Linder, Mira Dontcheva, Wilmot Li, and Bjoern Hartmann. 2013. Democut: generating concise instructional videos for physi- cal demonstrations. In Proceedings of the 26th annual ACM symposium on User interface software and technology . 141–150
2013
-
[14]
Guan-Jun Ding, TK Philip Hwang, and Pin-Chieh Kuo. 2020. Progressive dis- closure options for improving choice overload on home screen. In Advances in Usability, User Experience, Wearable and Assistive Technology: Proceedings of the AHFE 2020 Virtual Conferences on Usability a...
2020
-
[15]
Franz Faul, Edgar Erdfelder, Axel Buchner, and Albert-Georg Lang. 2009. Sta- tistical power analyses using G* Power 3.1: Tests for correlation and regression analyses. Behavior research methods 41, 4 (2009), 1149–1160
2009
-
[16]
C Ailie Fraser, Joy O Kim, Hijung Valentina Shin, Joel Brandt, and Mira Dontcheva
-
[17]
Google. 2024. NotebookLM. https://notebooklm.google/. Accessed: 2025-04-07
2024
-
[18]
Mingfei Han, Linjie Yang, Xiaojun Chang, and Heng Wang. 2023. Shot2story20k: A new benchmark for comprehensive understanding of multi-shot videos. arXiv preprint arXiv:2312.10300 (2023)
2023 arXiv
-
[19]
Wafa’A Hazaymeh and Moath Khalaf Alomery. 2022. The Effectiveness of Vi- sual Mind Mapping Strategy for Improving English Language Learners’ Critical Thinking Skills and Reading Ability. European Journal of Educational Research 11, 1 (2022), 141–150
2022
-
[20]
Markus H Hefter. 2024. Note-taking fosters distance video learning: smartphones as risk and intellectual values as protective factors. Scientific Reports 14, 1 (2024), 16962
2024
-
[21]
Aidan Hogan, Eva Blomqvist, Michael Cochez, Claudia d’Amato, Gerard De Melo, Claudio Gutierrez, Sabrina Kirrane, José Emilio Labra Gayo, Roberto Navigli, Sebastian Neumaier, et al. 2021. Knowledge graphs. ACM Computing Surveys (Csur) 54, 4 (2021), 1–37
2021
-
[22]
Bin Huang, Xin Wang, Hong Chen, Zihan Song, and Wenwu Zhu. 2024. VTimeLLM: Empower LLM to Grasp Video Moments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 14271– 14280
2024
-
[23]
Qirui Huang, Min Lu, Joel Lanir, Dani Lischinski, Daniel Cohen-Or, and Hui Huang. 2024. Graphimind: Llm-centric interface for information graphics design. arXiv preprint arXiv:2401.13245 (2024)
2024 arXiv
-
[24]
Zhihan Jiang, Handi Chen, Rui Zhou, Jing Deng, Xinchen Zhang, Running Zhao, Cong Xie, Yifang Wang, and Edith CH Ngai. 2023. Healthprism: a visual ana- lytics system for exploring children’s physical and mental health profiles with multimodal data. IEEE Transactions on Visualiz...
2023
-
[25]
Zhihan Jiang, Xin He, Chenhui Lu, Binbin Zhou, Xiaoliang Fan, Cheng Wang, Xiaojuan Ma, Edith CH Ngai, and Longbiao Chen. 2022. Understanding drivers’ visual and comprehension loads in traffic violation hotspots leveraging crowd- based driving simulation. IEEE transactions on i...
2022
-
[26]
Zhihan Jiang, Running Zhao, Lin Lin, Yue Yu, Handi Chen, Xinchen Zhang, Xuhai Xu, Yifang Wang, Xiaojuan Ma, and Edith CH Ngai. 2025. DietGlance: Dietary Monitoring and Personalized Analysis at a Glance with Knowledge-Empowered AI Assistant. arXiv preprint arXiv:2502.01317 (2025)
2025
-
[27]
Matthew Kam, Jingtao Wang, Alastair Iles, Eric Tse, Jane Chiu, Daniel Glaser, Orna Tarshish, and John Canny. 2005. Livenotes: a system for cooperative and augmented note-taking in lectures. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Portland...
2005
-
[28]
Kenneth A Kiewra, Nelson F DuBois, David Christian, Anne McShane, Michelle Meyerhoffer, and David Roskelley. 1991. Note-taking functions and techniques. Journal of educational psychology 83, 2 (1991), 240
1991
-
[29]
Guo, Robert C
Juho Kim, Phu Tran Nguyen, Sarah Weir, Philip J. Guo, Robert C. Miller, and Krzysztof Z. Gajos. 2014. Crowdsourcing step-by-step information extraction to enhance existing how-to videos. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, On...
2014
-
[30]
Ross A Knepper, Dishaan Ahuja, Geoffrey Lalonde, and Daniela Rus. 2014. Dis- tributed assembly with and/or graphs. In Workshop on AI Robotics at the Int. Conf. on Intelligent Robots and Systems (IROS)
2014
-
[31]
Jon Kolko. 2010. Exposing the Magic of Design: A Practitioner’s Guide to the Methods and Theory of Synthesis . Oxford University Press. doi:10.1093/acprof: oso/9780199744336.001.0001
2010
-
[33]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning . PMLR, 19730–19742
2023
-
[34]
Chuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng, Zhi Yu, and Cong Yao. 2024. LayoutLLM: Layout Instruction Tuning with Large Language Models for Docu- ment Understanding. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 15630–15640. doi:10.1109/CVP...
2024
-
[35]
Xiaojun Meng, Shengdong Zhao, and Darren Edge. 2016. HyNote: Integrated Concept Mapping and Notetaking. InProceedings of the International Working Con- ference on Advanced Visual Interfaces (Bari, Italy) (A VI ’16). Association for Com- puting Machinery, New York, NY, USA, 236...
2016
-
[36]
Megha Nawhal, Jacqueline B Lang, Greg Mori, and Parmit K Chilana. 2019. VideoWhiz: Non-Linear Interactive Overviews for Recipe Videos.. In Graphics Interface. 15–1
2019
-
[37]
Cuong Nguyen and Feng Liu. 2016. Gaze-based Notetaking for Learning from Lecture Videos. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (San Jose, California, USA) (CHI ’16) . Association for Computing Machinery, New York, NY, USA, 2093–2097. d...
2016 doi
-
[38]
NoteGPT. 2024. NoteGPT. https://notegpt.io/. Accessed: 2025-04-07
2024
-
[39]
OpenAI. 2024. GPT-4o System Card. arXiv:2410.21276 [cs.CL] https://arxiv.org/ abs/2410.21276
2024 arXiv
-
[40]
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po- Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabb...
2024
-
[41]
Liang-Ming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu, Chong-Wah Ngo, Min-Yen Kan, Yugang Jiang, and Tat-Seng Chua. 2020. Multi-modal Cooking Workflow Construction for Food Recipes. In Proceedings of the 28th ACM Interna- tional Conference on Multimedia(Seattle, WA, USA)(MM...
2020
-
[42]
Amy Pavel, Colorado Reed, Björn Hartmann, and Maneesh Agrawala. 2014. Video digests: a browsable, skimmable format for informational lecture videos.. InUIST, Vol. 10. Citeseer, 2642918–2647400
2014
-
[43]
Yingzhe Peng, Xiaoting Qin, Zhiyang Zhang, Jue Zhang, Qingwei Lin, Xu Yang, Dongmei Zhang, Saravan Rajmohan, and Qi Zhang. 2025. Navigating the Un- known: A Chat-Based Collaborative Interface for Personalized Exploratory Tasks. In Proceedings of the 30th International Conferen...
2025
-
[44]
Yi-Hao Peng, Peggy Chi, Anjuli Kannan, Meredith Ringel Morris, and Irfan Essa
-
[45]
Leping Qiu, Erin Seongyoon Kim, Sangho Suh, Ludwig Sidenmark, and Tovi Grossman. 2025. MaRginalia: Enabling In-person Lecture Capturing and Note- taking Through Mixed Reality. arXiv preprint arXiv:2501.16010 (2025)
2025 arXiv
-
[46]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[47]
In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23)
Slide Gestalt: Automatic Structure Extraction in Slide Decks for Non-Visual Access. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Article 829, 14 pages. doi:...
2023
-
[48]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning . PMLR, 28492–28518
2023
-
[49]
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun
-
[50]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[51]
Dalai S Ribeiro, Alysson Gomes de Sousa, Rodrigo B de Almeida, Pedro Henrique Thompson Furtado, Hélio Côrtes Vieira Lopes, and Simone Diniz Junqueira Barbosa. 2020. Exploring Ontology-Based Information Through the Progressive Disclosure of Visual Answers to Related Queries. In...
2020
-
[52]
Bernard Rosner, Robert J Glynn, and Mei-Ling T Lee. 2006. The Wilcoxon signed rank test for paired comparisons of clustered data. Biometrics 62, 1 (2006), 185– 192
2006
-
[53]
IEEE transactions on pattern analysis and machine intelligence 44, 3 (2020), 1623–1637
Towards robust monocular depth estimation: Mixing datasets for zero- shot cross-dataset transfer. IEEE transactions on pattern analysis and machine intelligence 44, 3 (2020), 1623–1637
2020
-
[54]
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al. 2024. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714 (2024)
2024 arXiv
-
[55]
Alessandra Semeraro and Laia Turmo Vidal. 2022. Visualizing Instructions for Physical Training: Exploring Visual Cues to Support Movement Learning from Instructional Videos. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) ...
2022
-
[56]
Xuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin, Bowen He, Xiaodong Han, Aixuan Li, Yuchao Dai, Lingpeng Kong, Meng Wang, et al. 2023. Fine-grained audible video description. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10585–10596
2023
-
[57]
Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. 2011. ORB: An efficient alternative to SIFT or SURF. In 2011 International conference on computer vision. Ieee, 2564–2571
2011
-
[58]
Ahmad Seifi and Amir Moshayeri. 2024. The Influence of Color Schemes and Aesthetics on User Satisfaction in Web Design: An Empirical Study.International Journal of Advanced Human Computer Interaction 2, 2 (2024), 33–43
2024
-
[59]
Bennett, and Jaime Teevan
Amanda Swearngin, Shamsi Iqbal, Victor Poznanski, Mark Encarnación, Paul N. Bennett, and Jaime Teevan. 2021. Scraps: Enabling Mobile Capture, Contex- tualization, and Use of Document Resources. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yo...
2021
-
[60]
Anh Truong, Peggy Chi, David Salesin, Irfan Essa, and Maneesh Agrawala. 2021. Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup Videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21)....
2021 doi
-
[61]
Hijung Valentina Shin, Floraine Berthouzoz, Wilmot Li, and Frédo Durand. 2015. Visual transcripts: lecture notes from blackboard-style lecture videos.ACM Trans. Graph. 34, 6, Article 240 (Nov. 2015), 10 pages. doi:10.1145/2816795.2818123
2015
-
[62]
Hariharan Subramonyam, Colleen Seifert, Priti Shah, and Eytan Adar. 2020. texSketch: Active Diagramming through Pen-and-Ink Annotations. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Mach...
2020
-
[63]
Bryan Wang, Meng Yu Yang, and Tovi Grossman. 2021. Soloist: Generating mixed-initiative tutorials from existing guitar instructional videos through au- dio processing. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–14
2021
-
[64]
Cheng-Yao Wang, Wei-Chen Chu, Hou-Ren Chen, Chun-Yen Hsu, and Mike Y. Chen. 2014. EverTutor: automatically creating interactive guided tutorials on smartphones by user demonstration. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, Ontari...
2014
-
[65]
Alexandre N Tuch, Sandra P Roth, Kasper Hornbæk, Klaus Opwis, and Javier A Bargas-Avila. 2012. Is beautiful really usable? Toward understanding the relation between usability, aesthetics, and affect in HCI. Computers in human behavior 28, 5 (2012), 1596–1607
2012
-
[66]
Sylvaine Tuncer, Barry Brown, and Oskar Lindwall. 2020. On Pause: How Online Instructional Videos are Used to Achieve Practical Tasks. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machi...
2020
-
[67]
Xiaohan Wang, Yuhui Zhang, Orr Zohar, and Serena Yeung-Levy. 2024. VideoA- gent: Long-form Video Understanding with Large Language Model as Agent. arXiv:2403.10517 [cs.CV] https://arxiv.org/abs/2403.10517
2024 arXiv
-
[68]
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing 13, 4 (2004), 600–612. UIST ’25, September 28-October 1, 2025, Busan, Republic of Korea R. Zhao,...
2004
-
[69]
Liang Wang, Gaofeng Che, Jiantuan Hu, and Lin Chen. 2024. Online review helpfulness and information overload: the roles of text, image, and video elements. Journal of Theoretical and Applied Electronic Commerce Research19, 2 (2024), 1243– 1266
2024
-
[70]
Mingqiu Wang, Izhak Shafran, Hagen Soltau, Wei Han, Yuan Cao, Dian Yu, and Laurent El Shafey. 2024. Retrieval Augmented End-to-End Spoken Dialog Models. arXiv:2402.01828 [cs.CL] https://arxiv.org/abs/2402.01828
2024 arXiv
-
[71]
Chengpei Xu, Wenjing Jia, Ruomei Wang, Xiangjian He, Baoquan Zhao, and Yuanfang Zhang. 2023. Semantic Navigation of PowerPoint-Based Lecture Video for AutoNote Generation. IEEE Transactions on Learning Technologies 16, 1 (2023), 1–17. doi:10.1109/TLT.2022.3216535
2023
-
[72]
Chengpei Xu, Ruomei Wang, Shujin Lin, Xiaonan Luo, Baoquan Zhao, Lijie Shao, and Mengqiu Hu. 2019. Lecture2Note: Automatic Generation of Lecture Notes from Slide-Based Educational Videos. In 2019 IEEE International Conference on Multimedia and Expo (ICME) . 898–903. doi:10.110...
2019
-
[73]
Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei, Xudong Lin, Shuohang Wang, Ziyi Yang, Chenguang Zhu, Derek Hoiem, et al. 2022. Language models with image descriptors are strong few-shot video-language learners. Advances in Neural Information Processing Systems ...
2022
-
[74]
Gajos, and Robert C
Sarah Weir, Juho Kim, Krzysztof Z. Gajos, and Robert C. Miller. 2015. Learn- ersourcing Subgoal Labels for How-to Videos. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing (Van- couver, BC, Canada) (CSCW ’15). Association for C...
2015
-
[75]
Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie
-
[76]
Saelyne Yang, Sangkyung Kwak, Tae Soo Kim, and Juho Kim. 2022. Improving Video Interfaces by Presenting Informational Units of Videos. CHI’22 Extended Abstracts. Association for Computing Machinery (2022)
2022
-
[77]
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont- Tuset, Ivan Laptev, Josef Sivic, and Cordelia Schmid. 2023. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. In Proceedings of the IEEE/CVF Conference on Computer ...
2023
-
[78]
Dongjie Yang, Suyuan Huang, Chengqiang Lu, Xiaodong Han, Haoxin Zhang, Yan Gao, Yao Hu, and Hai Zhao. 2024. Vript: A video is worth thousands of words. Advances in Neural Information Processing Systems 37 (2024), 57240–57261
2024
-
[79]
Andy Zeng, Maria Attarian, brian ichter, Krzysztof Marcin Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael S Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, and Pete Florence. 2023. Socratic Models: Composing Zero-Shot Multimodal Reasonin...
2023
-
[80]
Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-llama: An instruction- tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858 (2023)
2023 arXiv
-
[81]
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large lan- guage models. arXiv preprint arXiv:2304.10592 (2023)
2023 arXiv
-
[82]
Saelyne Yang, Sangkyung Kwak, Juhoon Lee, and Juho Kim. 2023. Beyond Instructions: A Taxonomy of Information Types in How-to Videos. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery...
2023
-
[83]
Saelyne Yang, Anh Truong, Juho Kim, and Dingzeyu Li. 2025. VideoMix: Aggregat- ing How-To Videos for Task-Oriented Learning. InProceedings of the 30th Interna- tional Conference on Intelligent User Interfaces (IUI ’25). Association for Computing Machinery, New York, NY, USA, 1...
2025
-
[2020]
In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems
Temporal segmentation of creative live streams. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–12
2020
-
[2023]
Video ChatCaptioner: Towards enriched spatiotemporal descriptions.arXiv preprint arXiv:2304.04227 (2023)
2023 arXiv
-
[2024]
arXiv preprint arXiv:2412.14171 (2024)
Thinking in space: How multimodal large language models see, remember, and recall spaces. arXiv preprint arXiv:2412.14171 (2024)
2024 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.