Pith. sign in

REVIEW 4 major objections 6 minor 137 references

Vid2Coach: Transforming How-To Videos into Task Assistants

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Vid2Coach turns how-to videos into hands-free cooking assistants for blind and low vision cooks, cutting errors by 58.5% in a home-kitchen study.

desk verdict A strong, well-grounded HCI systems paper whose core user-study result is credible; the main soft spot is that the proactive feedback mechanism is validated on only six actions, so the distinctive 'coach' claim is thinner than the full-system result. read the letter →

arxiv 2506.00717 v2 pith:OEGY4IES submitted 2025-05-31 cs.HC cs.CV

classification cs.HCcs.CV
keywords how-tovideosblindandlowvisionwearabletaskassistantsmartglassesvision-languagemodelretrieval-augmentedgenerationproceduralfeedbackaccessibility
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a how-to video itself contains enough latent task knowledge—demonstration details, tool and ingredient states, and visual signs of completion—for an AI system to transform it into an accessible, hands-free task assistant for blind and low vision (BLV) people. In a within-subjects kitchen study with eight BLV participants, the system reduced procedural and technique errors by 58.5% and let five of eight participants finish the recipe, versus one of eight with their usual workflow. The paper argues that the rich audio-visual content that benefits sighted learners can be repurposed into non-visual guidance, strengthening rather than replacing the users' own tactile, auditory, and olfactory skills.

What carries the argument

The load-bearing object is the set of per-action completion criteria generated from the source video—visual descriptions of what 'done' looks like (for example, 'the butter is a deep golden brown'), together with in-progress and mistake criteria. Actions are first typed as punctual, iterative, or durative so that feedback timing matches the action's temporal structure: punctual actions only receive completion confirmation, iterative actions get counting, and durative actions get progressive updates. A dual-model loop—a fast streaming vision-language model for immediate responsiveness and a slower batch model running every five seconds to classify status against the criteria—monitors the user's smart-glasses stream and decides when to pause, update, or prompt for confirmation.

What would settle it

Run a controlled transfer test: take one recipe video, generate completion criteria, and have several BLV cooks prepare the same dish in different kitchens while egocentric video is recorded, then measure per-frame agreement between the system's status labels and human annotations of 'in progress / complete / mistake.' The claim would be weakened if accuracy for punctual and iterative actions stays at or below the 0.53–0.60 narrow-field-of-view figures reported, because those are exactly the actions that proactive feedback depends on.

Watch

Extended reading notes

Core claim

Vid2Coach establishes that a how-to video can be converted into an interactive task assistant without additional human annotation. The pipeline transcribes and filters the narration into atomic actions, pairs each action with task-relevant frames, uses a vision-language model to write demonstration details and per-action completion criteria, and supplements steps with retrieval-augmented workarounds drawn from BLV-authored resources. During the task, a fast streaming model gives low-latency descriptions while a slower batch model judges the user's egocentric video against those criteria, pausing during unrelated actions, giving progress updates, and asking for confirmation before advancing. The user study then shows measurable error reduction and task-completion gains in real home kitchens, supporting the claim that proactive, video-grounded feedback is what drives the benefit.

Load-bearing premise

The visual completion criteria written from the source video's camera angles, lighting, and tools still apply to the user's own egocentric view in their own kitchen, where the camera is narrower, hands occlude objects, and lighting and equipment differ.

Editorial extensions

If this is right

  • BLV users could follow mainstream how-to videos without a sighted helper on call, because the video itself supplies the visual details and completion checks that are normally missing from narration.
  • The same pipeline should transfer to other hands-on domains where step outcomes have visible states, such as assembly, decoration, or crafts, as the paper's exploratory extension begins to show.
  • Error patterns would shift from missed steps and wrong measurements to finer-grained issues like localized doneness, pointing to where the next generation of assistive feedback should focus.
  • Dependence on remote visual interpreter services could drop for the most frequent 'is this done?' questions, since proactive feedback covers them in real time.
  • Because the system grounds its feedback in the video's own criteria, it can explain why a step looks complete, giving users confidence to proceed independently.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would be to let the system learn completion criteria from the user's own prior sessions, so that criteria adapt to their tools, lighting, and camera framing instead of relying solely on the source video.
  • The action-type taxonomy suggests a simple robustness heuristic: refuse proactive completion feedback for punctual actions unless the camera view can be stabilized, since narrow field-of-view made those actions the least reliable in the paper's technical evaluation.
  • The reported error reduction may depend on the particular vision-language models and prompt design; an ablation that swaps the batch model for an open-source model of comparable cost would reveal how much of the gain comes from the criteria abstraction versus raw model strength.
  • The egocentric, task-local monitoring design implies a privacy-friendly deployment path: on-device inference could make the assistant viable in kitchens where streaming video is undesirable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. Vid2Coach proposes a system that converts narrated how-to videos into step-wise verbal instructions augmented with demonstration details, BLV-specific tips retrieved via RAG, and real-time progress feedback from a smart-glasses camera. A formative observational study with three VRT-BLV pairs motivates six design goals. The system is evaluated in three parts: instruction-generation quality against state-of-the-art VLMs, progress-monitoring accuracy on six action videos, and a within-subjects study with eight BLV participants cooking in their own home kitchens, reporting 58.5% fewer errors, higher task completion, and lower workload. An exploratory extension with one participant covers flower arrangement and gingerbread-house assembly. The full-system result is promising, but the paper does not isolate the contribution of proactive progress feedback from the other components of the system.

Significance. If the full-system result holds, this is a valuable contribution to accessibility research: it provides real-world evidence that a wearable, VLM-driven assistant can reduce errors for BLV people following how-to videos, and the VRT observational study is a useful design resource. The pipeline is fully automatic and was not tuned on the user-study data, and the instruction-generation evaluation is a strength because it measures atomic-fact coverage and hallucination against strong baselines. The authors also plan to release their curated accessibility-resource datasets. However, the paper's central causal claim that proactive progress feedback drives the error reduction is not yet secured: the user study varies the whole system against participants' existing workflows, the technical evaluation of progress monitoring is limited to six action videos with low narrow-field-of-view accuracy for punctual and iterative actions, and the only direct rating of feedback is not statistically significant. These issues are load-bearing for the 'coach' contribution and require either additional evidence or a carefully scoped claim.

major comments (4)
  1. [§7.2 and §5.3] The paper attributes the error reduction to the system's 'real-time guidance,' but the user study compares the full Vid2Coach system to participants' existing workflows; there is no condition that disables or ablates progress feedback while keeping the generated step instructions, demonstration details, and RAG tips. The 58.5% error reduction could therefore be carried entirely by the accessible-instruction and workaround components, which are not the components that make Vid2Coach a 'coach.' This is not a hypothetical concern: §7.4 documents incorrect progress feedback (pancake doneness judged from the top surface, missed bacon pieces, over- and under-doneness), and §7.3 reports that the difference in participants' ratings of feedback was not statistically significant (Z=0.85, p>0.05). The revision should either add an ablation that removes or disables the feedback mechanism, or restrict the Section 7.2 causal claim to the full system and present the feedback mechanism as a design feature whose individual contribution remains untested.
  2. [§5.3 and §6.2] The claim that completion criteria extracted from the how-to video generalize to the user's egocentric camera stream is load-bearing for the proactive-feedback mechanism. The evidence in Table 2, however, rests on only 6 action videos and 90 frames, with no per-recipe or per-participant breakdown. On the narrow Meta-glasses field of view, per-frame accuracy is 0.53 for punctual actions and 0.60 for iterative actions, and the user-study recipes (V11, V12) contain many such actions. The failure cases in §7.4 confirm that these are not edge cases in deployment. The authors should validate the criteria on a substantially larger set of actions, report confusion matrices and confidence intervals, and ideally measure feedback precision and recall on the actual user-study egocentric videos to connect the technical metric to the reported user outcomes.
  3. [§7.1] The primary outcome, error count, is coded without reported inter-rater reliability. Section 7.1 describes the coding categories but does not state how many annotators coded the user-study videos, whether they were blinded to condition, or what agreement they achieved. With N=8, measurement noise in the error counts can materially affect the Wilcoxon results. In addition, because only one of eight participants completed the baseline task versus five of eight with Vid2Coach, raw error counts are not normalized by the number of steps attempted or by cooking time; a participant who stops early has a different exposure to error opportunities than one who finishes. Please report inter-rater reliability, describe the coding procedure in detail, and provide a sensitivity analysis using errors per completed step or per minute.
  4. [§7.3] The only direct rating of the feature that is claimed to drive the main effect—feedback—is not statistically significant, and the reported summary statistics appear to contain a data-reporting issue: the text gives identical means (μ=1.28 vs. μ=1.28) with different standard deviations (σ=4.75 vs. σ=5.25) on a 1–7 scale. If this is a typo, it should be corrected; if the feedback-helpfulness rating was indeed equivalent between conditions, the qualitative quotes in §7.3 do not compensate for the lack of a quantitative signal. The revision should report the correct values and discuss what the nonsignificant feedback rating means for the claim that progress feedback is what reduces errors.
minor comments (6)
  1. [§7.2] The reported 58.5% error reduction does not match the reported means: (11.00 − 4.38) / 11.00 = 60.2%; please clarify how the percentage was computed.
  2. [§6.2] Table 2 reports average per-frame accuracy without per-action counts, standard deviations, or confidence intervals; adding these would make the small-sample comparison more interpretable.
  3. [§8 and Table 4] The extension study text says participant P5 completed the flower-arrangement and gingerbread-house tasks, but Table 4 lists the extension-study participant as 'PX'; please align the label.
  4. [§9.6] The text expands VRT as 'Virtual Reality Therapy,' but the paper consistently uses VRT to mean vision rehabilitation therapist; please correct this expansion.
  5. [§10] The conclusion contains a typo: 'sclable' should be 'scalable.'
  6. [§5.1] The CLIP frame-similarity threshold range (0.27–0.30) and the ±15-second sampling window are described as heuristics without a sensitivity analysis; a sentence noting their robustness would be helpful.

Circularity Check

0 steps flagged · score 0.0 of 10

The central claims rest on external user-study and held-out evaluations, not on fitted or self-referential inputs.

full rationale

The paper's primary outcome is measured in a within-subjects user study with BLV participants in their own kitchens (Section 7), which compares the full Vid2Coach system against the participants' typical workflow. No system parameter is fitted to the error counts or task-completion data from that study, and the reported 58.5% error reduction is an external empirical result rather than a quantity derived from the system's own definitions. The progress-monitoring component is evaluated separately on held-out egocentric action videos (Section 6.2), where criteria extracted from how-to videos are applied to user streams and compared against baseline prompts and CLIP; this is a generalization test, not a fit. The only notable author self-citation appears in Section 5.1, where the targeted prompting strategy follows Huh et al. [43], but that citation is not load-bearing for the central claim: the user study evaluates the whole assistant, and the prompting strategy is an implementation choice rather than the evidence for the reported error reduction. The paper's own limitations (Section 7.4, Table 2) show that progress feedback is imperfect under narrow field-of-view and occlusion, and that the user study does not isolate the feedback component; those are evidence-strength and causal-attribution concerns, not circularity. No step was found where a prediction is equivalent to its input by construction, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

This is an empirical systems paper without mathematical derivations. The main assumptions are about the reliability and generality of AI components (Whisper, Gemini, CLIP, RAG) and the representativeness of the datasets and study population. No new physical or conceptual entities are postulated.

free parameters (3)
  • CLIP frame similarity threshold range = 0.27-0.30 (adaptive per action)
    Used to filter task-relevant frames in Section 5.1. The range is hand-chosen and adjusted based on frame density, not derived from a held-out set.
  • Temporal sampling window = ±15 seconds around action timestamps
    Frames are sampled at 1 fps within a 15-second window around each action. This window size is a design choice affecting which visual demonstrations are captured.
  • RAG top-k = 3
    Retrieves top 3 text chunks to generate tips/workarounds. This k is chosen by the authors and not ablated in the paper.
assumptions (4)
  • domain assumption How-to videos with spoken narration contain sufficient task-relevant information to generate accessible instructions.
    The entire pipeline depends on Whisper transcription of narration. Videos that rely purely on visuals or ambient audio are explicitly excluded in Section 9.1, limiting scope.
  • domain assumption VLM-generated completion criteria from the source video generalize to the user's egocentric view.
    The progress monitor uses abstracted criteria rather than frame matching. Generalization is assumed despite differences in camera, lighting, kitchen, and tools; only small-scale ablation evidence is provided in Section 6.2.
  • domain assumption Curated accessibility resources are representative enough for RAG to surface useful workarounds.
    The dataset contains 100 cooking videos and 100 resources, but its coverage is not measured against the space of cooking steps and no retrieval quality evaluation is reported.
  • domain assumption The baseline, typical workflow (video, transcript, human/AI assistance), is a strong and fair comparator.
    The user study compares Vid2Coach to participants' current practices, which vary widely and include unreliable human agents; this comparison may inflate the apparent benefit.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Vid2Coach: Transforming How-To Videos into Task Assistants." pith.science (2026). https://pith.science/paper/OEGY4IES

@misc{pith2026250600717,
  author       = {Pith},
  title        = {Pith review of: Vid2Coach: Transforming How-To Videos into Task Assistants},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OEGY4IES}},
  note         = {Machine review of arXiv:2506.00717}
}
read the original abstract

People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our observations of visual rehabilitation therapists (VRTs) guiding BLV people to follow how-to videos revealed that VRTs provide both proactive and responsive support including detailed descriptions, non-visual workarounds, and progress feedback. We propose Vid2Coach, a system that transforms how-to videos into wearable camera-based assistants that provide accessible instructions and mixed-initiative feedback. From the video, Vid2Coach generates accessible instructions by augmenting narrated instructions with demonstration details and completion criteria for each step. It then uses retrieval-augmented-generation to extract relevant non-visual workarounds from BLV-specific resources. Vid2Coach then monitors user progress with a camera embedded in commercial smart glasses to provide context-aware instructions, proactive feedback, and answers to user questions. BLV participants (N=8) using Vid2Coach completed cooking tasks with 58.5\% fewer errors than when using their typical workflow and wanted to use Vid2Coach in their daily lives. Vid2Coach demonstrates an opportunity for AI visual assistance that strengthens rather than replaces non-visual expertise.

Figures

Figures reproduced from arXiv: 2506.00717 by the authors.

Figure 1
Figure 1. Vid2Coach is a system that transforms how-to videos into a wearable camera-based task assistant that provides [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. We observed how VRTs deliver real-time remote [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 4
Figure 4. Vid2Coach generates step instructions from a how [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figures from the paper (7 more)
Figure 5
Figure 5. Figure 5: From the how-to video, Vid2Coach generates crite [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Qualitative comparisons of Vid2Coach descriptions with SOTA VLMs on 2 action sequences. These VLM descriptions [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: In the baseline condition, participants used their [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: Step completion (left) and user-initiated interactions (right) visualized across participants, grouped by system condition [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 10
Figure 10. Figure 10: Final dishes from cooking tasks with Vid2Coach [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: In our extension study, P5 explored the use of [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: In the user study, participants used Vid2Coach to receive real-time feedback and ask free-form questions during [PITH_FULL_IMAGE:figures/full_fig_p024_12.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

137 extracted references · 50 canonical work pages

  1. [1]

    https://www.bemyeyes.com/

    Last visited: 2025. https://www.bemyeyes.com/

  2. [2]

    https://aira.io/

    Last visited: 2025. https://aira.io/

  3. [3]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  4. [4]

    Google Cloud Vertex AI. [n. d.]. Multimodal Live API. https://cloud.google. com/vertex-ai/generative-ai/docs/multimodal-live-api

  5. [5]

    Rahaf Alharbi, Pa Lor, Jaylin Herskovitz, Sarita Schoenebeck, and Robin N Brewer. 2024. Misfitting With AI: How Blind People Verify and Contest AI Errors. In Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility. 1–17

  6. [6]

    Riku Arakawa, Jill Fain Lehman, and Mayank Goel. 2024. Prism-q&a: Step-aware voice assistant on a smartwatch enabled by multimodal procedure tracking and large language models. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 4 (2024), 1–26

  7. [7]

    Riku Arakawa, Hiromu Yakura, and Mayank Goel. 2024. PrISM-Observer: Intervention agent to help users perform everyday procedures sensed using a smartwatch. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–16

  8. [8]

    Rob Argent, Ailish Daly, and Brian Caulfield. 2018. Patient involvement with home-based exercise programs: can connected health interventions influence adherence? JMIR mHealth and uHealth 6, 3 (2018), e8518

Show all 137 references
  1. [9]

    Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani, and Kristen Grauman. 2024. ExpertAF: Expert actionable feedback from video.arXiv preprint arXiv:2408.00672 (2024)

  2. [10]

    Kumar Ashutosh, Santhosh Kumar Ramakrishnan, Triantafyllos Afouras, and Kristen Grauman. 2023. Video-mined task graphs for keystep recognition in instructional videos. Advances in Neural Information Processing Systems 36 (2023), 67833–67846

  3. [11]

    Kumar Ashutosh, Zihui Xue, Tushar Nagarajan, and Kristen Grauman. 2024. Detours for navigating instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18804–18815

  4. [12]

    Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, and Mike Zheng Shou. 2024. Hallucination of multimodal large language models: A survey. arXiv preprint arXiv:2404.18930 (2024)

  5. [13]

    Nikola Banovic, Tovi Grossman, Justin Matejka, and George Fitzmaurice. 2012. Waken: reverse engineering usage information and interface structure from software videos. In Proceedings of the 25th annual ACM symposium on User interface software and technology . 83–92

  6. [14]

    BBC. 2018. Mary Berry’s tasty eggs Benedict Florentine - Classic Mary Berry - BBC. https://www.youtube.com/watch?v=YybJTrdwWQk

  7. [15]

    Hugh Beyer and Karen Holtzblatt. 1999. Contextual design. interactions 6, 1 (1999), 32–42

  8. [16]

    Jeffrey P Bigham, Chandrika Jayant, Hanjie Ji, Greg Little, Andrew Miller, Robert C Miller, Robin Miller, Aubrey Tatarowicz, Brandyn White, Samual White, et al . 2010. Vizwiz: nearly real-time answers to visual questions. In Proceedings of the 23nd annual ACM symposium on User...

  9. [17]

    Marie Claire Bilyk, Jessica M Sontrop, Gwen E Chapman, Susan I Barr, and Linda Mamer. 2009. Food experiences and eating patterns of visually impaired and blind people. Canadian Journal of Dietetic practice and research 70, 1 (2009), 13–18

  10. [18]

    Chi-Min Chan, Chunpu Xu, Ruibin Yuan, Hongyin Luo, Wei Xue, Yike Guo, and Jie Fu. 2024. Rq-rag: Learning to refine queries for retrieval augmented generation. arXiv preprint arXiv:2404.00610 (2024)

  11. [19]

    Minsuk Chang, Mina Huh, and Juho Kim. 2021. Rubyslippers: Supporting content-based voice navigation for how-to videos. In Proceedings of the 2021 CHI conference on human factors in computing systems . 1–14

  12. [20]

    Ruei-Che Chang, Yuxuan Liu, and Anhong Guo. 2024. WorldScribe: Towards Context-Aware Live Visual Descriptions. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–18

  13. [21]

    Pei-yu Chi, Jen-hao Chen, Hao-hua Chu, and Bing-Yu Chen. 2007. Enabling nutrition-aware cooking in a smart kitchen. In CHI’07 extended abstracts on Human factors in computing systems . 2333–2338

  14. [22]

    Pei-Yu Chi, Joyce Liu, Jason Linder, Mira Dontcheva, Wilmot Li, and Bjoern Hartmann. 2013. Democut: generating concise instructional videos for physical demonstrations. In Proceedings of the 26th annual ACM symposium on User interface software and technology . 141–150

  15. [23]

    Ken Click. 2024. Crisp Tortilla Pizza! https://www.youtube.com/watch?v= 2U45pP3i85g

  16. [24]

    Rob Comber, Jettie Hoonhout, Aart Van Halteren, Paula Moynihan, and Patrick Olivier. 2013. Food practices as situated action: exploring and designing for everyday food practices with households. InProceedings of the SIGCHI conference on human factors in computing systems . 245...

  17. [25]

    Elyse M Connors, Polly M Abbott, Daniel E Norris, Jennifer J Ottowitz, and Brigitte N Morren. 2023. The Perspectives of Vision Rehabilitation Therapists on the State of the Profession: A Time for Action? Journal of Visual Impairment & Blindness 117, 4 (2023), 303–313

  18. [26]

    Cook! Stacey Cook. 2024. Beef And Broccoli Stir Fry | Beef Stir Fry With Vegetables. https://www.youtube.com/watch?v=BBABeZjlRM8

  19. [27]

    Crouton Crackerjacks. 2014. How to Make Tiramisu!! Classic Italian Dessert Recipe. https://www.youtube.com/watch?v=bvVH4Mk2ku4

  20. [28]

    Crouton Crackerjacks. 2017. How to Make Strawberry Jam!! Homemade Small Batch Preserves Recipe. https://www.youtube.com/watch?v=F5LhDkAfxA8

  21. [29]

    Google Deepmind. [n. d.]. Project Astra. https://deepmind.google/models/ project-astra/

  22. [30]

    Nikita Dvornik, Isma Hadji, Ran Zhang, Konstantinos G Derpanis, Richard P Wildes, and Allan D Jepson. 2023. Stepformer: Self-supervised step discovery and localization in instructional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition...

  23. [31]

    Fallow. 2023. POV: How to Make an Omelette Like a Chef. https://www. youtube.com/watch?v=fqqwFWqxUr4

  24. [32]

    C Ailie Fraser, Tricia J Ngoon, Mira Dontcheva, and Scott Klemmer. 2019. Re- Play: contextually presenting learning videos across software applications. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . 1–13

  25. [33]

    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. 2024. Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. arXiv preprint arXiv:2405.21075 (2024)

  26. [34]

    Charles Goodwin and John Heritage. 1990. Conversation analysis. Annual review of anthropology 19 (1990), 283–307

  27. [35]

    Tovi Grossman, Justin Matejka, and George Fitzmaurice. 2010. Chronicle: cap- ture, exploration, and playback of document workflow histories. In Proceedings of the 23nd annual ACM symposium on User interface software and technology . 143–152

  28. [36]

    Sandra G Hart and Lowell E Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Advances in psychology. Vol. 52. Elsevier, 139–183

  29. [37]

    HoneySuckle. [n. d.]. World’s Best CHOCOLATE CHIP COOKIES Recipe: Crunchy Outside, Soft & Chewy Inside. https://www.youtube.com/watch?v=f- M3JN_7LGU

  30. [38]

    Honeysuckle. 2020. World’s Best CHOCOLATE CHIP COOKIES Recipe: Crunchy Outside, Soft & Chewy Inside. https://www.youtube.com/watch?v=f-M3JN_ 7LGU

  31. [39]

    Baixiang Huang, Canyu Chen, and Kai Shu. 2024. Can large language models identify authorship? arXiv preprint arXiv:2403.08213 (2024)

  32. [40]

    Ting-Hao Huang, Joseph Chee Chang, and Jeffrey P Bigham. 2018. Evorus: A crowd-powered conversational assistant built to automate itself over time. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–13

  33. [41]

    Mina Huh, YunJung Lee, Dasom Choi, Haesoo Kim, Uran Oh, and Juho Kim. 2022. Cocomix: utilizing comments to improve non-visual webtoon accessibility. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–18

  34. [42]

    Mina Huh and Amy Pavel. 2024. DesignChecker: Visual Design Support for Blind and Low Vision Web Developers. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–19

  35. [43]

    Mina Huh, Yi-Hao Peng, and Amy Pavel. 2023. GenAssist: Making image generation accessible. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–17

  36. [44]

    Mina Huh, Fangyuan Xu, Yi-Hao Peng, Chongyan Chen, Hansika Murugu, Danna Gurari, Eunsol Choi, and Amy Pavel. 2024. Long-Form Answers to Visual Questions from Blind and Low Vision People. arXiv preprint arXiv:2408.06303 (2024)

  37. [45]

    Mina Huh, Saelyne Yang, Yi-Hao Peng, Xiang’Anthony’ Chen, Young-Ho Kim, and Amy Pavel. 2023. Avscript: Accessible video editing with audio-visual scripts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17

  38. [46]

    Alyssa Hwang, Natasha Oza, Chris Callison-Burch, and Andrew Head. 2023. Rewriting the script: Adapting text instructions for voice interaction. In Pro- ceedings of the 2023 ACM designing interactive systems conference . 2233–2248

  39. [47]

    Ifrah Idrees, Tian Yun, Naveen Sharma, Yunxin Deng, Nakul Gopalan, George Konidaris, and Stefanie Tellex. 2023. Improved Inference of Human Intent by Combining Plan Recognition and Language Feedback. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (...

  40. [48]

    Junichi Ishikiriyama and Kenji Suzuki. 2017. An interactive virtual mirror to support makeup for visually impaired persons. In 2017 IEEE International Conference on Systems, Man, and Cybernetics (SMC) . IEEE, 1393–1398

  41. [49]

    Razan Jaber, Sabrina Zhong, Sanna Kuoppamäki, Aida Hosseini, Iona Gessinger, Duncan P Brumby, Benjamin R Cowan, and Donald Mcmillan. 2024. Cooking With Agents: Designing Context-aware Voice Interaction. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Sy...

  42. [50]

    Lucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl, and Shiri Azenkot

  43. [51]

    Liqiang Jing, Ruosen Li, Yunmo Chen, and Xinya Du. 2023. FaithScore: Fine- grained Evaluations of Hallucinations in Large Vision-Language Models. arXiv preprint arXiv:2311.01477 (2023)

  44. [52]

    Johnny. 2020. How to Make the Meat filling for Dumplings (Mandu). https: //www.youtube.com/watch?v=61bEy8CX53c

  45. [53]

    The Wallstreet Journal. [n. d.]. Meta’s AI-Powered Ray-Bans Are Life-Enhancing for the Blind. https://www.wsj.com/tech/ai/metas-ai-powered-ray-bans-are- life-enhancing-for-the-blind-3ae38026

  46. [54]

    Wendy Ju, Rebecca Hurwitz, Tilke Judd, and Bonny Lee. 2001. CounterActive: an interactive cookbook for the kitchen counter. In CHI’01 extended abstracts on Human factors in computing systems . 269–270

  47. [55]

    Jeongyeon Kim, Daeun Choi, Nicole Lee, Matt Beane, and Juho Kim. 2023. Surch: Enabling structural search and comparison for surgical videos. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–17

  48. [56]

    Juho Kim, Phu Tran Nguyen, Sarah Weir, Philip J Guo, Robert C Miller, and Krzysztof Z Gajos. 2014. Crowdsourcing step-by-step information extraction to enhance existing how-to videos. In Proceedings of the SIGCHI conference on human factors in computing systems . 4017–4026

  49. [57]

    Natashas Kitchen. [n. d.]. Dessert: Easy Mini Pavlovas - Homemade Meringues Recipe. https://www.youtube.com/watch?v=Zo5ATW4eq8o

  50. [58]

    Natasha’s Kitchen. 2016. Dessert: Easy Mini Pavlovas - Homemade Meringues Recipe. https://www.youtube.com/watch?v=Zo5ATW4eq8o

  51. [59]

    Preppy Kitchen. 2021. Mashed Potatoes Recipe. https://www.youtube.com/ watch?v=HfdFlenF6XI

  52. [60]

    KQED. 2020. Bread Flapjacks | Jacques Pépin Cooking At Home | KQED. https: //www.youtube.com/watch?v=86CeN5AFMG0

  53. [61]

    Aleksandra Królak, Weiqin Chen, Norun C Sanderson, and Siri Kessel. 2017. The accessibility of MOOCs for blind learners. InProceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility . 401–402

  54. [62]

    Jaewook Lee, Andrew D Tjahjadi, Jiho Kim, Junpu Yu, Minji Park, Jiawen Zhang, Jon E Froehlich, Yapeng Tian, and Yuhang Zhao. 2024. CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low Vision. In Proceedings of the 37th Annual...

  55. [63]

    Jaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu, Sebastian S Rodriguez, and Jon E Froehlich. 2024. GazePointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality. In Proceedings of the 2024 CHI Conference on Human Factors in C...

  56. [64]

    Yibin Lei, Yu Cao, Tianyi Zhou, Tao Shen, and Andrew Yates. 2024. Corpus- steered query expansion with large language models. arXiv preprint arXiv:2402.18031 (2024)

  57. [65]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing...

  58. [66]

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. 2024. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326 (2024)

  59. [67]

    Chaoyu Li, Sid Padmanabhuni, Maryam Cheema, Hasti Seifi, and Pooyan Fazli

  60. [68]

    Franklin Mingzhe Li, Jamie Dorst, Peter Cederberg, and Patrick Carrington. 2021. Non-visual cooking: exploring practices and challenges of meal preparation by people with visual impairments. In Proceedings of the 23rd International ACM SIGACCESS Conference on Computers and Acc...

  61. [69]

    Franklin Mingzhe Li, Michael Xieyang Liu, Shaun K Kane, and Patrick Carring- ton. 2024. A Contextual Inquiry of People with Vision Impairments in Cooking. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–14

  62. [70]

    Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, and Patrick Carrington. 2025. OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking. arXiv preprint arXiv:2503.05962 (2025)

  63. [71]

    It Feels Like Taking a Gamble

    Franklin Mingzhe Li, Franchesca Spektor, Meng Xia, Mina Huh, Peter Cederberg, Yuqi Gong, Kristen Shinohara, and Patrick Carrington. 2022. “It Feels Like Taking a Gamble”: Exploring Perceptions, Practices, and Challenges of Using Makeup and Cosmetics for People with Visual Impa...

  64. [72]

    Franklin Mingzhe Li, Ashley Wang, Patrick Carrington, and Shaun K Kane. 2024. A Recipe for Success? Exploring Strategies for Improving Non-Visual Access to Cooking Instructions. In Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility . ...

  65. [73]

    Ching Liu, Juho Kim, and Hao-Chuan Wang. 2018. ConceptScape: Collaborative concept mapping for video learning. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–12

  66. [74]

    Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti, Lei Ding, Yi Zhang, and Leilani H Gilpin. 2024. Right this way: Can VLMs Guide Us to See More to Answer Questions? arXiv preprint arXiv:2411.00394 (2024)

  67. [75]

    Xingyu Liu, Patrick Carrington, Xiang’Anthony’ Chen, and Amy Pavel. 2021. What makes videos accessible to blind and visually impaired people?. In Pro- ceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–14

  68. [76]

    Xingyu" Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang Anthony Chen, and Amy Pavel. 2022. Crossa11y: Identifying video accessibility issues via cross-modal grounding. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–14

  69. [77]

    Microsoft. [n. d.]. Microsoft AI Audio Descriptions. https://github.com/ microsoft/ai-audio-descriptions

  70. [78]

    Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. 2019. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF international conference on computer vision ....

  71. [79]

    Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. Factscore: Fine-grained atomic evaluation of factual precision in long form text generation. arXiv preprint arXiv:2305.14251 (2023)

  72. [80]

    Alok Mysore and Philip J Guo. 2018. Porta: Profiling software tutorials using operating-system-wide activity tracing. In Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology . 201–212

  73. [81]

    Archana Narayanan, Erzhen Hu, and Seongkook Heo. 2022. Enabling Remote Hand Guidance in Video Calls Using Directional Force Illusion. In Companion Publication of the 2022 Conference on Computer Supported Cooperative Work and Social Computing. 135–139

  74. [82]

    Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, and Ko- taro Hara. 2024. Audio description customization. In Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility . 1–19

  75. [83]

    Cuong Nguyen and Feng Liu. 2015. Making software tutorial video responsive. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems. 1565–1568

  76. [84]

    Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: interactive video content exploration through augmented audio descriptions for blind or low-vision view- ers. In Proceedings of the 2024 CHI Conferenc...

  77. [85]

    OpenAI. [n. d.]. OpenAI Whisper. https://openai.com/index/whisper/

  78. [86]

    Korinn Ostrow and Neil Heffernan. 2014. Testing the multimedia principle in the real world: a comparison of video vs. Text feedback in authentic middle school math assignments. In Educational Data Mining 2014

  79. [87]

    Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Recasens, Larisa Mar- keeva, Dylan Banarse, Skanda Koppula, Mateusz Malinowski, Yi Yang, Carl Doersch, et al. 2023. Perception test: A diagnostic benchmark for multimodal video models. Advances in Neural Information Process...

  80. [88]

    Amy Pavel, Colorado Reed, Björn Hartmann, and Maneesh Agrawala. 2014. Video digests: a browsable, skimmable format for informational lecture videos.. In UIST, Vol. 10. Citeseer, 2642918–2647400

  81. [89]

    Amy Pavel, Gabriel Reyes, and Jeffrey P Bigham. 2020. Rescribe: Authoring and automatically editing audio descriptions. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology . 747–759

  82. [90]

    Rohith Peddi, Shivvrat Arya, Bharath Challa, Likhitha Pallapothula, Akshay Vyas, Bhavya Gouripeddi, Qifan Zhang, Jikai Wang, Vasundhara Komaragiri, Eric Ragan, et al. 2024. CaptainCook4D: A dataset for understanding errors in procedural activities. Advances in Neural Informati...

  83. [91]

    Yi-Hao Peng, Jeffrey P Bigham, and Amy Pavel. 2021. Slidecho: Flexible non- visual exploration of presentation videos. InProceedings of the 23rd International ACM SIGACCESS Conference on Computers and Accessibility . 1–12

  84. [92]

    Yi-Hao Peng, JiWoong Jang, Jeffrey P Bigham, and Amy Pavel. 2021. Say it all: Feedback for improving non-visual presentation accessibility. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–12

  85. [93]

    Ethan Prihar, Aaron Haim, Tracy Shen, Adam Sales, Dongwon Lee, Xintao Wu, and Neil Heffernan. 2023. Investigating the Impact of Skill-Related Videos on Online Learning. In Proceedings of the Tenth ACM Conference on Learning@ Scale. 4–13

  86. [94]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al

  87. [95]

    Gordon Ramsay. [n. d.]. How To Cook Eggs Benedict | Gordon Ramsay. https: //www.youtube.com/watch?v=gBJjRYk0yC0

  88. [96]

    Gordon Ramsey. 2018. How To Cook Eggs Benedict | Gordon Ramsay. https: //www.youtube.com/watch?v=f-M3JN_7LGU

  89. [97]

    Kyle Rector, Cynthia L Bennett, and Julie A Kientz. 2013. Eyes-free yoga: an exergame using depth cameras for blind & low vision exercise. In Proceedings of the 15th international acm sigaccess conference on computers and accessibility . 1–8

  90. [98]

    Book Sadprasid, Carl Gutwin, and Scott Bateman. 2024. Improving Video Navigation for Spatial Task Tutorials by Spatially Segmenting and Situating How-To Videos. In Proceedings of the 2024 ACM Symposium on Spatial User Interaction. 1–13

  91. [99]

    Woosuk Seo and Hyunggu Jung. 2018. Understanding blind or visually impaired people on youtube through qualitative analysis of videos. In Proceedings of the 2018 ACM International Conference on Interactive Experiences for TV and Online Video. 191–196

  92. [100]

    Yale Song, Eugene Byrne, Tushar Nagarajan, Huiyu Wang, Miguel Martin, and Lorenzo Torresani. 2024. Ego4d goal-step: Toward hierarchical understanding of procedural activities. Advances in Neural Information Processing Systems 36 (2024)

  93. [101]

    Sweetology. [n. d.]. Gingerbread House Assembly. https://www.youtube.com/ watch?v=ZW6plAZ-dkY

  94. [102]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805 (2023)

  95. [103]

    Just Jordan Things. [n. d.]. How To Build a Flower Arrangement In Only 10 Minutes. https://www.youtube.com/watch?v=9PVFYYLjN-w

  96. [104]

    Balasaravanan Thoravi Kumaravel, Cuong Nguyen, Stephen DiVerdi, and Björn Hartmann. 2019. TutoriVR: A video-based tutorial system for design applications in virtual reality. In Proceedings of the 2019 CHI conference on human factors in computing systems. 1–12

  97. [105]

    TikTok. [n. d.]. TikTok. https://tiktok.com

  98. [106]

    Tram Thi Minh Tran, Shane Brown, Oliver Weidlich, Soojeong Yoo, and Callum Parker. 2025. Wearable AR in Everyday Contexts: Insights from a Digital Ethnography of YouTube Videos. arXiv preprint arXiv:2502.06191 (2025)

  99. [107]

    Anh Truong, Peggy Chi, David Salesin, Irfan Essa, and Maneesh Agrawala. 2021. Automatic generation of two-level hierarchical tutorials from instructional makeup videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–16

  100. [108]

    Tess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C Derry, Mina Huh, and Amy Pavel. 2024. Making short-form videos accessible with hierarchical video sum- maries. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–17

  101. [109]

    Marynel Vázquez and Aaron Steinfeld. 2014. An assisted photography frame- work to help visually impaired users properly aim a camera. ACM Transactions on Computer-Human Interaction (TOCHI) 21, 5 (2014), 1–29

  102. [110]

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv preprint arXiv:2409.12191 (2024)

  103. [111]

    Yujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang, Dingzeyu Li, and Lap- Fai Yu. 2021. Toward automatic audio description generation for accessible videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–12

  104. [112]

    Maximiliane Windl and Sven Mayer. 2022. The skewed privacy concerns of bystanders in smart environments. Proceedings of the ACM on Human-Computer Interaction 6, MHCI (2022), 1–21

  105. [113]

    Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis

  106. [114]

    Shuchang Xu, Xiaofu Jin, Huamin Qu, and Yukang Yan. 2025. DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions. arXiv preprint arXiv:2501.15711 (2025)

  107. [115]

    Zihui Xue, Joungbin An, Xitong Yang, and Kristen Grauman. 2024. Progress- Aware Video Frame Captioning. arXiv preprint arXiv:2412.02071 (2024)

  108. [116]

    Zihui Xue, Kumar Ashutosh, and Kristen Grauman. 2024. Learning object state changes in videos: An open-world perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18493–18503

  109. [117]

    Saelyne Yang, Sangkyung Kwak, Juhoon Lee, and Juho Kim. 2023. Beyond In- structions: A Taxonomy of Information Types in How-to Videos. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–21

  110. [118]

    Saelyne Yang, Anh Truong, Juho Kim, and Dingzeyu Li. 2025. VideoMix: Aggre- gating How-To Videos for Task-Oriented Learning. InProceedings of the 30th International Conference on Intelligent User Interfaces . 1564–1580

  111. [119]

    Saelyne Yang, Jisu Yim, Aitolkyn Baigutanova, Seoyoung Kim, Minsuk Chang, and Juho Kim. 2022. SoftVideo: Improving the Learning Experience of Software Tutorial Videos with Collective Interaction Data. In Proceedings of the 27th International Conference on Intelligent User Inte...

  112. [120]

    YouDescribe. [n. d.]. YouDescribe - Audio Description for YouTube Videos. https://youdescribe.org/

  113. [121]

    YouTube. [n. d.]. YouTube. https://www.youtube.com/

  114. [122]

    Zhuohao Zhang and Jacob O Wobbrock. 2023. A11yboard: making digital artboards accessible to blind and low-vision users. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–17

  115. [123]

    Rewind to the Jiggling Meat Part

    Yaxi Zhao, Razan Jaber, Donald McMillan, and Cosmin Munteanu. 2022. “Rewind to the Jiggling Meat Part”: Understanding Voice Control of Instructional Videos in Everyday Tasks. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–11

  116. [124]

    Yaoyao Zhong, Junbin Xiao, Wei Ji, Yicong Li, Weihong Deng, and Tat-Seng Chua. 2022. Video question answering: Datasets, algorithms and challenges. arXiv preprint arXiv:2203.01225 (2022)

  117. [125]

    You are going to add flour. One and a half cups of flour

    Honglu Zhou, Roberto Martín-Martín, Mubbasir Kapadia, Silvio Savarese, and Juan Carlos Niebles. 2023. Procedure-aware pretraining for instructional video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 10727–10738. UIST ’25,...

  118. [130]

    Hey, what’s up you guys, Chef [...] here

    Greeting Opening: Starting remarks and instructor/channel introductions. Example: "Hey, what’s up you guys, Chef [...] here." Closing: Parting remarks and wrap-up. Example: "Stay tuned, we’ll catch you all later."

  119. [131]

    Today, I’ll show you a special technique which is totally special and about image pressing

    Overview Goal: Main purpose of the video and its descriptions. Example: "Today, I’ll show you a special technique which is totally special and about image pressing." Motivation: Reasons or background information on why the video was created. Example: "[...] Someone is making a...

  120. [132]

    Now for the intricate layer that will give me the final webbing look

    Method Subgoal: Objective of a subsection. Example: "Now for the intricate layer that will give me the final webbing look." Instruction: Actions that the instructor performs to complete the task. Example: "We’re going to pour that into our silicone baking cups." Tool: Introduc...

  121. [133]

    I find that it’s easier to do just a couple of layers at a time instead of all four layers at a time

    Supplementary Tip: Additional instructions or information that makes instructions easier, faster, or more efficient. Example: "I find that it’s easier to do just a couple of layers at a time instead of all four layers at a time." Warning: Actions that should be avoided. Exampl...

  122. [134]

    Because every time we wear our contact lenses, makeup and even dirt particles [...] might harm our eyes directly

    Explanation Justification: Reasons why the instruction was performed. Example: "Because every time we wear our contact lenses, makeup and even dirt particles [...] might harm our eyes directly." Effect: Consequences of the instruction. Example: "And these will overhang a littl...

  123. [135]

    Something sticky and dirty all through the back seat

    Description Status: Descriptions of the current state of the target object. Example: "Something sticky and dirty all through the back seat." Context: Descriptions of the method or the setting. Example: "[...] The process of putting on a tip by hand [...] takes a lot of patienc...

  124. [136]

    And now we have a dinosaur taggy blanket that wrinkles, so a fun gift for any baby on your gift giving list

    Conclusion Outcome: Descriptions of the final results of the procedure. Example: "And now we have a dinosaur taggy blanket that wrinkles, so a fun gift for any baby on your gift giving list." Reflection: Summary, evaluation, and suggestions for the future about the overall pro...

  125. [137]

    Tristan is back from basketball - He made it on the team so it’s pretty exciting

    Miscellaneous Side Note: Personal stories, jokes, user engagement, and advertisements. Example: "Tristan is back from basketball - He made it on the team so it’s pretty exciting." Self-promotion: Promotion of the instructor of the channel (i.e. likes, subscription, notificatio...

  126. [2021]

    In International conference on machine learning

    Learning transferable visual models from natural language supervision. In International conference on machine learning . PmLR, 8748–8763

  127. [2023]

    arXiv preprint arXiv:2309.17453 (2023)

    Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453 (2023)

  128. [2024]

    It’s Kind of Context Dependent

    “It’s Kind of Context Dependent”: Understanding Blind and Low Vision People’s Video Accessibility Preferences Across Viewing Scenarios. In Proceed- ings of the CHI Conference on Human Factors in Computing Systems . 1–20

  129. [2025]

    arXiv preprint arXiv:2502.20480 (2025)

    VideoA11y: Method and Dataset for Accessible Video Description. arXiv preprint arXiv:2502.20480 (2025)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.