Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:27:18.614377Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 3 inbound Pith citation observations for arXiv:2501.12231.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:27:18.614377Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T22:00:28.350003Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T17:27:15.161769Z
91 of 91 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0f054d61-0dba-499f-bdf7-2471f4a6d132 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc7e294-1f2e-42a4-baa8-c94c11f1c69b · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Ht-step: Aligning instructional articles with how-to videos
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd5688d-92c4-439c-9b05-de697e4f6142 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Unsuper- vised learning from narrated instruction videos
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a441990-3e0c-4b3b-9787-2254fbc2ab74 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Flamingo: a visual language model for few-shot learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9fb148e-45cb-41a0-b77a-feb5ee36bcc4 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Localizing moments in video with natural language
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f91d33f1-4595-4952-b0cb-f65cf552446a · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models https://www.anthropic.com/news/developing- computer-use, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f612e43d-53cd-4fbc-b5f4-4c457b0e5899 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Video-mined task graphs for keystep recognition in instructional videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd3ec09d-363e-4712-a532-65b93d9943f6 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Detours for navigating instructional videos
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a74ed87-5bf5-4e06-b027-655d79e4f239 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Is space-time attention all you need for video understanding? In ICML, page 4, 2021
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9d14d8-935c-4e9b-a745-08498ce1091a · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure planning in instructional videos via contextual modeling and model-based policy learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe6c2e33-9f1f-4e96-bbd2-b2531c6b019d · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Activitynet: A large-scale video bench- mark for human activity understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef2fe9c7-d606-4e31-9992-1f3a3d34a5f7 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure planning in instructional videos
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c129299-5f51-4509-88bb-069ccd546940 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporally grounding natural sentence in video
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0564b7c4-b66d-427d-b20f-8925c9539acf · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Videollm-online: Online video large language model for streaming video
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 86bbecf3-2c85-4636-a1de-28804f6ceac4 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Semantic proposal for activity localization in videos via sentence query
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f00acf7c-0e9b-4e87-bb24-a4de1732b50c · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models KGPT: Knowledge-grounded pre-training for data-to-text generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c78474aa-08c1-4bfe-9001-d22221ae367e · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2ed87957-4a6a-47db-acb3-e2d79d793410 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models From Local to Global: A Graph RAG Approach to Query-Focused Summarization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04cc3f4-08c4-41c6-bd1a-7323530553ce · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b72e409f-5eea-483d-8074-0b85a7d1c6be · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Slowfast networks for video recognition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 248d6cc9-cd8e-4fe7-aeb6-f54c9509f8ed · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Tall: Temporal activity localization via language query
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629d54ac-d2a0-4cf0-a721-22f9663b7c5e · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc755edc-ea2f-4e1c-8cd0-93431f93e330 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mac: Mining activity concepts for language-based temporal local- ization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537d9060-b12a-4da7-8cec-e84dec96dfc6 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Imagebind: One embedding space to bind them all
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d1ae0f-f85d-46d6-b357-d1cc5c09400d · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Radar: automated task planning for proactive decision support
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 170d0ab5-d947-44b1-9229-e48d410d753d · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval augmented language model pre- training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 246d6b37-dad7-46cf-83e6-2d983773ddd9 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fd43daab-b8d6-491e-b28b-8a999a1cfcae · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Denoising diffu- sion probabilistic models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d50f50e-e48c-4099-b735-611dcdb025ad · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models LoRA: Low-rank adaptation of large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 048154c4-394f-48b6-a5e4-456912107e24 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Audiogpt: Understanding and generating speech, music, sound, and talking head
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b5ed1157-9431-4931-a58e-1f8b0c20658a · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language is not all you need: Aligning perception with language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 300761ce-e88f-4c10-981d-959ac6b77067 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Leveraging passage retrieval with generative models for open domain question answering
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3ef6a6de-5536-4722-b939-7b113c4ef4fa · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mistral 7B
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0310a70-554e-4f9b-9f44-de75ffca7a9a · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Gen- erating images with multimodal language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aa40df6-34c4-4764-9a09-ecfe61dc25a9 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models WikiHow: A Large Scale Text Summarization Dataset
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bad06d2-7f42-44ae-96b8-95d227ef9d82 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models A survey on temporal sentence grounding in videos
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4f859fb8-ec1e-47e5-bdac-e4c9aa136e71 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Tvr: A large-scale dataset for video-subtitle moment retrieval
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b3dabebf-3f94-4b74-a5b4-6ec0448383ea · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Less is more: Clipbert for video-and-language learning via sparse sampling
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 596d4a41-5f75-49d6-be14-de21ac0339ef · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Chatting makes perfect: Chat-based image retrieval
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e9112a-b8ab-44e2-92bf-84a31838eade · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval- augmented generation for knowledge-intensive nlp tasks
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8f42b6c-0d31-470d-8001-d8f90161d8e3 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b40416-83e2-4efc-bbd4-0a23843c2bbb · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Skip-plan: Procedure planning in instructional videos via condensed action space learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a7f9a262-5fb1-419d-b595-0fca782911c5 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Cohn, and Janet B
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 40b457fb-a7cd-44a8-a55b-8708313d02b8 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning to recognize procedural activities with distant supervision
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4713edc8-2e9f-4a33-acec-590ae14c0c9d · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Jointly cross-and self-modal graph attention network for query-based moment localization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d8bfad8e-0392-4876-af25-c9613c35bd01 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Improved baselines with visual instruction tuning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8aec459-d859-4fbd-9f72-631649fb012d · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Visual instruction tuning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46da08f4-c521-47d8-8139-6a60145d8163 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Attentive moment retrieval in videos
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 26decaed-decf-42f3-85d4-e172b0f4d7a6 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language models of code are few-shot commonsense learners
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ebe4190f-1994-4be1-b33b-a277349828bf · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models What’s cookin’? interpreting cooking videos using text, speech and vision
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8abe8ae2-99f4-4a9e-8f58-fab46928dc7c · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0f3350a2-a0a3-4501-ad92-c2c0d25d5f24 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models End-to-end learn- ing of visual representations from uncurated instructional videos
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ad815325-c0ba-41c6-a0dd-6595078478a7 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning and Verification of Task Structure in Instructional Videos
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae29bc64-05ae-4610-b9f1-c961e70fb6e0 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models SCHEMA: State CHanges MAtter for procedure planning in instructional videos
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 46042b3c-b577-4cd8-a223-dec461577ac3 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Virtualhome: Sim- ulating household activities via programs
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1d462551-7fc9-4673-824b-d3ef509b4255 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning transferable visual models from natural language supervi- sion
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f28cd14f-2aa5-49f4-b2f8-1edd37df038a · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Grounding action descriptions in videos
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8f268d0b-267a-4bcd-b7ac-7362cd2a945d · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models FLAP: Flow-adhering planning with constrained decoding in LLMs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 69608492-668a-46ee-a34e-8077fe24d1a5 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models proScript: Par- tially ordered scripts generation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 88d099c9-b89b-4ee9-b1a2-b817958e3567 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 374b58e8-0949-41aa-ae82-225144ebb97e · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5c00df61-fe98-42ae-aa1e-54de161d9b76 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Moviechat: From dense token to sparse memory for long video understanding
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ebf6b249-9a9b-4921-b8c6-45eb5f078a1a · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mpnet: Masked and permuted pre-training for language understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c3a0b2e7-5820-41df-a9ef-2d276b799ea7 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language Models Can See: Plugging Visual Controls in Text Generation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6019f744-9b5d-4a78-9201-ffe1265ca81f · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models PandaGPT: One model to instruction-follow them all
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f9f1a374-9afe-4810-83be-9ba1448966c0 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Plate: Visually-grounded planning with transformers in procedural tasks
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 76415b0a-33bb-4fbd-98bb-489517242b8a · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Coin: A large-scale dataset for comprehensive instructional video analysis
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a2a0d350-3c6b-4c88-ad2a-6c639902c601 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models On the planning abilities of large language models-a critical investigation
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ecf651f8-0f62-4594-85cc-a4041e7331b5 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d7d46fa4-61ed-4a8f-8c20-036ce1a5e244 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Event-guided procedure planning from in- structional videos with text supervision
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aa1704bf-798f-4e21-9317-a7a01628461f · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Pdpp: Projected diffusion for procedure planning in instructional videos
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 641a44df-1b47-42d3-b018-956b5133109d · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporal segment networks: Towards good practices for deep action recognition
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5155ce14-7170-4533-a1bc-32587b290c6d · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f82821f-3eaa-4d58-9597-5dec1e9624c7 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models NExt-GPT: Any-to-any multimodal LLM
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bfe34a2f-97a6-47ab-8583-2618443845c0 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b4329a38-a98a-4794-bac9-b97886dcf609 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Translating Natural Language to Planning Goals with Large-Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d198cda2-5116-439d-ac92-f0bafaac0ec3 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Multilevel language and vision integration for text-to-clip retrieval
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bb148e8c-adcf-4f31-b9d8-0bd1fa4bb170 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models VideoCLIP: Contrastive pre- training for zero-shot video-text understanding
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6cc9b393-3050-4d91-a6f8-f455fe2de311 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval- augmented generation with knowledge graphs for customer service question answering
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f320f8fd-dd08-47a9-9add-799ba20f929d · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4a9e3b04-e34b-4f3f-bd07-db480e1e6fe9 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Distilling script knowledge from large language models for constrained language planning
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1e5d354d-f8e9-4e08-88f6-567a49c7bad2 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e47a14f2-1409-470a-b597-6372db2921c9 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models SpeechGPT: Empowering large language models with intrinsic cross-modal conversa- tional abilities
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 65d682a9-ad24-4ef0-9aff-985af49ebc05 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d628f934-164d-4b8b-a513-33d7a7b43b42 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporal sentence grounding in videos: A survey and fu- ture directions
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 703bc398-48a0-48c7-9795-3cb775a7a0ae · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models P3iv: Probabilistic procedure planning from instructional videos with weak su- pervision
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1bfea070-78ba-47be-a292-a0398e67453e · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning procedure-aware video repre- sentation from instructional videos and their narrations
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cd6107ab-a413-41f7-ad15-d3bcb5ac5f94 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure-aware pretraining for instructional video understanding
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7901844d-b6bf-424c-b788-b3bd49dd1263 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Towards auto- matic learning of procedures from web instructional videos
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d130331b-ff3b-4241-903e-a583413649ee · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models MiniGPT-4: Enhancing vision-language understanding with advanced large language models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f8de0d6b-998d-4e0c-bf16-2b794b0714b4 · outbound
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Cross- task weakly supervised learning from instructional videos
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1a817673-2a56-4d5d-916d-4a9b06e41a2b · inbound
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 94abaaac-ea99-486a-89b4-b240171202ed · inbound
VisionClaw: Always-On AI Agents through Smart Glasses InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3b18f4b1-b9d9-41b9-9fb7-55667d3f9008 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
Reference 256
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.