Pith. sign in

Paper Citation Record · LEDGER

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

As of 13 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 3 inbound Pith citation observations for arXiv:2501.12231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12231 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:27:18.614377Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T22:00:28.350003Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.161769Z

Reference resolution

91 of 91 outbound references displayed

  • verified exact1
  • verified fuzzy60
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f054d61-0dba-499f-bdf7-2471f4a6d132 · outbound

This paper cites GPT-4 Technical Report.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.147094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.147094Z digest=sha256:8e44e475eb38fca335aad72e51d8afaf4f538743dbc211ea00cda5de2c6e32e7

Observation bbc7e294-1f2e-42a4-baa8-c94c11f1c69b · outbound

This paper cites Ht-step: Aligning instructional articles with how-to videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Ht-step: Aligning instructional articles with how-to videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.153382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.153382Z digest=sha256:d7f14e65d12e88fb59fd388b6eaf33acd015a09657775fe3f5b517fa4ed4a70c

Observation 0fd5688d-92c4-439c-9b05-de697e4f6142 · outbound

This paper cites Unsuper- vised learning from narrated instruction videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Unsuper- vised learning from narrated instruction videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.159361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.159361Z digest=sha256:d2538af250a0535a6e839ef95bb7401a06c38f0424e083f8fc419fcb4af292bb

Observation 3a441990-3e0c-4b3b-9787-2254fbc2ab74 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.164584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.164584Z digest=sha256:a68440bbf39f7a9e09f7641be3d09b32158ee19b08139938b5d9a09fb5a97fd8

Observation c9fb148e-45cb-41a0-b77a-feb5ee36bcc4 · outbound

This paper cites Localizing moments in video with natural language.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Localizing moments in video with natural language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.169866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.169866Z digest=sha256:64470f433c46c5d857552a27c52fbba0c5da9be1f87765d000c45e0d64cb396e

Observation f91d33f1-4595-4952-b0cb-f65cf552446a · outbound

This paper cites https://www.anthropic.com/news/developing- computer-use, 2024.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models https://www.anthropic.com/news/developing- computer-use, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.175183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.175183Z digest=sha256:9d1ed092eab768a78596a468b1cfbbd63fe5b79aa1a442ad064620f2f1a29d7f

Observation f612e43d-53cd-4fbc-b5f4-4c457b0e5899 · outbound

This paper cites Video-mined task graphs for keystep recognition in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Video-mined task graphs for keystep recognition in instructional videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.180832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.180832Z digest=sha256:ec5741e3cb85a9327a0a83c53cb33b14926fd7fc4514b1825f149ddea388f2d4

Observation bd3ec09d-363e-4712-a532-65b93d9943f6 · outbound

This paper cites Detours for navigating instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Detours for navigating instructional videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.186508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.186508Z digest=sha256:e7299770add36721f6c1517e00a2a846110dad209b521314c146c9c6f388036d

Observation 6a74ed87-5bf5-4e06-b027-655d79e4f239 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.191609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.191609Z digest=sha256:0c3010d2edd1a46dfe8cf6cbd9fbca7e5b48027f88e8c59ba894e1e979475512

Observation 5a9d14d8-935c-4e9b-a745-08498ce1091a · outbound

This paper cites Procedure planning in instructional videos via contextual modeling and model-based policy learning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure planning in instructional videos via contextual modeling and model-based policy learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.196739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.196739Z digest=sha256:518329387aaa7fecaa7c1907d3bdeba86c66fd27bf16b0b66cc319bd87548f1c

Observation fe6c2e33-9f1f-4e96-bbd2-b2531c6b019d · outbound

This paper cites Activitynet: A large-scale video bench- mark for human activity understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Activitynet: A large-scale video bench- mark for human activity understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.201677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.201677Z digest=sha256:621affe1d8432e057612ba8e0d9a698b068374cf666c93262c6876ad7daead0f

Observation ef2fe9c7-d606-4e31-9992-1f3a3d34a5f7 · outbound

This paper cites Procedure planning in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure planning in instructional videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.206432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.206432Z digest=sha256:a41b9eeba4fb3f7112ae48298b3a46710ce82924dc97012ea1fc5f8a82c85f8c

Observation 9c129299-5f51-4509-88bb-069ccd546940 · outbound

This paper cites Temporally grounding natural sentence in video.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporally grounding natural sentence in video

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.077125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.211601Z digest=sha256:6e71f6f58128fb1d27d852e6b55cf6367ff1d68e959c20c40cc9b69ba08b92df

Observation 0564b7c4-b66d-427d-b20f-8925c9539acf · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Videollm-online: Online video large language model for streaming video

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.051969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.216331Z digest=sha256:e02fc32a08a3e82c481af523bffe15c626a2c2517b61d9b0a01a09c2bbb616ec

Observation 86bbecf3-2c85-4636-a1de-28804f6ceac4 · outbound

This paper cites Semantic proposal for activity localization in videos via sentence query.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Semantic proposal for activity localization in videos via sentence query

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.032877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.221166Z digest=sha256:c6371d7200fda83c11c365da88f9277eba8ac5efc983c1486fa83709a7956f56

Observation f00acf7c-0e9b-4e87-bb24-a4de1732b50c · outbound

This paper cites KGPT: Knowledge-grounded pre-training for data-to-text generation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models KGPT: Knowledge-grounded pre-training for data-to-text generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:20.014111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.226036Z digest=sha256:c16a14382864f4cebcf1c92308c44225c209e95670a70e4471a0ec506d037b80

Observation c78474aa-08c1-4bfe-9001-d22221ae367e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.995691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.230690Z digest=sha256:b956596cb7586c444fbe252c4d1b2da0b624c146fc85af43fd3da438ef096b5b

Observation 2ed87957-4a6a-47db-acb3-e2d79d793410 · outbound

This paper cites From Local to Global: A Graph RAG Approach to Query-Focused Summarization.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models From Local to Global: A Graph RAG Approach to Query-Focused Summarization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.235459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.235459Z digest=sha256:c45cdc43dce0b1803f26766dc6f13deb0f0e394b11e069853b491f31b55ae7e3

Observation d04cc3f4-08c4-41c6-bd1a-7323530553ce · outbound

This paper cites Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:27:18.797148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.240608Z digest=sha256:b28a04aab7e108013dc79ae688e39df4b362f02ad7746e9f7591c994a055a78c

Observation b72e409f-5eea-483d-8074-0b85a7d1c6be · outbound

This paper cites Slowfast networks for video recognition.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Slowfast networks for video recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.977834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.246444Z digest=sha256:f2a3ec80abcf763d332e2f307ddecf1b1e6caa7e2f3af6a6e311bcdfeb3be626

Observation 248d6cc9-cd8e-4fe7-aeb6-f54c9509f8ed · outbound

This paper cites Tall: Temporal activity localization via language query.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Tall: Temporal activity localization via language query

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.251469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.251469Z digest=sha256:e7b55d84812c27ca5d44bfd271451fb6c8de4abc866baeaf86b6fe3ff42c997f

Observation 629d54ac-d2a0-4cf0-a721-22f9663b7c5e · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.256148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.256148Z digest=sha256:d13804c0cc43dd959da9c9bdfe7c048e7398317ef1bbe0891554cabd1911b915

Observation cc755edc-ea2f-4e1c-8cd0-93431f93e330 · outbound

This paper cites Mac: Mining activity concepts for language-based temporal local- ization.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mac: Mining activity concepts for language-based temporal local- ization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.261755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.261755Z digest=sha256:eb2c605346b7d327d185c8bd08b018e3db50b1701eb1a2c8d7e26fb71e2c5b05

Observation 537d9060-b12a-4da7-8cec-e84dec96dfc6 · outbound

This paper cites Imagebind: One embedding space to bind them all.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Imagebind: One embedding space to bind them all

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.266660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.266660Z digest=sha256:7abd8ec96b9edb6655688820ef44cecb74d394d89cf71fc9754e7c93fc797f23

Observation 81d1ae0f-f85d-46d6-b357-d1cc5c09400d · outbound

This paper cites Radar: automated task planning for proactive decision support.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Radar: automated task planning for proactive decision support

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.928432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.271492Z digest=sha256:7d893c280c260c6aa6d6d878458aeeadc337c77bf34951490bf8b0361e87e73c

Observation 170d0ab5-d947-44b1-9229-e48d410d753d · outbound

This paper cites Retrieval augmented language model pre- training.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval augmented language model pre- training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.911731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.276299Z digest=sha256:078d87e3f49ea4435e36078aa06ea14901d92dcf231c0c983beae8fe9363ef64

Observation 246d6b37-dad7-46cf-83e6-2d983773ddd9 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.892002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.281075Z digest=sha256:76759ef0cdcaee8a571fc6214982927810b84804605ff2e68c992d1e62b4dbea

Observation fd43daab-b8d6-491e-b28b-8a999a1cfcae · outbound

This paper cites Denoising diffu- sion probabilistic models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Denoising diffu- sion probabilistic models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.286179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.286179Z digest=sha256:7a17dba6f471b2eeea0820d8804d68fc3e46b8b30ee5f54a36092cc234749b19

Observation 2d50f50e-e48c-4099-b735-611dcdb025ad · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models LoRA: Low-rank adaptation of large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.864426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.291143Z digest=sha256:3c5871a5dfdebe417a7411a5ef629d3355250dc5c22c97d3f584b39be9ede06a

Observation 048154c4-394f-48b6-a5e4-456912107e24 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.846787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.296118Z digest=sha256:a2c8e300e2db95d419f9db663b85bcaac9c5826137c1cbab1b55809bafd891c8

Observation b5ed1157-9431-4931-a58e-1f8b0c20658a · outbound

This paper cites Language is not all you need: Aligning perception with language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language is not all you need: Aligning perception with language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.829639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.301317Z digest=sha256:86f54890e2f2ec62889bf3af4d8dd5d35f5ba5f22d948316d9662caf8c005bc6

Observation 300761ce-e88f-4c10-981d-959ac6b77067 · outbound

This paper cites Leveraging passage retrieval with generative models for open domain question answering.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Leveraging passage retrieval with generative models for open domain question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.813004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.306835Z digest=sha256:63998221660f9ff014dfa04718a115c4b905e1ba195cf349c78bf3e8ea333c81

Observation 3ef6a6de-5536-4722-b939-7b113c4ef4fa · outbound

This paper cites Mistral 7B.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mistral 7B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.311567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.311567Z digest=sha256:db93df32559562eeacd7fa6372e5068a7373e7d3b62da56623345e2efddaa48c

Observation a0310a70-554e-4f9b-9f44-de75ffca7a9a · outbound

This paper cites Gen- erating images with multimodal language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Gen- erating images with multimodal language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.317157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.317157Z digest=sha256:c74de3f03b29c0135fc9b8ba77771fb6c1d702314b76edd5956e1cb6adbf7a87

Observation 1aa40df6-34c4-4764-9a09-ecfe61dc25a9 · outbound

This paper cites WikiHow: A Large Scale Text Summarization Dataset.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models WikiHow: A Large Scale Text Summarization Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.322049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.322049Z digest=sha256:0e3de341f9764c59ff47f3207e9eb680e61b41ce92d14fdfebc77533759b2d24

Observation 9bad06d2-7f42-44ae-96b8-95d227ef9d82 · outbound

This paper cites A survey on temporal sentence grounding in videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models A survey on temporal sentence grounding in videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.785086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.328096Z digest=sha256:663b476ff773900d6167433769a7da9ad3023e70f710849f31136936fcbbeae9

Observation 4f859fb8-ec1e-47e5-bdac-e4c9aa136e71 · outbound

This paper cites Tvr: A large-scale dataset for video-subtitle moment retrieval.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Tvr: A large-scale dataset for video-subtitle moment retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.767356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.333044Z digest=sha256:358abe77bea313d5da86e9f5c4dfb5c67fe6eb30db2b8519a6c8a299aeaa3af3

Observation b3dabebf-3f94-4b74-a5b4-6ec0448383ea · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.749524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.338050Z digest=sha256:3f4e1eb7643547130fbde825848652f1d5ff0be394fa74cc83400d9d9a3eccb0

Observation 596d4a41-5f75-49d6-be14-de21ac0339ef · outbound

This paper cites Chatting makes perfect: Chat-based image retrieval.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Chatting makes perfect: Chat-based image retrieval

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.343420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.343420Z digest=sha256:df4c3923208235de5e3e13d0785ac89fa447815b8034243628ddc9a5c4a79143

Observation c7e9112a-b8ab-44e2-92bf-84a31838eade · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.721183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.348350Z digest=sha256:0aa7c9070460511088f65668269ff43f0b1da94d91d56d45e94bea03defb95dc

Observation e8f42b6c-0d31-470d-8001-d8f90161d8e3 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.353346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.353346Z digest=sha256:b6ce1d9408b9c0d0ce73e9f61cf8d42464a546f632494364876294f5a7e814de

Observation 07b40416-83e2-4efc-bbd4-0a23843c2bbb · outbound

This paper cites Skip-plan: Procedure planning in instructional videos via condensed action space learning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Skip-plan: Procedure planning in instructional videos via condensed action space learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.690149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.358932Z digest=sha256:f24b1adf2af91c4fcc647f85138aef815fbea038d353840d53d628c9c2f37842

Observation a7f9a262-5fb1-419d-b595-0fca782911c5 · outbound

This paper cites Cohn, and Janet B.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Cohn, and Janet B

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.665546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.363704Z digest=sha256:2ac87f3c0fa2a22ae0954c590d4629795cb9f8f7bddf9ae78efd3da7fb777c41

Observation 40b457fb-a7cd-44a8-a55b-8708313d02b8 · outbound

This paper cites Learning to recognize procedural activities with distant supervision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning to recognize procedural activities with distant supervision

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.643121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.368912Z digest=sha256:3fd6d78e0f64c6ad4beae1623f8eea7bba794afeae5f8664374917bd88102a8f

Observation 4713edc8-2e9f-4a33-acec-590ae14c0c9d · outbound

This paper cites Jointly cross-and self-modal graph attention network for query-based moment localization.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Jointly cross-and self-modal graph attention network for query-based moment localization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.617330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.374456Z digest=sha256:70d05dfe04a472a24addd8bcfc2e5712bc752b72ca56cb36fc80c889eebffe1c

Observation d8bfad8e-0392-4876-af25-c9613c35bd01 · outbound

This paper cites Improved baselines with visual instruction tuning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Improved baselines with visual instruction tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.380143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.380143Z digest=sha256:3ca06069781f7e937b8b360917deb378ec260fc96c8cd67852ab07c2edbb09dc

Observation b8aec459-d859-4fbd-9f72-631649fb012d · outbound

This paper cites Visual instruction tuning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Visual instruction tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.385014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.385014Z digest=sha256:eb394aa75459143947eb6d9fb8dd6598a18303203eda92fd20240510977a0fc7

Observation 46da08f4-c521-47d8-8139-6a60145d8163 · outbound

This paper cites Attentive moment retrieval in videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Attentive moment retrieval in videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.569112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.390493Z digest=sha256:d5cc7caec7f8ddf3ab6051926273843e1c53646b567b67e9fcf53b4f2b4c653b

Observation 26decaed-decf-42f3-85d4-e172b0f4d7a6 · outbound

This paper cites Language models of code are few-shot commonsense learners.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language models of code are few-shot commonsense learners

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.545580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.395669Z digest=sha256:c6274c2affe7690b341e4e6064b04d28493fa655cc233c2ecad8f91e2c4d3b82

Observation ebe4190f-1994-4be1-b33b-a277349828bf · outbound

This paper cites What’s cookin’? interpreting cooking videos using text, speech and vision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models What’s cookin’? interpreting cooking videos using text, speech and vision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.517289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.400639Z digest=sha256:43ca06036020378c7702aad03746f97dd3e1a9414fb09f3b50b052c634d43fab

Observation 8abe8ae2-99f4-4a9e-8f58-fab46928dc7c · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.496609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.405575Z digest=sha256:60ec6a4600c8c8af4df9e120e2b89b66fe0e5e0835864fed38535ba4c93ad24e

Observation 0f3350a2-a0a3-4501-ad92-c2c0d25d5f24 · outbound

This paper cites End-to-end learn- ing of visual representations from uncurated instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models End-to-end learn- ing of visual representations from uncurated instructional videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.470230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.410996Z digest=sha256:3a4384dd6003d4fb82efb30eb867995f029f205de41708beb844e50cb40d0f5c

Observation ad815325-c0ba-41c6-a0dd-6595078478a7 · outbound

This paper cites Learning and Verification of Task Structure in Instructional Videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning and Verification of Task Structure in Instructional Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.415992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.415992Z digest=sha256:f6ddc330e1b18fe1c744e8ed0a81b96954e3ff1619ec28fd062fc6407051ea81

Observation ae29bc64-05ae-4610-b9f1-c961e70fb6e0 · outbound

This paper cites SCHEMA: State CHanges MAtter for procedure planning in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models SCHEMA: State CHanges MAtter for procedure planning in instructional videos

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.451365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.421653Z digest=sha256:35583093048d3d3c28fe8fdb3a97fc91819d3f5c342de85b52b35971143c78d7

Observation 46042b3c-b577-4cd8-a223-dec461577ac3 · outbound

This paper cites Virtualhome: Sim- ulating household activities via programs.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Virtualhome: Sim- ulating household activities via programs

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.430224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.427256Z digest=sha256:b0d4c0a9fac63b591555bc48dfc4e564a95456dd011bfa5fd6195266b4e17621

Observation 1d462551-7fc9-4673-824b-d3ef509b4255 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning transferable visual models from natural language supervi- sion

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.432722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.432722Z digest=sha256:e0ba703bb124c0e4185a963a62d4f2c11737f5ccce8b6a48000bb7d75d6604d3

Observation f28cd14f-2aa5-49f4-b2f8-1edd37df038a · outbound

This paper cites Grounding action descriptions in videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Grounding action descriptions in videos

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.398146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.437992Z digest=sha256:2931dd67e13091d33ffd92d729a66c1bd84b2c347c649a52c581a42a8ec3076a

Observation 8f268d0b-267a-4bcd-b7ac-7362cd2a945d · outbound

This paper cites FLAP: Flow-adhering planning with constrained decoding in LLMs.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models FLAP: Flow-adhering planning with constrained decoding in LLMs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.380232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.443152Z digest=sha256:37ce1691f6d8525660de69ac49256bc20f33b3ff1b1e333ed473342ce3ecf273

Observation 69608492-668a-46ee-a34e-8077fe24d1a5 · outbound

This paper cites proScript: Par- tially ordered scripts generation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models proScript: Par- tially ordered scripts generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.363004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.448309Z digest=sha256:7eee648a0f69cef6f6443c555b22bf32500ad3c3afa50e04b81d5ebe4671d13e

Observation 88d099c9-b89b-4ee9-b1a2-b817958e3567 · outbound

This paper cites As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.345705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.453462Z digest=sha256:0c1112f10f39ce1be0bf8ed3c8336f4edd772d64e5dc38792ee6f06464342d4b

Observation 374b58e8-0949-41aa-ae82-225144ebb97e · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.327594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.458789Z digest=sha256:73f15b68de33bc10397633fa99238dfe2a7aa7994eda3803b9ec5ef680452c48

Observation 5c00df61-fe98-42ae-aa1e-54de161d9b76 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.310703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.464824Z digest=sha256:321a036e7c256d7c3e57f7a29061db3c8590f4bdf17b96634ffccca3371ff655

Observation ebf6b249-9a9b-4921-b8c6-45eb5f078a1a · outbound

This paper cites Mpnet: Masked and permuted pre-training for language understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Mpnet: Masked and permuted pre-training for language understanding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.293438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.469567Z digest=sha256:9f2344abd6e7a0a298ce4eb5c23e705c5c946df5ff648f1006e769127d052b44

Observation c3a0b2e7-5820-41df-a9ef-2d276b799ea7 · outbound

This paper cites Language Models Can See: Plugging Visual Controls in Text Generation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Language Models Can See: Plugging Visual Controls in Text Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.474723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.474723Z digest=sha256:c4646196cabc65577b247861b1c3fa41d9546bb2e258cda35365682c36264e59

Observation 6019f744-9b5d-4a78-9201-ffe1265ca81f · outbound

This paper cites PandaGPT: One model to instruction-follow them all.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models PandaGPT: One model to instruction-follow them all

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.276653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.479891Z digest=sha256:d56d3ab32514ea2be9c810800c1111a92cbb6614706198822340c82e7fef2fee

Observation f9f1a374-9afe-4810-83be-9ba1448966c0 · outbound

This paper cites Plate: Visually-grounded planning with transformers in procedural tasks.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Plate: Visually-grounded planning with transformers in procedural tasks

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.258622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.484978Z digest=sha256:0c48aeb00434649f7cbc7cb0ecfe89f5b5c117820bb18eeee8d8037cff828165

Observation 76415b0a-33bb-4fbd-98bb-489517242b8a · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.239508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.489713Z digest=sha256:ad0eb55bd089b3da0ffccddb1f9a8cba42dbf2684a731b2a9d5fc2183f382b91

Observation a2a0d350-3c6b-4c88-ad2a-6c639902c601 · outbound

This paper cites On the planning abilities of large language models-a critical investigation.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models On the planning abilities of large language models-a critical investigation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.222261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.495158Z digest=sha256:b8a389fcfeb2ab458fc10ee0ceee17e2e3334d0e49431095a66696b4e2efa6d7

Observation ecf651f8-0f62-4594-85cc-a4041e7331b5 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.203062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.500019Z digest=sha256:4a76478bce2849b4a733439ac5134847615d006af231c812c3de96ee9a83b1be

Observation d7d46fa4-61ed-4a8f-8c20-036ce1a5e244 · outbound

This paper cites Event-guided procedure planning from in- structional videos with text supervision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Event-guided procedure planning from in- structional videos with text supervision

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.185574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.504726Z digest=sha256:d9bb80aa58db62c831c5c04661f6d3efd10a8d6eb16cf82e44bcd09266f36bac

Observation aa1704bf-798f-4e21-9317-a7a01628461f · outbound

This paper cites Pdpp: Projected diffusion for procedure planning in instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Pdpp: Projected diffusion for procedure planning in instructional videos

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.168578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.509417Z digest=sha256:fd7ba8d44609459ca80f6b552201e140c568d1c03532161e4d9e86cba31516c9

Observation 641a44df-1b47-42d3-b018-956b5133109d · outbound

This paper cites Temporal segment networks: Towards good practices for deep action recognition.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporal segment networks: Towards good practices for deep action recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.150365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.514360Z digest=sha256:235567d4acbbdebed0385b0cdcc062093e6988cba3a697945431739b0267d1e8

Observation 5155ce14-7170-4533-a1bc-32587b290c6d · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.519058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.519058Z digest=sha256:a458e805ca20ae8dde8c57a8fc1252c0ee95faed7367363ec488a5c83aee9500

Observation 9f82821f-3eaa-4d58-9597-5dec1e9624c7 · outbound

This paper cites NExt-GPT: Any-to-any multimodal LLM.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models NExt-GPT: Any-to-any multimodal LLM

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.133657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.524104Z digest=sha256:6e6d7697c0902aac063003f8bce0b354227ce1f701255eb8a961cf2bc7c65d53

Observation bfe34a2f-97a6-47ab-8583-2618443845c0 · outbound

This paper cites Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.116500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.528664Z digest=sha256:76df7cc4a4f8b35ed56e540c0e652945d308f759cf0d91b542c85c80783338b3

Observation b4329a38-a98a-4794-bac9-b97886dcf609 · outbound

This paper cites Translating Natural Language to Planning Goals with Large-Language Models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Translating Natural Language to Planning Goals with Large-Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:18.533727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:18.533727Z digest=sha256:5b39f766db5dfb091522f4465775aaa9614d11dde837b66cb48f894fc871d193

Observation d198cda2-5116-439d-ac92-f0bafaac0ec3 · outbound

This paper cites Multilevel language and vision integration for text-to-clip retrieval.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Multilevel language and vision integration for text-to-clip retrieval

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.097622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.539015Z digest=sha256:9653e9f0678f121db09d6c8bd7d286b5e73edb5eb2c6c2fa1ffb900f1955dd4e

Observation bb148e8c-adcf-4f31-b9d8-0bd1fa4bb170 · outbound

This paper cites VideoCLIP: Contrastive pre- training for zero-shot video-text understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models VideoCLIP: Contrastive pre- training for zero-shot video-text understanding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.079625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.544244Z digest=sha256:e6fb5945018860036e0159b3a04265d7122414e0736f777b9aac769c0d5d436a

Observation 6cc9b393-3050-4d91-a6f8-f455fe2de311 · outbound

This paper cites Retrieval- augmented generation with knowledge graphs for customer service question answering.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Retrieval- augmented generation with knowledge graphs for customer service question answering

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.060275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.549134Z digest=sha256:e0f27b880085609555f83c5766e21d23efb22ecc842d86a44f73ac4c3e1fad94

Observation f320f8fd-dd08-47a9-9add-799ba20f929d · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.043613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.554112Z digest=sha256:25db091702fed348454ce59ebca63bb4382bf9a54474a496639880fbc4fc83c5

Observation 4a9e3b04-e34b-4f3f-bd07-db480e1e6fe9 · outbound

This paper cites Distilling script knowledge from large language models for constrained language planning.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Distilling script knowledge from large language models for constrained language planning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.025535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.559096Z digest=sha256:1722a26e971a384d081df27e4812cc5d4cd0734e187ad4d0a28e0c94fd4b3eba

Observation 1e5d354d-f8e9-4e08-88f6-567a49c7bad2 · outbound

This paper cites Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:19.007476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.563941Z digest=sha256:43345bc24f075cccb7db6bcd6a1b5db6e608f564df3cdbc62b7b36875266df1d

Observation e47a14f2-1409-470a-b597-6372db2921c9 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversa- tional abilities.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models SpeechGPT: Empowering large language models with intrinsic cross-modal conversa- tional abilities

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.990834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.568729Z digest=sha256:4427dccb4a24e45e48d1ec2ae1e0b391db4568e940e3637708c80b185e766cea

Observation 65d682a9-ad24-4ef0-9aff-985af49ebc05 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.974322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.574278Z digest=sha256:de627a889c1a70fd05c7a4e0c00e9007e8e412173c9880fd621a4cff1684f3bb

Observation d628f934-164d-4b8b-a513-33d7a7b43b42 · outbound

This paper cites Temporal sentence grounding in videos: A survey and fu- ture directions.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Temporal sentence grounding in videos: A survey and fu- ture directions

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.957496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.579525Z digest=sha256:74f08a536c7b3649362ec190312f303d950ae59180a9c6d210fd0f2c7940c8c4

Observation 703bc398-48a0-48c7-9795-3cb775a7a0ae · outbound

This paper cites P3iv: Probabilistic procedure planning from instructional videos with weak su- pervision.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models P3iv: Probabilistic procedure planning from instructional videos with weak su- pervision

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.939426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.584443Z digest=sha256:2fef275e8b0cc6db263cc286ad3911c032e07f7dbd4f796d28753579627bcbfa

Observation 1bfea070-78ba-47be-a292-a0398e67453e · outbound

This paper cites Learning procedure-aware video repre- sentation from instructional videos and their narrations.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Learning procedure-aware video repre- sentation from instructional videos and their narrations

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.921158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.590544Z digest=sha256:9034756285809691bb526f2b21f499ad450e41df76df13ae97d1f9ac75ee1094

Observation cd6107ab-a413-41f7-ad15-d3bcb5ac5f94 · outbound

This paper cites Procedure-aware pretraining for instructional video understanding.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Procedure-aware pretraining for instructional video understanding

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.904145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.596552Z digest=sha256:f9fb6262f8389f10f8b4e7c147c62d86d9c821d2086caf23f6bfba8db66a2502

Observation 7901844d-b6bf-424c-b788-b3bd49dd1263 · outbound

This paper cites Towards auto- matic learning of procedures from web instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Towards auto- matic learning of procedures from web instructional videos

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.887341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.602763Z digest=sha256:181f07f95bafb251fbcacf61be221729df36313b2a923db3cf710906fc3cd123

Observation d130331b-ff3b-4241-903e-a583413649ee · outbound

This paper cites MiniGPT-4: Enhancing vision-language understanding with advanced large language models.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models MiniGPT-4: Enhancing vision-language understanding with advanced large language models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.869455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.608913Z digest=sha256:13247e807d0ee220b1bff7c2241818c306dcc4778d1b00e8a3a84c43fbca7f6c

Observation f8de0d6b-998d-4e0c-bf16-2b794b0714b4 · outbound

This paper cites Cross- task weakly supervised learning from instructional videos.

InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models Cross- task weakly supervised learning from instructional videos

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:27:18.851520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T17:27:18.614377Z digest=sha256:714be3d532751ad925cf9868b1e8dd7c2082a73a2314cb4a4680fc2d655fb5ee

Pith citing papers

Observation 1a817673-2a56-4d5d-916d-4a9b06e41a2b · inbound

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting cites this paper.

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:54:18.278334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T17:53:19.002603Z digest=sha256:ca118445b47030e6248af1a6192e8eca1d3bb12b0173915a70cdd749c9d40dbe

Observation 94abaaac-ea99-486a-89b4-b240171202ed · inbound

VisionClaw: Always-On AI Agents through Smart Glasses cites this paper.

VisionClaw: Always-On AI Agents through Smart Glasses InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:13:05.498964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T18:12:34.183092Z digest=sha256:d868ae45f1f1b3dc2136c9b4cb1bdbd0102c5e1386683eb629ba2e4b9dfe4db2

Observation 3b18f4b1-b9d9-41b9-9fb7-55667d3f9008 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models

Reference 256

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.163371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:a681bdc5a83e1cbd3727a3e35ce1ec23f47f51f58f907bf260ba42bbafa28b08