Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:46:34.180369Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 7 inbound Pith citation observations for arXiv:2411.14505.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:46:34.180369Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:52.859271Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T17:27:14.988080Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 29aea9b4-9f57-4e3d-9a78-1ba0dd04daee · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e137d13-e512-4518-8e9e-ad2bc1e43c5e · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Localizing mo- ments in video with natural language
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74001564-cdce-4479-9e12-72b00375fe3b · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els, 2023
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation df9a1c82-2d1e-4c25-bfc5-f7fdc5c9745f · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval CTRN: Class-Temporal Relational Network for Action Detection
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 49dd12b1-cddf-4b10-8a3f-e4b8a4f74c17 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Ms-tct: Multi-scale temporal con- vtransformer for action detection
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 52de721a-5de5-496d-8f39-41286100187e · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Instructblip: Towards general- purpose vision-language models with instruction tuning,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30268258-4e5a-45fa-8a3a-6bb5ef239cde · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Tall: Temporal activity localization via language query
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0b4df684-294d-4393-afa5-0363afeb3d15 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Video action transformer network
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9aedc72a-7358-42f7-bd3f-c848bd9d4586 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval LoRA: Low-Rank Adaptation of Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 557fffbd-95b8-4070-9a83-357119e67536 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Knowing where to focus: Event-aware transformer for video grounding, 2023
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3488fcf5-b860-4968-b767-a0b37bab20db · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Efficient multimodal large language models: A survey
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2adcebb4-0b78-45eb-a74f-b2fb462a1e8c · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Dense-captioning events in videos,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40028e2e-2684-48b2-aea1-10f5e0b6ea3d · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Temporal convolutional networks for ac- tion segmentation and detection
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1351cd3b-6637-480e-94f8-918d8740b673 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Berg, and Mohit Bansal
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation caaaa026-fd5e-4e66-b865-ec6984da59e1 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Detecting mo- ments and highlights in videos via natural language queries
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7d3f1942-2f5b-44ea-a7a4-3c544590db18 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Mimic-it: Multi-modal in-context instruction tuning, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fb2af8c-fcc8-4136-9ccf-17de71e758c7 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d2f88f-bc4d-4379-bd57-59a0eb3a210b · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval A survey on benchmarks of multimodal large language models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7bc2fa66-4b8e-400c-b54a-8c4a5a35b30a · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Videochat: Chat-centric video understanding, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d95690e-881a-43eb-8591-e22017a0498d · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Fast learning of temporal action proposal via dense boundary generator
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c8b8761-daf7-43ea-880d-fa85cf37a993 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Univtg: Towards unified video- language temporal grounding, 2023
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71aee2b2-1ce9-4d9b-901d-e25aa33cfbf3 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Bsn: Boundary sensitive network for temporal action proposal generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c4996de-1346-4b70-a7d6-06e9f565ddd7 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Visual instruction tuning, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1174f455-31c5-4ed0-948c-9c7b642a567f · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71cbe1a5-0956-4492-b4c4-7d6b114981f4 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Decoupled weight decay regularization, 2019
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246e3dfe-457e-4f15-ae4a-7c4edb6b7597 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Valley: Video assistant with large language model enhanced ability, 2023
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19f0d48a-eb54-4200-943a-d996aed2b9e5 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 53832ca4-6755-4aa6-bf0a-05ee22e796cf · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval The surprising effectiveness of multimodal large language models for video moment retrieval, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b2bb8faa-3cb8-4b76-9c4a-eb24a7479be3 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Query-dependent video representa- tion for moment retrieval and highlight detection, 2023
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fd5efcf6-9441-4802-8a47-4cacf6a24796 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Correlation-guided query-dependency calibration for video temporal grounding, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8770941d-1da9-4a76-9b6a-9742bc443b62 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Local- global video-text interactions for temporal grounding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 04bc2f47-7523-40fb-8322-ba84ceab1104 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Pat: Position-aware transformer for dense multi-label action detection
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0fd47542-75b4-408d-8759-a2762961bf45 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Temporal action localization in untrimmed videos via multi-stage cnns
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06cf80aa-567e-4be8-b43b-a2f20b646461 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Vlg-net: Video-language graph matching network for video grounding, 2021
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2c06cbe0-62ce-4cf3-9857-2154ddbc99f1 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Learning grounded vision-language representation for versatile understanding in untrimmed videos, 2023
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a557d774-50bd-472c-a27d-9ebbf57e4a13 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Internvideo2: Scaling foundation models for multimodal video understanding, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70727b54-4d97-42c4-8ec7-9bdd134efb16 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unloc: A unified framework for video localization tasks, 2023
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b175bed0-7f32-4ac1-836d-2ab0537235bd · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unloc: A unified framework for video localization tasks, 2023
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 99495fa3-ae96-4f4d-b2a3-a1b32e1025b3 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Self-chained image-language model for video localization and question answering, 2023
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 53683c38-2ee5-42a1-8334-a5ccb8a990e7 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3c4a4bbc-5325-4133-a3cd-864e8e3d07e4 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Graph con- volutional networks for temporal action localization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation abe0109e-633f-4829-b1b3-f1ac251df33b · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Dense regression network for video grounding, 2020
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16ead2ca-91f5-4d79-b4ce-19ee17466dc8 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unimd: Towards unifying moment retrieval and temporal ac- tion detection, 2024
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ffe070f-1deb-4aa9-bb1e-af6a2e1b46bf · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Actionformer: Localizing moments of actions with transformers
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 88ccc7a2-1faf-4697-a385-08e9d72ead34 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fb139d1d-1b80-4db8-8346-44fa833f7abb · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Video-llama: An instruction-tuned audio-visual language model for video un- derstanding, 2023
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d70a4ef0-1d32-41b6-810f-abe85ea0bf88 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Llama-adapter: Efficient fine-tuning of language models with zero-init attention, 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 30d3c546-e572-4a63-9f0f-244e97c6613a · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Learning 2d temporal adjacent networks for moment local- ization with natural language
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce4def2-55cd-43a5-945e-ae64f5ff87d1 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Temporal action detection with structured segment networks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8b0f7f74-0fa9-4704-9eb2-de4b3c8e6e6d · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f7d3ba55-d9ba-42ae-ad0c-cea89989a227 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Enriching local and global contexts for temporal action localization
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a8e99c9d-6d27-43a4-8ffe-5dd551089ade · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5eab3430-918f-4c82-ae81-aed16842eb78 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval [[-1, -1]]
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7c2b5e37-9af4-4a0b-9e1d-3aae182d0e2f · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval automated devices operating in a modern factory
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8500c787-c420-47b0-81e9-08b5dc4e9d88 · outbound
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval This section explores po- tential future directions for enhancing the performance of MLLMs in moment retrieval tasks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a8c5eb12-65ec-487d-8628-b0f2bed9b148 · inbound
DisTime: Distribution-based Time Representation for Video Large Language Models LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b3b1a3-7981-425a-a102-7663f633fb81 · inbound
Sparse-Dense Side-Tuner for efficient Video Temporal Grounding LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4436d2b-52f0-4d91-9374-14e51158659b · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66daadb5-120d-4ba9-985d-3d3e1ae18f45 · inbound
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b3d4251-7538-4694-85e7-b6fa0d1c2d9c · inbound
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 905d3179-9bae-429f-9fb5-bafc0ad4fb4e · inbound
Towards One-to-Many Temporal Grounding LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0b2b1842-8324-47dd-9e39-7e9d9215ef89 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.