Pith. sign in

Paper Citation Record · LEDGER

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

As of 12 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 7 inbound Pith citation observations for arXiv:2411.14505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14505 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:46:34.180369Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:32:52.859271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:14.988080Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29aea9b4-9f57-4e3d-9a78-1ba0dd04daee · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:33.987486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:33.987486Z digest=sha256:4468ee2dbdbc18767abe918fe452d6dec8c70802ee1a61e1737dbfda9f668cf0

Observation 5e137d13-e512-4518-8e9e-ad2bc1e43c5e · outbound

This paper cites Localizing mo- ments in video with natural language.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Localizing mo- ments in video with natural language

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.826376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:33.992151Z digest=sha256:39fb93bdf9436a5d3bc434d4235e5646999cdad8259478d49482cb80fe53aa85

Observation 74001564-cdce-4479-9e12-72b00375fe3b · outbound

This paper cites Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.815024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:33.995449Z digest=sha256:a51408e160a859cfd490391c899aa16c9895442b6c49d000523815de6d6bfaeb

Observation df9a1c82-2d1e-4c25-bfc5-f7fdc5c9745f · outbound

This paper cites CTRN: Class-Temporal Relational Network for Action Detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval CTRN: Class-Temporal Relational Network for Action Detection

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:46:34.294023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:33.999034Z digest=sha256:3fc97b74d12e2ddd469fbeaecfe7a4f700f0458150c7545d527dd4d1d3a1f7e3

Observation 49dd12b1-cddf-4b10-8a3f-e4b8a4f74c17 · outbound

This paper cites Ms-tct: Multi-scale temporal con- vtransformer for action detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Ms-tct: Multi-scale temporal con- vtransformer for action detection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.803597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.003000Z digest=sha256:4ef978f8904d61a9944a392a7f2f4e5714fc16cfac20ada4446748642cefd431

Observation 52de721a-5de5-496d-8f39-41286100187e · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.006592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.006592Z digest=sha256:7c0128fb08560a1f0aceadf5349d812ced831a73a553c95fa21bb03a0fa2c320

Observation 30268258-4e5a-45fa-8a3a-6bb5ef239cde · outbound

This paper cites Tall: Temporal activity localization via language query.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Tall: Temporal activity localization via language query

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.784282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.010324Z digest=sha256:742f02aa0edbe5304bfa1d59e0db6eb580c77889637885b90fd42c88c1c11c08

Observation 0b4df684-294d-4393-afa5-0363afeb3d15 · outbound

This paper cites Video action transformer network.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Video action transformer network

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.772785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.014040Z digest=sha256:67eb8b35aa434462e7165b37892a02058c4a9d5f99e1f8349c56b84b4ab36b77

Observation 9aedc72a-7358-42f7-bd3f-c848bd9d4586 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval LoRA: Low-Rank Adaptation of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.017653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.017653Z digest=sha256:c558e8e1e58b8bc83ef87e4da628e4092fda09b220852800c0ff545f6ffa5524

Observation 557fffbd-95b8-4070-9a83-357119e67536 · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Knowing where to focus: Event-aware transformer for video grounding, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.761744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.021559Z digest=sha256:e2001d278653ca3ad450a5bc94317839db5cf3db9c12279ce984b6c25c48b541

Observation 3488fcf5-b860-4968-b767-a0b37bab20db · outbound

This paper cites Efficient multimodal large language models: A survey.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Efficient multimodal large language models: A survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.025371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.025371Z digest=sha256:7cd36f2d12c4be9c42d7476b14b1d92bbcedd700c1f4aa4e891cb5b0b2700f7d

Observation 2adcebb4-0b78-45eb-a74f-b2fb462a1e8c · outbound

This paper cites Dense-captioning events in videos,.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Dense-captioning events in videos,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.029064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.029064Z digest=sha256:0bd1a8940188382c75c49d935c0768f985f1fd48277d383d1640953681e9e479

Observation 40028e2e-2684-48b2-aea1-10f5e0b6ea3d · outbound

This paper cites Temporal convolutional networks for ac- tion segmentation and detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Temporal convolutional networks for ac- tion segmentation and detection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.743835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.032763Z digest=sha256:3047b166e728ee4a06e0b424074212412eb6e77a4a7dcf4144d00f04ebf3cee1

Observation 1351cd3b-6637-480e-94f8-918d8740b673 · outbound

This paper cites Berg, and Mohit Bansal.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Berg, and Mohit Bansal

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.732984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.036780Z digest=sha256:4f7a6dbe3d84473d20ade2a8bdffb15790073c7a5e1bf54278308afc866a944c

Observation caaaa026-fd5e-4e66-b865-ec6984da59e1 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Detecting mo- ments and highlights in videos via natural language queries

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.721157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.040363Z digest=sha256:5546551de78ee347c91ba4ee60fa0d6a3a0f2358a5f6a8506a0766373363dee7

Observation 7d3f1942-2f5b-44ea-a7a4-3c544590db18 · outbound

This paper cites Mimic-it: Multi-modal in-context instruction tuning, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Mimic-it: Multi-modal in-context instruction tuning, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.043808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.043808Z digest=sha256:e43b4bf89b086d4af108004c12b21bcc19f3cd50de2c2bd5a4a68b23b7963c38

Observation 4fb2af8c-fcc8-4136-9ccf-17de71e758c7 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.047579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.047579Z digest=sha256:76066f6e0f916957e5e9e657ad986b339e60fad8e7eee8bb1fcc35a38e79480a

Observation 48d2f88f-bc4d-4379-bd57-59a0eb3a210b · outbound

This paper cites A survey on benchmarks of multimodal large language models,.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval A survey on benchmarks of multimodal large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.695108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.051139Z digest=sha256:4461336174690ee55b34aefc779e470b9fe28563df1502281a860a73d12f701d

Observation 7bc2fa66-4b8e-400c-b54a-8c4a5a35b30a · outbound

This paper cites Videochat: Chat-centric video understanding, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Videochat: Chat-centric video understanding, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.054794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.054794Z digest=sha256:844713eeb2cc2c241d1119f8d552aaa073718db130a99b2b835c10d2907f7de5

Observation 5d95690e-881a-43eb-8591-e22017a0498d · outbound

This paper cites Fast learning of temporal action proposal via dense boundary generator.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Fast learning of temporal action proposal via dense boundary generator

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.676050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.058301Z digest=sha256:96600c933040f5c96414fce37f10dc31fdb499bf59307b92fca8f44b4821996d

Observation 9c8b8761-daf7-43ea-880d-fa85cf37a993 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Univtg: Towards unified video- language temporal grounding, 2023

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.664660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.061972Z digest=sha256:97516f115d2f0647a85a11c793a12a32244e0ed7c8681555601c58d9d47b71ca

Observation 71aee2b2-1ce9-4d9b-901d-e25aa33cfbf3 · outbound

This paper cites Bsn: Boundary sensitive network for temporal action proposal generation.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Bsn: Boundary sensitive network for temporal action proposal generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.654031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.066083Z digest=sha256:e0be2b00fe75f5c979c693001a30ee34fcbf51fd947f16ab5dce9168a1e3862c

Observation 8c4996de-1346-4b70-a7d6-06e9f565ddd7 · outbound

This paper cites Visual instruction tuning, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Visual instruction tuning, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.069591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.069591Z digest=sha256:96b47941b48a0de4d5b1e9acb2a2ab11289c14f1d3b7c1fc0dde26879737bf57

Observation 1174f455-31c5-4ed0-948c-9c7b642a567f · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.635444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.072992Z digest=sha256:5fa668f5f2ed48fc6e027876923df7041b9c7dce4862b655c49b38a572dc4c4e

Observation 71cbe1a5-0956-4492-b4c4-7d6b114981f4 · outbound

This paper cites Decoupled weight decay regularization, 2019.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Decoupled weight decay regularization, 2019

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.076530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.076530Z digest=sha256:3d140ef0bd4ef80e40522d53d4384b9c3915f5426fbaea46af58566393605da0

Observation 246e3dfe-457e-4f15-ae4a-7c4edb6b7597 · outbound

This paper cites Valley: Video assistant with large language model enhanced ability, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Valley: Video assistant with large language model enhanced ability, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.079946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.079946Z digest=sha256:e411f9cd1e420a2120bd2cee77dc046a841a4f665c8c0a1eda3253fb9cbc282d

Observation 19f0d48a-eb54-4200-943a-d996aed2b9e5 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Video-chatgpt: Towards detailed video understanding via large vision and language models, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.609308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.083458Z digest=sha256:81277d6b52661a6d72d9fd253a28a3bf46768dadda98185ec71ec25db7dde8d5

Observation 53832ca4-6755-4aa6-bf0a-05ee22e796cf · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval The surprising effectiveness of multimodal large language models for video moment retrieval, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.598219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.086941Z digest=sha256:b3132f2a994e5de7a1c81eeb50aeec2709717be800cf4a5ddea89e6ef8e7ea73

Observation b2bb8faa-3cb8-4b76-9c4a-eb24a7479be3 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Query-dependent video representa- tion for moment retrieval and highlight detection, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.586889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.090381Z digest=sha256:a368f08df296d2acc4c56a676af0b93269aad60d15b7240861eb074e10c72eb9

Observation fd5efcf6-9441-4802-8a47-4cacf6a24796 · outbound

This paper cites Correlation-guided query-dependency calibration for video temporal grounding, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Correlation-guided query-dependency calibration for video temporal grounding, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.575592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.094038Z digest=sha256:0d60c606abbb43e3bdbd5df96d77a3f63ca370a8ae4e7da39c0f4c4ab9c0efba

Observation 8770941d-1da9-4a76-9b6a-9742bc443b62 · outbound

This paper cites Local- global video-text interactions for temporal grounding.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Local- global video-text interactions for temporal grounding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.564426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.097533Z digest=sha256:fe81a64757ccf045d95a1a47ca601d576fef56beeba442b5d34c46f9ceff88e2

Observation 04bc2f47-7523-40fb-8322-ba84ceab1104 · outbound

This paper cites Pat: Position-aware transformer for dense multi-label action detection.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Pat: Position-aware transformer for dense multi-label action detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.553015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.101057Z digest=sha256:c918e7aa7eee8cb26a22ad64b73e118a7cecb3ae11cdfc1241850c001c384921

Observation 0fd47542-75b4-408d-8759-a2762961bf45 · outbound

This paper cites Temporal action localization in untrimmed videos via multi-stage cnns.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Temporal action localization in untrimmed videos via multi-stage cnns

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.541275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.104973Z digest=sha256:f651c2fe0534081992648dd0503c0d299c6db07429f938fc646f0eeea14f196a

Observation 06cf80aa-567e-4be8-b43b-a2f20b646461 · outbound

This paper cites Vlg-net: Video-language graph matching network for video grounding, 2021.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Vlg-net: Video-language graph matching network for video grounding, 2021

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.530151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.108323Z digest=sha256:037070bbb9ffbe698c0ef00302f27ead344036e2e991fe4da30bd5cfd1bf9000

Observation 2c06cbe0-62ce-4cf3-9857-2154ddbc99f1 · outbound

This paper cites Learning grounded vision-language representation for versatile understanding in untrimmed videos, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Learning grounded vision-language representation for versatile understanding in untrimmed videos, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.519074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.111758Z digest=sha256:8a3d0ccf725d3480ba7bc8549367d7b47663cf656035089ccf40d232a312e706

Observation a557d774-50bd-472c-a27d-9ebbf57e4a13 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Internvideo2: Scaling foundation models for multimodal video understanding, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.508781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.115164Z digest=sha256:afbb055648aadc2d044cf37acd0ba4781c27acf4c03b8c20f03768711825c3ef

Observation 70727b54-4d97-42c4-8ec7-9bdd134efb16 · outbound

This paper cites Unloc: A unified framework for video localization tasks, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unloc: A unified framework for video localization tasks, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.498670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.118300Z digest=sha256:59abbc321fa578c7f8b4c9b7214f16324bdc79072fbfcd534f52607f4cf69096

Observation b175bed0-7f32-4ac1-836d-2ab0537235bd · outbound

This paper cites Unloc: A unified framework for video localization tasks, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unloc: A unified framework for video localization tasks, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.489049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.121669Z digest=sha256:2b3689781cea36497a88c9dc823b56b911e5fa33469b6112ed5b2048cd95fac5

Observation 99495fa3-ae96-4f4d-b2a3-a1b32e1025b3 · outbound

This paper cites Self-chained image-language model for video localization and question answering, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Self-chained image-language model for video localization and question answering, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.478605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.124891Z digest=sha256:f7f21778d650de30f2140e50477d8cd78109d5869800400f2162ec4275598727

Observation 53683c38-2ee5-42a1-8334-a5ccb8a990e7 · outbound

This paper cites Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.466924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.128070Z digest=sha256:7a9879ff85faf09f296b9b44144b8ba41fc8fc697c42438f0e77bbf21e2e4ea8

Observation 3c4a4bbc-5325-4133-a3cd-864e8e3d07e4 · outbound

This paper cites Graph con- volutional networks for temporal action localization.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Graph con- volutional networks for temporal action localization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.454748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.131043Z digest=sha256:6454db9b0477cb27b10496eaa12e77251cc5fd98ac557ab7e37b4826dd07f14e

Observation abe0109e-633f-4829-b1b3-f1ac251df33b · outbound

This paper cites Dense regression network for video grounding, 2020.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Dense regression network for video grounding, 2020

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.443580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.134371Z digest=sha256:c2f0227fec69f5e7ee9df1d7b39a8919170028b672d313fdfce813df740d7574

Observation 16ead2ca-91f5-4d79-b4ce-19ee17466dc8 · outbound

This paper cites Unimd: Towards unifying moment retrieval and temporal ac- tion detection, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unimd: Towards unifying moment retrieval and temporal ac- tion detection, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.432557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.137745Z digest=sha256:605b6d1e1caea4cbc92a760ab6f82bf5aca0a8c9b547d9a6d1a603a31fc21814

Observation 6ffe070f-1deb-4aa9-bb1e-af6a2e1b46bf · outbound

This paper cites Actionformer: Localizing moments of actions with transformers.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Actionformer: Localizing moments of actions with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.420874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.141472Z digest=sha256:161a8dada9d4ad6b1831d69d009ec646c8d324d17b696c73b42ce57d2d646f35

Observation 88ccc7a2-1faf-4697-a385-08e9d72ead34 · outbound

This paper cites Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.409042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.145028Z digest=sha256:3cb106e158a0af02998f208ad38d9288a84e1b6f2b4e610fbd800641a3128b88

Observation fb139d1d-1b80-4db8-8346-44fa833f7abb · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Video-llama: An instruction-tuned audio-visual language model for video un- derstanding, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.148600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.148600Z digest=sha256:cd906a6d8499ff95c4bc2a51371e765f7df7ed6fcb029861945fdcfbdabb8a87

Observation d70a4ef0-1d32-41b6-810f-abe85ea0bf88 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of language models with zero-init attention, 2024.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Llama-adapter: Efficient fine-tuning of language models with zero-init attention, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.390918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.151866Z digest=sha256:9d7c2402133a6ace87114e012ec0df9a9bec2d65dfa37815a603f374c6e99f6c

Observation 30d3c546-e572-4a63-9f0f-244e97c6613a · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:34.155307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:34.155307Z digest=sha256:665f6efe640ee4b1bf763f928409f46be459b4d5bfab58672b23be25f4c52085

Observation 0ce4def2-55cd-43a5-945e-ae64f5ff87d1 · outbound

This paper cites Temporal action detection with structured segment networks.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Temporal action detection with structured segment networks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.372984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.158791Z digest=sha256:a3eb7a645faaa14e2adaf342ce521d9725f796aa04328e13274495c22487131d

Observation 8b0f7f74-0fa9-4704-9eb2-de4b3c8e6e6d · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.361880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.162165Z digest=sha256:2b03623ba903997cb2b22e99289a14577eeb033ad35e84561cdb9c10664841cb

Observation f7d3ba55-d9ba-42ae-ad0c-cea89989a227 · outbound

This paper cites Enriching local and global contexts for temporal action localization.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Enriching local and global contexts for temporal action localization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.350436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.165559Z digest=sha256:bbf52e54cbe00b3b88aaeec6d35afef0eb83748c99856dff9d62091e0c08a1ac

Observation a8e99c9d-6d27-43a4-8ffe-5dd551089ade · outbound

This paper cites an unresolved cited work.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:46:34.339358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.168963Z digest=sha256:a38622751166496f2aa9e849115bff4883f45b34e8e833c4d76ff1bc34bcd816

Observation 5eab3430-918f-4c82-ae81-aed16842eb78 · outbound

This paper cites [[-1, -1]].

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval [[-1, -1]]

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.328503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.172906Z digest=sha256:c8ce88389baa86ee73c917bc0d53417521c599b13ae40a7d561d142a81f219cd

Observation 7c2b5e37-9af4-4a0b-9e1d-3aae182d0e2f · outbound

This paper cites automated devices operating in a modern factory.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval automated devices operating in a modern factory

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.317235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.176858Z digest=sha256:a0f25da99386faf3d51811dd3f88b083b009d912d057e19b9118dd1c2da77589

Observation 8500c787-c420-47b0-81e9-08b5dc4e9d88 · outbound

This paper cites This section explores po- tential future directions for enhancing the performance of MLLMs in moment retrieval tasks.

LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval This section explores po- tential future directions for enhancing the performance of MLLMs in moment retrieval tasks

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:46:34.305958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:46:34.180369Z digest=sha256:ff4fa2426d66b85140dbe2af865ee778c6d8397109aa2b4949c40be5ef6cd6ac

Pith citing papers

Observation a8c5eb12-65ec-487d-8628-b0f2bed9b148 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.859271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.859271Z digest=sha256:cffffebee92f3a548efc231f34df6a0b5512e0bf1cdacb36ef759cfebb2a7761

Observation 34b3b1a3-7981-425a-a102-7663f633fb81 · inbound

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding cites this paper.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.771931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.771931Z digest=sha256:46b757089ecfd42f58d3e9ab56cd5fbcdf3481a963ee433a577038c010c4880d

Observation e4436d2b-52f0-4d91-9374-14e51158659b · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.938393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.938393Z digest=sha256:7a13d51a208b18c0ddce99003a7439b02a1768826cabb38fbfb1450327622064

Observation 66daadb5-120d-4ba9-985d-3d3e1ae18f45 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:45.907432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:45.907432Z digest=sha256:3582f12e4765d0013982590063f4ae9ca642c2379924d67bf8b878e16326c102

Observation 4b3d4251-7538-4694-85e7-b6fa0d1c2d9c · inbound

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos cites this paper.

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:30:58.574381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T17:37:40.373211Z digest=sha256:c5c3a5d901eaa18b4d548cc00cada5d69cc7bdbbaf535fdfa956b126fb709c47

Observation 905d3179-9bae-429f-9fb5-bafc0ad4fb4e · inbound

Towards One-to-Many Temporal Grounding cites this paper.

Towards One-to-Many Temporal Grounding LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.671960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T02:11:48.455492Z digest=sha256:2a93dfff5c103db75ca87247f82f27f01335197db6673f5b4cffcdfb02b594f6

Observation 0b2b1842-8324-47dd-9e39-7e9d9215ef89 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:14.990641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:445d24eefaa6c6200c22cbbd9f9b35f54450575afae6fd576bbf3e013baf4fef