Pith. sign in

Paper Citation Record · LEDGER

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2411.14901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14901 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:51:17.742891Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:32:17.962997Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T23:32:18.789296Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy37
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6de1641-e8d3-49f1-9180-66134e2807f3 · outbound

This paper cites needle in a haystack.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos needle in a haystack

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.633581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.480428Z digest=sha256:06f021236586e45476633909c3ba84b5021bb63d79eead2182ed9771336a640b

Observation 14f8428d-d7b1-4c95-9845-40a18e4758a1 · outbound

This paper cites needle in a haystack.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos needle in a haystack

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.485527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.485527Z digest=sha256:fa3656d74a0ca37e1498a725804a34b44186cedc8d645687928a45b7d634fcfb

Observation 932cbd73-d287-489b-86e0-771313651af9 · outbound

This paper cites https://sharegpt.com/.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos https://sharegpt.com/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.621707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.489695Z digest=sha256:afddfe4bed6653f6f9ec561b5a521353e3da168c13ed8629abe2f5c20024a277

Observation 2494fcf0-7316-44fd-8728-03845a06f407 · outbound

This paper cites https://github.com/gkamradt/LLMTest_ NeedleInAHaystack.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos https://github.com/gkamradt/LLMTest_ NeedleInAHaystack

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.610355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.493940Z digest=sha256:b6c199cf055f72c2ff5822a424db1ec021318e769f7b76852ee01d63228818b2

Observation 514194ae-6741-4e06-b166-858bb214e022 · outbound

This paper cites Lo- calizing moments in long video via multimodal guidance.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Lo- calizing moments in long video via multimodal guidance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.598205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.497775Z digest=sha256:dc21866ab713604793421b70aac39a7dd88331b500dd48f5373369c187b31053

Observation 70c6f739-fbc3-4dfb-9a5b-18bf90ef7e20 · outbound

This paper cites Functional brain organization of preparatory attentional control in visual search.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Functional brain organization of preparatory attentional control in visual search

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.586723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.501746Z digest=sha256:369b0f478dd8b3a6644940d168cd298707ad5cc2a078177159520c84e35055e4

Observation c3a58dae-6d8d-4df8-9152-0553858fa38f · outbound

This paper cites End-to- end object detection with transformers.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos End-to- end object detection with transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.505724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.505724Z digest=sha256:44e4a705ff71ce82740111e8ad6d11d47635c46be25c8cb9904ba5ee8d52153d

Observation f6a37a7e-6852-4567-ba07-10d991286069 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos VideoLLM: Modeling Video Sequence with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.509494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.509494Z digest=sha256:e965795b8e0b6ed01e8ea6294d9839ca293b9de5729173722efea779652d4d5f

Observation b8ee4b52-d852-41af-b32a-ee00329f0aad · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.567451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.513431Z digest=sha256:bfabbd5cf37d367c690d627f16c10a982b79691417d306fc2d350582be13214a

Observation 8425f23a-9f25-42db-b812-904a74ca0259 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.556051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.516640Z digest=sha256:b3965cf08bc3bf8a7904f934a6aa12ee0eb71448ce811e53cb69a16b8b19439e

Observation 7bfc6ce1-d9db-4510-b89a-3b9e40b894c4 · outbound

This paper cites Uatvr: Uncertainty-adaptive text-video retrieval,.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Uatvr: Uncertainty-adaptive text-video retrieval,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.543891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.519848Z digest=sha256:c673a4703a9632c89dca95744f89ad84a41bc77b308e9e6ede496a0ca7dff681

Observation b54cee9a-901a-4e69-b862-fad754abe3e1 · outbound

This paper cites Multi-modal transformer for video retrieval.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Multi-modal transformer for video retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.531878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.523083Z digest=sha256:c959e12301032c502f3431888535c9f957b0113589735fc15b673dbad0d4eebb

Observation d59df371-f9dc-4faf-9877-95bd9c219341 · outbound

This paper cites AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.526246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.526246Z digest=sha256:e158da6bffbf41c4817c5682b7ed07248e95f8ea1897fb6c8e3aa8ea100a6008

Observation c23fb26c-f5c4-464b-8e35-9534ba353132 · outbound

This paper cites X-pool: Cross-modal language-video attention for text- video retrieval, 2022.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos X-pool: Cross-modal language-video attention for text- video retrieval, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.520335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.529565Z digest=sha256:6bdf9a79dbdb4693e1ec3207dcacab59883bb518f228319172b695bff741ccad

Observation bc23c703-d543-4919-8f23-5065e7fcfed0 · outbound

This paper cites RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:51:18.046182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.532531Z digest=sha256:a168e4cab9b80c0b4d087621501dae4afff547d2cada8af5ed5c0cbe6e924e60

Observation 255d90f0-9d27-4813-ad21-2f1b032353c9 · outbound

This paper cites CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.536220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.536220Z digest=sha256:a8fd05a772ccd7e1bd6a0e342ffa6673df8bf6ba89484eb89b9507537f272f99

Observation 4395324c-3dd4-4766-adfa-b8ea1ccf4037 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.539848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.539848Z digest=sha256:3082fc80bf99d1ace1ff39b1c5e5d43d03d322935ab385e07a2dfe1e48d5998a

Observation 857ffaaa-1626-4923-8342-37e56075e9d4 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vtimellm: Empower llm to grasp video moments

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.509395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.543649Z digest=sha256:12fba7b98fb5f0a8562e00d8cbae09292a8ef2b52449b046b718595c8dfcbd0b

Observation 5760acec-746b-45f6-bfbc-d5a67ef8412f · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LITA: Language Instructed Temporal-Localization Assistant

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.547199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.547199Z digest=sha256:b1f2dc40846a5c702674505472b760a1d9497f4d5feba5ead76c25b67dad5317

Observation 6e092eaf-b30b-4139-ac67-d61fcb73725d · outbound

This paper cites Audio- enhanced text-to-video retrieval using text-conditioned fea- ture alignment, 2023.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Audio- enhanced text-to-video retrieval using text-conditioned fea- ture alignment, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.498330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.551108Z digest=sha256:72d0752efc18a9d3ec31e8325533bf3cc500f8ef5056f1b0dc719319f29212d2

Observation 5fdfe32c-e0ae-4d09-97d9-3009ec7483f9 · outbound

This paper cites Efficient long- text understanding with short-text models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Efficient long- text understanding with short-text models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.487239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.554425Z digest=sha256:54430d058a0b8e8b15ed1010781883a848798edbcc22437bf6e8a41094670a7c

Observation 859ddc87-c4f6-4c55-a492-ae3650642b4b · outbound

This paper cites Diffusionret: Generative text-video retrieval with diffusion model, 2023.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Diffusionret: Generative text-video retrieval with diffusion model, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.476521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.557860Z digest=sha256:d712675a8fb65240f2752b737e36e7b53e48ac32fb19c0a5fcc5ae53269be99b

Observation 7e84c5c3-f984-47f6-8364-31ec71538ba3 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.561282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.561282Z digest=sha256:dcd98953b0994c182edebc4f25f32b880b50d6368267bae1382d45d7bbf7d907

Observation f31af705-9116-4cde-8097-d2707e6b1939 · outbound

This paper cites Large Language Models Must Be Taught to Know What They Don't Know.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Large Language Models Must Be Taught to Know What They Don't Know

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.565722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.565722Z digest=sha256:c47d1cd8669da6e03207ec67f8229fc83ac517c685f0685ab5a47075163a6cca

Observation 4f91482b-c689-4763-a06a-5d7a9054805b · outbound

This paper cites Uncertainty-Aware Evaluation for Vision-Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Uncertainty-Aware Evaluation for Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.569740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.569740Z digest=sha256:7ffcb2e54266be255a397ff1164055ba6434c95dfe864fa6a1b0d1cfeab783be

Observation 1793bfdc-bdeb-44ae-980b-102e7a7189f6 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Detecting mo- ments and highlights in videos via natural language queries

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.459089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.574528Z digest=sha256:0323afef9bef407f670649e75e8ab30d2379c6a676c69578974a77f71202faa4

Observation 46ec3526-134e-46ce-bf8d-cdbd49dce8be · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.578279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.578279Z digest=sha256:86739a699c71e523280a664952d924346fde1418a671b9b2fffd2f9640342964

Observation 76d82182-d59a-49bd-8510-a33aeb440e24 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.582466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.582466Z digest=sha256:36393faf1dc93277c2b7f1ca2e53aa04c0a14904c22aa3c94c05e7b7a609b948

Observation 2eb60991-0dce-4779-9f93-9019603535f1 · outbound

This paper cites Ground- inggpt: Language enhanced multi-modal grounding model.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Ground- inggpt: Language enhanced multi-modal grounding model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.441298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.586499Z digest=sha256:7eae3472ebeee6df52af2bb71907612d9fb62abe9a8116d3ae4c5e5b24133172

Observation cd4f3baa-467a-4ef6-bce8-b42ab4540ad7 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.590376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.590376Z digest=sha256:0876090e69f2d82bd10adfbee629389b9925d616564911506a1cbf54ed5b189b

Observation 05cbee80-bb5a-4306-b1f2-fb7b8f0f0578 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Univtg: Towards unified video- language temporal grounding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.430557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.594249Z digest=sha256:ee41583be6ef465863123eb6aa0eb5cb1b623c13be134771a8ba219c591a0df8

Observation 39bb5fd4-231b-40aa-a575-3eb309801b08 · outbound

This paper cites Visual Instruction Tuning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.597917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.597917Z digest=sha256:7f7e2134c23eea32370855459654b8d5d9cb8ef52a40a058ea930270aa49bc47

Observation 34c89be5-11c2-4918-a3c3-3aba6a6ddf88 · outbound

This paper cites ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.602209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.602209Z digest=sha256:4454d72d152f6ef2dcd37fd3bb56f3be5dc9f23f4dc06cb6266909db6e989576

Observation 96bdaf0f-68b8-4ce2-b281-74fd7486bc94 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Lost in the Middle: How Language Models Use Long Contexts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.605648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.605648Z digest=sha256:22106f77a1009b9f8f9fc208644711eadcec0c6abab8782c57e80e9f9efe3540

Observation 353f11b3-9fc8-4847-b641-021352ca4cbd · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.420410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.609270Z digest=sha256:a89d67bf4863b06b5f708c7228574a1f2be78cd642a30f47444d1d7653d7bf23

Observation ab253223-2440-4d2e-9dcf-5f41221f69d6 · outbound

This paper cites Decoupled weight decay regularization, 2019.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Decoupled weight decay regularization, 2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.409303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.612626Z digest=sha256:03c6b0beec4dfa510b5b5519b838a6d310992706b346a7647341879d2d95231e

Observation 37251822-9afd-464b-9f2f-b3981631e3be · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.615743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.615743Z digest=sha256:a9936e6f2c7dc68e639d259c1333189ca2026d86e8ceff89a8e4d05a469bb57a

Observation 19ff04d7-30bc-4f41-b283-d45d25fb7c41 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.397946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.618988Z digest=sha256:757a805570e27e121808638c89831bfc296009bbc26f1634e2d5f86ea4040496

Observation ebb83402-3994-44ec-9a7b-18b8f6713bf9 · outbound

This paper cites Snag: Scalable and accurate video grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Snag: Scalable and accurate video grounding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.386923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.622614Z digest=sha256:cb08b5a0f99259f733af41c65030452df1ced1d5e8213e4c92d6698225f4e66e

Observation 68f00f37-6ea9-412c-aede-02cffb35ab00 · outbound

This paper cites Towards Calibrated Robust Fine-Tuning of Vision-Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Towards Calibrated Robust Fine-Tuning of Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.626259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.626259Z digest=sha256:085e990987f0048da706c81dfd156e473db9fa8163c1be034bdfa3a57781c253

Observation 2d85814d-1246-4674-b31f-fe6f5d6d28f6 · outbound

This paper cites Obtaining well calibrated probabilities using bayesian binning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Obtaining well calibrated probabilities using bayesian binning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.375813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.630348Z digest=sha256:36d26e8497ddd8a72bddfb702b3144761db8adbe7da8e3f40527b3fb2001beab

Observation 8c9c6ddd-4e66-4dc0-bd29-dfb737e37993 · outbound

This paper cites Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.634152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.634152Z digest=sha256:3ee7f29bd439f6431a801898acee2fdc778c588b844478fdb8f6d1525f3387ba

Observation f894d7f6-1da3-4fff-8285-c0f3687dff1e · outbound

This paper cites Momen- tor: Advancing video large language model with fine-grained temporal reasoning, 2024.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Momen- tor: Advancing video large language model with fine-grained temporal reasoning, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.363249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.638044Z digest=sha256:dcfcd55594eaedb94c745cf8fad747871a16417fb3bab254b2e6ccb7d6608064

Observation 385ffb20-aa04-4efe-9481-fc8712c9a1ea · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.351289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.641966Z digest=sha256:0a89e846b69a4845b23c496fdc48644eb12a754dad3a62b89f6ba2b7f5b62ffb

Observation 369d7e0f-0aba-4b96-8864-7d0504029881 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learn- ing transferable visual models from natural language super- vision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.339416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.645735Z digest=sha256:9c8f1038efa3f7266b125da3833ad233764fb1f8776c872ab00114b320a0bf47

Observation bd3e9ee6-103f-49aa-92b8-11e2a1e79dad · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.327828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.649213Z digest=sha256:c3befd927e23ef104f67df96a64074fb8fa7b7fbc3deaacd9eaa606665efac11

Observation fa1786fc-b54f-462c-9a64-d3e8d146a890 · outbound

This paper cites Vlg-net: Video-language graph matching network for video grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vlg-net: Video-language graph matching network for video grounding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.315273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.652731Z digest=sha256:6e8da9b14087a7ad67a5fe69d0b60d53f2da7b7f2b80bec2ae0eaaa70fdf52da

Observation 3cb809b6-898d-4661-918b-d83729b0dc55 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Mad: A scalable dataset for language grounding in videos from movie audio descriptions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.656723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.656723Z digest=sha256:c105825deb6cfe95dc5b4543fa16c05f9227bd36a84d7ed46bdbd2d0c475fe3f

Observation 961da660-c1f6-4d3b-9626-7d6cb1786d99 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.660574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.660574Z digest=sha256:5a75de7dda593a0891d25e723b1bf6829d3478eac96c99848fe176e22b48ef76

Observation e7322f6a-c423-4ba0-b032-eb1cf4246819 · outbound

This paper cites Avicuna: Audio-visual llm with interleaver and context- boundary alignment for temporal referential dialogue.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Avicuna: Audio-visual llm with interleaver and context- boundary alignment for temporal referential dialogue

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.664377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.664377Z digest=sha256:794687fcf7fe11a44bfa14aba5d22236954e7de55b034c7202c7c7c2b6eb6782

Observation ec48ea8d-2526-42e0-8254-dd4948a38cd6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LLaMA: Open and Efficient Foundation Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.667958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.667958Z digest=sha256:6b11d223ca4c0ea34408d19691c387bd4456cf4fcdd83bbf733fed8998161b9f

Observation 65b3fd40-def7-460e-bfd8-1db356291c92 · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.672191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.672191Z digest=sha256:7261cdd493a54d0c0b49c82e1c7d0026b18edcc472b3e6cfbb38022c5b4f0704

Observation a4943ab4-0d48-4b0a-9490-36a41f56a9a5 · outbound

This paper cites Omnivid: A generative framework for universal video understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Omnivid: A generative framework for universal video understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.675817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.675817Z digest=sha256:370a06bbf0ba4347d9e82fc267c4bf001f9a6c323cb5095920d6eb5c69fbccc3

Observation f34862bb-cd24-414a-b9a5-ddde6d1215ee · outbound

This paper cites Text is mass: Modeling as stochastic embedding for text-video retrieval, 2024.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Text is mass: Modeling as stochastic embedding for text-video retrieval, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.283505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.679428Z digest=sha256:5954fc406f5871f4e82cdc94e44fb4e7dcf33063b2b82d2be59d053138b1c103

Observation 09b2e3e7-88b7-464a-aabb-3bfc16f0a067 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.683167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.683167Z digest=sha256:dda6518182f2f8009af991208c6713b0cb27cf781a50232aafcfba1de8955985

Observation b96e8328-abb2-4c62-89d9-c9ad92851365 · outbound

This paper cites Five factors that guide attention in visual search.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Five factors that guide attention in visual search

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.270664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.687158Z digest=sha256:e08eafd9f4e9caae60f47d0ee02a12b6fbf5198612d55fc401fd87b9942d92d8

Observation 33b0a0e2-eeea-45ab-ba58-bc38164d8fb1 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Msr-vtt: A large video description dataset for bridging video and language

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.258309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.690819Z digest=sha256:43533344ea4357c54e92999a2e88d0c99f4e295cd66cdcca599377f2666edc4e

Observation 1a570ca5-d499-4696-9f06-b73cd1f2e861 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.695172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.695172Z digest=sha256:db0e3a9b812ee50df748eabc89ecc784a37302416981297141fa39efbd28c407

Observation 3d11ef86-3c09-4298-adee-355fb136229e · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.699205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.699205Z digest=sha256:49f9824d27b30dd75d300426c733badea752ebced248c1b0eb2db2936bcecdb8

Observation 9e40798d-db98-4886-944f-82aedba8fb24 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Clip-vip: Adapting pre- trained image-text model to video-language representation alignment

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.247005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.703027Z digest=sha256:afeac72e308d1bda04afd81b67ba8917464bd447e28a25070367b5bd63938dba

Observation 6290a9e5-877f-4046-b121-d32a102f6e91 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment, 2023.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Clip-vip: Adapting pre- trained image-text model to video-language representation alignment, 2023

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.235064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.706082Z digest=sha256:d7a0e59d87e08620d78f7008843b7959c3a131d143ba22552d66f2f44b0e4646

Observation 39f0e8a5-21c9-470e-b485-db4f92b1617e · outbound

This paper cites Vidchapters-7m: Video chapters at scale,.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vidchapters-7m: Video chapters at scale,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.223591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.709255Z digest=sha256:e9e1dae203725f1ccd4d5a136015ebcd5e2a5b777f263ce5b484d4d0102efc3e

Observation c9001ed8-d3d1-4cf2-89d8-19583cb0605e · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.212240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.712563Z digest=sha256:5aad7a1312e4fa4e9c221427b741d9a84b8ccdbcf7727cd09a5a79b2d08f2dd2

Observation fe54dfd1-c547-4831-b9f9-96fa5c24a3a1 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Self-chained image-language model for video localization and question answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.715867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.715867Z digest=sha256:67a9d0801f6517b42b2acb2056eaa9414bd8057f8eb2ce8ca75f35d79a2fd0f8

Observation 6dc5565c-50d5-4e59-9bab-48ea7cf2f19e · outbound

This paper cites A joint se- quence fusion model for video question answering and re- trieval.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos A joint se- quence fusion model for video question answering and re- trieval

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.718994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.718994Z digest=sha256:8c0c413ab89f6ae9b77f82f752595c78a4ab6810d7bb46930ef85e8f38aaf227

Observation 4db7b359-bc0c-425b-8507-1cbba8219f49 · outbound

This paper cites A sim- ple llm framework for long-range video question-answering,.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos A sim- ple llm framework for long-range video question-answering,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.722170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.722170Z digest=sha256:719b4465f220e9224c1c90ea6a1e95b0ea89709bb746d0de33a5320a35816a38

Observation 11d10cb2-8c7c-4929-a6b0-fef46608327a · outbound

This paper cites Span-based Localizing Network for Natural Language Video Localization.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Span-based Localizing Network for Natural Language Video Localization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.726089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.726089Z digest=sha256:eaa3eb9c3a2d67509cb4bee2a9c55aea0e91e4ef72428fbe0c8c32bc88767d5d

Observation b79c7ce4-cf4f-45e0-be7b-fba95f2d2af5 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.730863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.730863Z digest=sha256:ad1633088f4b78d72518dbb8f2f6e26cfbd85cd09f73a4f2dc36a0de0d4a3a09

Observation ed8e90c1-123a-42a6-8552-8f642f137f6f · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.179532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.734969Z digest=sha256:d36efde24c08f2ee0b212b9462b59c648743f6b54a4b446070b2d82dc0ff248e

Observation 4458d16d-b41f-4419-8d00-ddaebcc8bdbc · outbound

This paper cites Learning video representations from large lan- guage models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learning video representations from large lan- guage models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.167474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.739263Z digest=sha256:1aa6f86246f97182d3efcd61d7b3acb8d5708ca1c6969238af83c7f4cd91603d

Observation 6de24dff-be19-4a52-86a4-6078f8679cf5 · outbound

This paper cites <video> Does the <event> happen in the video? Answer yes or no.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos <video> Does the <event> happen in the video? Answer yes or no

Reference 4096

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.154762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.742891Z digest=sha256:33a7c2835e7031d412c6dab3d5f5d77a152ba457fe09caaae87bd76b9aa33869

Pith citing papers

Observation fbe52eca-4221-4173-a058-749e81744064 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:32:18.794543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T23:32:17.962997Z digest=sha256:319067b7665f94f2cb23ee0a4dd9c27d86eb1eda0f861216d187ccb92746921a