Pith. sign in

Paper Citation Record · LEDGER

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2411.14901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14901 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:51:17.742891Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:32:17.962997Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T23:32:18.789296Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact1
  • verified fuzzy37
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6de1641-e8d3-49f1-9180-66134e2807f3 · outbound

This paper cites needle in a haystack.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos needle in a haystack

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.633581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.480428Z digest=sha256:bd9935641a6440266b42b6ffcd446f5450d68d93a5b8da833056685bb725b603

Observation 14f8428d-d7b1-4c95-9845-40a18e4758a1 · outbound

This paper cites needle in a haystack.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos needle in a haystack

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.485527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.485527Z digest=sha256:1661a23b05c3a09a9d3e64d1b63b7aaba55d65e1825de35def566d6e5da1ad52

Observation 932cbd73-d287-489b-86e0-771313651af9 · outbound

This paper cites https://sharegpt.com/.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos https://sharegpt.com/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.621707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.489695Z digest=sha256:eb9f02f957bbfdf0bb3efcaafef230c2abe4c5468053cc6bdf2691f1e25088ea

Observation 2494fcf0-7316-44fd-8728-03845a06f407 · outbound

This paper cites https://github.com/gkamradt/LLMTest_ NeedleInAHaystack.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos https://github.com/gkamradt/LLMTest_ NeedleInAHaystack

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.610355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.493940Z digest=sha256:f27369339369e9e3e920823b9a0332d17a57a259ee91f42ade98be218e4495c0

Observation 514194ae-6741-4e06-b166-858bb214e022 · outbound

This paper cites Lo- calizing moments in long video via multimodal guidance.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Lo- calizing moments in long video via multimodal guidance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.598205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.497775Z digest=sha256:bc316b270e36fe15980da51c2f41a8559ebf62c9bca3508579317609b9693876

Observation 70c6f739-fbc3-4dfb-9a5b-18bf90ef7e20 · outbound

This paper cites Functional brain organization of preparatory attentional control in visual search.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Functional brain organization of preparatory attentional control in visual search

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.586723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.501746Z digest=sha256:b5578488b4fcbc13a05884c5d4c1745d8c9b46971a5323506b3009874516540c

Observation c3a58dae-6d8d-4df8-9152-0553858fa38f · outbound

This paper cites End-to- end object detection with transformers.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos End-to- end object detection with transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.505724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.505724Z digest=sha256:7deb3b801b3ff84475e716bc39d0c8deab36e210d718af9f1b9e5d9e6d841838

Observation f6a37a7e-6852-4567-ba07-10d991286069 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos VideoLLM: Modeling Video Sequence with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.509494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.509494Z digest=sha256:a944e63524aa9c030078d2bfec3e0040b3ec4758bee6da33a2ff048b13c11b9d

Observation b8ee4b52-d852-41af-b32a-ee00329f0aad · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.567451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.513431Z digest=sha256:6693bcf65da2443527e749fc062d1a79f1d836bdf5ec7d7743e70ec047bdb601

Observation 8425f23a-9f25-42db-b812-904a74ca0259 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.556051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.516640Z digest=sha256:26bd920ca519c68c171668eb4656b120e2fee72ff2ab15ec4086220f8a77ef8b

Observation 7bfc6ce1-d9db-4510-b89a-3b9e40b894c4 · outbound

This paper cites Uatvr: Uncertainty-adaptive text-video retrieval,.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Uatvr: Uncertainty-adaptive text-video retrieval,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.543891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.519848Z digest=sha256:03037f28719259579bd422916dd8b3bad072daf151f0a392edc824eaa266552a

Observation b54cee9a-901a-4e69-b862-fad754abe3e1 · outbound

This paper cites Multi-modal transformer for video retrieval.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Multi-modal transformer for video retrieval

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.531878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.523083Z digest=sha256:4a4b537e51f7f60b19fc5f3603fcdc67cf2b4ba8f0e6d0f2901ed918ed20dd54

Observation d59df371-f9dc-4faf-9877-95bd9c219341 · outbound

This paper cites AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.526246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.526246Z digest=sha256:ab3af3920f8e6412e1c47abfe7b2909f82aad74b46767805f44b3f78c60c1b73

Observation c23fb26c-f5c4-464b-8e35-9534ba353132 · outbound

This paper cites X-pool: Cross-modal language-video attention for text- video retrieval, 2022.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos X-pool: Cross-modal language-video attention for text- video retrieval, 2022

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.520335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.529565Z digest=sha256:1ca85ed11bb883fcd4972e8a48cf3e7278348203c9811d1cb902eb1880837f7a

Observation bc23c703-d543-4919-8f23-5065e7fcfed0 · outbound

This paper cites RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:51:18.046182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.532531Z digest=sha256:cfacd8ac488f1265d5bf0a9371d06b74a6e411d605d1bfd777d72488d08d60e4

Observation 255d90f0-9d27-4813-ad21-2f1b032353c9 · outbound

This paper cites CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.536220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.536220Z digest=sha256:22d90f7cf43c4e690ce2d8c1097c42551eab74a0b151f4f5346fccd39a66097a

Observation 4395324c-3dd4-4766-adfa-b8ea1ccf4037 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.539848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.539848Z digest=sha256:f7ac354c2e97e05577e4471e7d727111eb4ae1f03a08b6ea74d5b712ab300f58

Observation 857ffaaa-1626-4923-8342-37e56075e9d4 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vtimellm: Empower llm to grasp video moments

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.509395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.543649Z digest=sha256:58a473e1183f9a214de8fd66b25330f912081c18b6be73f31a5d80b854b1add9

Observation 5760acec-746b-45f6-bfbc-d5a67ef8412f · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LITA: Language Instructed Temporal-Localization Assistant

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.547199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.547199Z digest=sha256:a749514f99b9e6841068dfbcc7e513e0ee889ce3bdbb4de8a5a400662b2751a8

Observation 6e092eaf-b30b-4139-ac67-d61fcb73725d · outbound

This paper cites Audio- enhanced text-to-video retrieval using text-conditioned fea- ture alignment, 2023.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Audio- enhanced text-to-video retrieval using text-conditioned fea- ture alignment, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.498330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.551108Z digest=sha256:54bc94c1b255a57a54bafb25ca32c202a944cd5345cce61a9c4e135c150d4c22

Observation 5fdfe32c-e0ae-4d09-97d9-3009ec7483f9 · outbound

This paper cites Efficient long- text understanding with short-text models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Efficient long- text understanding with short-text models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.487239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.554425Z digest=sha256:bc76cb0eac564779e05bd6d76a3abe437c77e07788be381a2efecab4edeb12f5

Observation 859ddc87-c4f6-4c55-a492-ae3650642b4b · outbound

This paper cites Diffusionret: Generative text-video retrieval with diffusion model, 2023.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Diffusionret: Generative text-video retrieval with diffusion model, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.476521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.557860Z digest=sha256:28b515a98de60b7405649b8de104decdf6dbcd8ad03fafbd141ee77af977b921

Observation 7e84c5c3-f984-47f6-8364-31ec71538ba3 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.561282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.561282Z digest=sha256:92ab29b28d34589fa6bc79fffda2e5488bfc43a81ddc745a983db7aef4e6d527

Observation f31af705-9116-4cde-8097-d2707e6b1939 · outbound

This paper cites Large Language Models Must Be Taught to Know What They Don't Know.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Large Language Models Must Be Taught to Know What They Don't Know

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.565722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.565722Z digest=sha256:1a8f2ae5647ead93cc70d4a4fc5bde06a22d76fb57e5206462cc1f066b01488f

Observation 4f91482b-c689-4763-a06a-5d7a9054805b · outbound

This paper cites Uncertainty-Aware Evaluation for Vision-Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Uncertainty-Aware Evaluation for Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.569740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.569740Z digest=sha256:f3bd4c3b0660b857f8607c7ea8df8089a980428b17dcf89d14e16a2c06405f15

Observation 1793bfdc-bdeb-44ae-980b-102e7a7189f6 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Detecting mo- ments and highlights in videos via natural language queries

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.459089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.574528Z digest=sha256:928c170233455070fe46812e360f73438b0206b50890ee986436bb16434a737c

Observation 46ec3526-134e-46ce-bf8d-cdbd49dce8be · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.578279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.578279Z digest=sha256:d06985a11ef29bf437bd0b90e94553324b662486698240641656d26b644ecf31

Observation 76d82182-d59a-49bd-8510-a33aeb440e24 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.582466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.582466Z digest=sha256:3f8a6832cd55a2c7db3591e6f6ccaab606e43ade3a90775a40b48973834622e3

Observation 2eb60991-0dce-4779-9f93-9019603535f1 · outbound

This paper cites Ground- inggpt: Language enhanced multi-modal grounding model.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Ground- inggpt: Language enhanced multi-modal grounding model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.441298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.586499Z digest=sha256:fe5ff2dd78f9b492dec222edaa1bf42c4531b8eb187c8b6b2a8c7bfda9d4d6d2

Observation cd4f3baa-467a-4ef6-bce8-b42ab4540ad7 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.590376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.590376Z digest=sha256:da2b6a34fca2865c11f9c7281474dd3052626026017d83cd3176ef6f3096939d

Observation 05cbee80-bb5a-4306-b1f2-fb7b8f0f0578 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Univtg: Towards unified video- language temporal grounding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.430557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.594249Z digest=sha256:0cc5b39eccf4d9a7600b55617d264b17e606897d1f595c648214cf7a5fbf126d

Observation 39bb5fd4-231b-40aa-a575-3eb309801b08 · outbound

This paper cites Visual Instruction Tuning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.597917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.597917Z digest=sha256:5a040e4e02f4164043b8a4aaa3b27d78eef8b23f73086cb8ec4ab2865848d5ea

Observation 34c89be5-11c2-4918-a3c3-3aba6a6ddf88 · outbound

This paper cites ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.602209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.602209Z digest=sha256:ea65fb52d5e7ddafefa430c9afd7c4063a2b4f93cd7777dec2c841e29461e3c2

Observation 96bdaf0f-68b8-4ce2-b281-74fd7486bc94 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Lost in the Middle: How Language Models Use Long Contexts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.605648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.605648Z digest=sha256:0f657beb44cf9a8339c29335f48c503694aacd0150e3f7948a35a458ac4341c2

Observation 353f11b3-9fc8-4847-b641-021352ca4cbd · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.420410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.609270Z digest=sha256:9166d5be4c0547fbe0420aa0b1fbbb37b6cd36f76cb78f78c182b6d224a10714

Observation ab253223-2440-4d2e-9dcf-5f41221f69d6 · outbound

This paper cites Decoupled weight decay regularization, 2019.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Decoupled weight decay regularization, 2019

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.409303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.612626Z digest=sha256:7e49e56013ad7ac44fc952f12b44dc732c586518bf3ddaffa007c5e7b724df40

Observation 37251822-9afd-464b-9f2f-b3981631e3be · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.615743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.615743Z digest=sha256:b18d5d02c7094fae014910d89868181de2338e7dbc09184247ed413b8d235ce5

Observation 19ff04d7-30bc-4f41-b283-d45d25fb7c41 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.397946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.618988Z digest=sha256:49ac5ccde727012bbf11b36c4bde64b0ba6ea62f8e77d1d510df00b6fa03f141

Observation ebb83402-3994-44ec-9a7b-18b8f6713bf9 · outbound

This paper cites Snag: Scalable and accurate video grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Snag: Scalable and accurate video grounding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.386923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.622614Z digest=sha256:f89af6f7ff46b4151fc0c44abaccd86f2e4c9d54fec17815088d801cd8cf5fd6

Observation 68f00f37-6ea9-412c-aede-02cffb35ab00 · outbound

This paper cites Towards Calibrated Robust Fine-Tuning of Vision-Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Towards Calibrated Robust Fine-Tuning of Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.626259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.626259Z digest=sha256:4910da00d8c58f6617ae751c86fe5b74eb948f031c13d25ba6e40b995d52b75c

Observation 2d85814d-1246-4674-b31f-fe6f5d6d28f6 · outbound

This paper cites Obtaining well calibrated probabilities using bayesian binning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Obtaining well calibrated probabilities using bayesian binning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.375813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.630348Z digest=sha256:0e11ecc177910ca257fc75d7acc991f8b43eb63b9190d41d17cab7839ef609a0

Observation 8c9c6ddd-4e66-4dc0-bd29-dfb737e37993 · outbound

This paper cites Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.634152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.634152Z digest=sha256:ecca4ea9b3b0a78d7c6b10040950d49688cdc84aedc4917762cc91da066912f9

Observation f894d7f6-1da3-4fff-8285-c0f3687dff1e · outbound

This paper cites Momen- tor: Advancing video large language model with fine-grained temporal reasoning, 2024.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Momen- tor: Advancing video large language model with fine-grained temporal reasoning, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.363249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.638044Z digest=sha256:cae7d5bbb30c4a2e8ac0ef0c1caf372f381b8e048736d0f5a3d969fd56ee26ee

Observation 385ffb20-aa04-4efe-9481-fc8712c9a1ea · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.351289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.641966Z digest=sha256:98ffb775921d016ce7c0908c609523497e8eab9ddac37399957b25491cba274d

Observation 369d7e0f-0aba-4b96-8864-7d0504029881 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learn- ing transferable visual models from natural language super- vision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.339416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.645735Z digest=sha256:8449669505330a755538fa97b934dde9e7289d586d2e82c00e6bdc13521d82ea

Observation bd3e9ee6-103f-49aa-92b8-11e2a1e79dad · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.327828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.649213Z digest=sha256:3f78430e0113423445b7e1f1cc24ffb701be53b313fb364ededdb6ed25cfac1e

Observation fa1786fc-b54f-462c-9a64-d3e8d146a890 · outbound

This paper cites Vlg-net: Video-language graph matching network for video grounding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vlg-net: Video-language graph matching network for video grounding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.315273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.652731Z digest=sha256:f5fdf012ad8911c0257372fc0bf313aac42c1688bbf0c7f9fb0a7d55a0525da7

Observation 3cb809b6-898d-4661-918b-d83729b0dc55 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Mad: A scalable dataset for language grounding in videos from movie audio descriptions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.656723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.656723Z digest=sha256:3363137cb3195d454964563aa0898b5fe5a518e4cd68d3cda40b3b63eac5ac49

Observation 961da660-c1f6-4d3b-9626-7d6cb1786d99 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Moviechat: From dense token to sparse memory for long video understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.660574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.660574Z digest=sha256:b807e81e87b90f6634f7ace9c99f829daa4cc5db4878b72fb7d74ac842a172dc

Observation e7322f6a-c423-4ba0-b032-eb1cf4246819 · outbound

This paper cites Avicuna: Audio-visual llm with interleaver and context- boundary alignment for temporal referential dialogue.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Avicuna: Audio-visual llm with interleaver and context- boundary alignment for temporal referential dialogue

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.664377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.664377Z digest=sha256:a998dcb7df240eccb5bdc5d3e71caf0045d899122de0d72e3a6e2fc4d56b2990

Observation ec48ea8d-2526-42e0-8254-dd4948a38cd6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LLaMA: Open and Efficient Foundation Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.667958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.667958Z digest=sha256:0995a2373857ae06a39e02a0764a8e59c0f3004a1560e3655129dacc2cd223d3

Observation 65b3fd40-def7-460e-bfd8-1db356291c92 · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.672191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.672191Z digest=sha256:268cddd783111c91d57d417aea1daab0d0ed6dacc6acb48bf592c7f7ab183f53

Observation a4943ab4-0d48-4b0a-9490-36a41f56a9a5 · outbound

This paper cites Omnivid: A generative framework for universal video understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Omnivid: A generative framework for universal video understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.675817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.675817Z digest=sha256:91ef583a58ca13ce4b14ed35dbcb657bc9f03d7de80714f1519952ecd75a9a29

Observation f34862bb-cd24-414a-b9a5-ddde6d1215ee · outbound

This paper cites Text is mass: Modeling as stochastic embedding for text-video retrieval, 2024.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Text is mass: Modeling as stochastic embedding for text-video retrieval, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.283505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.679428Z digest=sha256:15efd26157de5eb85792238211bf37363b732ea4b2ebcc639c9946b8943f90fe

Observation 09b2e3e7-88b7-464a-aabb-3bfc16f0a067 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.683167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.683167Z digest=sha256:b05ed801d57c0f5ebf4780a5a60e6b9364a972a4eecffe60a2285cb591456fcf

Observation b96e8328-abb2-4c62-89d9-c9ad92851365 · outbound

This paper cites Five factors that guide attention in visual search.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Five factors that guide attention in visual search

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.270664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.687158Z digest=sha256:6aae81389e58f1fc40d72ebc044a9033eab28d6fe4824150cdc22237003a1cd5

Observation 33b0a0e2-eeea-45ab-ba58-bc38164d8fb1 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Msr-vtt: A large video description dataset for bridging video and language

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.258309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.690819Z digest=sha256:f602e21f5739f0e79c45fdcb7ddd828b58a7a3d31c6c4c3f6884a90ae8be0074

Observation 1a570ca5-d499-4696-9f06-b73cd1f2e861 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.695172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.695172Z digest=sha256:dd714dcc92ef9c822d66ffe8dc77b769894385c834892145062b97b1c8014d4e

Observation 3d11ef86-3c09-4298-adee-355fb136229e · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.699205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.699205Z digest=sha256:aabc9c7793a0bad8799d165129431aa4a9234f836a8969de527023dbda67bdbe

Observation 9e40798d-db98-4886-944f-82aedba8fb24 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Clip-vip: Adapting pre- trained image-text model to video-language representation alignment

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.247005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.703027Z digest=sha256:4dabdec2247dac09d788d06ccb2321bfe96e6c5462a1491d32cfb3d1f888f425

Observation 6290a9e5-877f-4046-b121-d32a102f6e91 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment, 2023.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Clip-vip: Adapting pre- trained image-text model to video-language representation alignment, 2023

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.235064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.706082Z digest=sha256:87c7c356228b80f65185f5e278714131f3472a5396c46cb3b4e2129c016b4bcc

Observation 39f0e8a5-21c9-470e-b485-db4f92b1617e · outbound

This paper cites Vidchapters-7m: Video chapters at scale,.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vidchapters-7m: Video chapters at scale,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.223591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.709255Z digest=sha256:0b79969edceb314f8ae3562fa56bead2fc700a7490ba5f43fc36b9aa5eca40cf

Observation c9001ed8-d3d1-4cf2-89d8-19583cb0605e · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.212240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.712563Z digest=sha256:3907538d861fe1ff1a9ac95b23a57ba74a9f66a877e466358689987f1c75d6aa

Observation fe54dfd1-c547-4831-b9f9-96fa5c24a3a1 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Self-chained image-language model for video localization and question answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.715867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.715867Z digest=sha256:bbafc1d5905ed9f4f4277b18a6ee59f2ec8b20a80a1ba97d13ff9d5db1e9f200

Observation 6dc5565c-50d5-4e59-9bab-48ea7cf2f19e · outbound

This paper cites A joint se- quence fusion model for video question answering and re- trieval.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos A joint se- quence fusion model for video question answering and re- trieval

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.718994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.718994Z digest=sha256:271531d77a71e4127e88e5a37a1b0b9085a6c45c98a512dbfde670ca6fc60fde

Observation 4db7b359-bc0c-425b-8507-1cbba8219f49 · outbound

This paper cites A sim- ple llm framework for long-range video question-answering,.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos A sim- ple llm framework for long-range video question-answering,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.722170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.722170Z digest=sha256:78a8b19940b5de0375e8242a7a6fd17d08f316d05a78a1b913d96abb6469e9eb

Observation 11d10cb2-8c7c-4929-a6b0-fef46608327a · outbound

This paper cites Span-based Localizing Network for Natural Language Video Localization.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Span-based Localizing Network for Natural Language Video Localization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.726089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.726089Z digest=sha256:ffbe6f7116365464b464c8ed0380d2a4ec5d18756806e2991aa1479b5f8cb5a5

Observation b79c7ce4-cf4f-45e0-be7b-fba95f2d2af5 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T14:51:17.730863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:51:17.730863Z digest=sha256:9f05c17967ccebd75d8307548b026710ead4e7da5c141bce7bdbf4248eb847a4

Observation ed8e90c1-123a-42a6-8552-8f642f137f6f · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.179532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.734969Z digest=sha256:071f1daeaceabc68d7bde38b4710bae5c3bb2a32e133ff355a3182e14b331177

Observation 4458d16d-b41f-4419-8d00-ddaebcc8bdbc · outbound

This paper cites Learning video representations from large lan- guage models.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learning video representations from large lan- guage models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.167474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.739263Z digest=sha256:d4b8301b376ec17645ce49a81658ac33549cddfbc3269419cc571d423475ae33

Observation 6de24dff-be19-4a52-86a4-6078f8679cf5 · outbound

This paper cites <video> Does the <event> happen in the video? Answer yes or no.

ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos <video> Does the <event> happen in the video? Answer yes or no

Reference 4096

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:51:18.154762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:51:17.742891Z digest=sha256:1010de89ba918c24365e32c730791377fd8871677c0e5307c7ccf25b16575cba

Pith citing papers

Observation fbe52eca-4221-4173-a058-749e81744064 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:32:18.794543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-05T23:32:17.962997Z digest=sha256:cfea3f9280514b93b91d696cc10ef6517ad57ee2a7653180f65bb30beb45ee26