Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:51:17.742891Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2411.14901.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:51:17.742891Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:32:17.962997Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T23:32:18.789296Z
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f6de1641-e8d3-49f1-9180-66134e2807f3 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos needle in a haystack
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 14f8428d-d7b1-4c95-9845-40a18e4758a1 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos needle in a haystack
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 932cbd73-d287-489b-86e0-771313651af9 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos https://sharegpt.com/
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2494fcf0-7316-44fd-8728-03845a06f407 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos https://github.com/gkamradt/LLMTest_ NeedleInAHaystack
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 514194ae-6741-4e06-b166-858bb214e022 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Lo- calizing moments in long video via multimodal guidance
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 70c6f739-fbc3-4dfb-9a5b-18bf90ef7e20 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Functional brain organization of preparatory attentional control in visual search
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c3a58dae-6d8d-4df8-9152-0553858fa38f · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos End-to- end object detection with transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6a37a7e-6852-4567-ba07-10d991286069 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos VideoLLM: Modeling Video Sequence with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ee4b52-d852-41af-b32a-ee00329f0aad · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Gonzalez, Ion Stoica, and Eric P
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8425f23a-9f25-42db-b812-904a74ca0259 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Bert: Pre-training of deep bidirectional trans- formers for language understanding, 2019
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7bfc6ce1-d9db-4510-b89a-3b9e40b894c4 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Uatvr: Uncertainty-adaptive text-video retrieval,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b54cee9a-901a-4e69-b862-fad754abe3e1 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Multi-modal transformer for video retrieval
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d59df371-f9dc-4faf-9877-95bd9c219341 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23fb26c-f5c4-464b-8e35-9534ba353132 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos X-pool: Cross-modal language-video attention for text- video retrieval, 2022
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bc23c703-d543-4919-8f23-5065e7fcfed0 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 255d90f0-9d27-4813-ad21-2f1b032353c9 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4395324c-3dd4-4766-adfa-b8ea1ccf4037 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LoRA: Low-Rank Adaptation of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 857ffaaa-1626-4923-8342-37e56075e9d4 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vtimellm: Empower llm to grasp video moments
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5760acec-746b-45f6-bfbc-d5a67ef8412f · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LITA: Language Instructed Temporal-Localization Assistant
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e092eaf-b30b-4139-ac67-d61fcb73725d · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Audio- enhanced text-to-video retrieval using text-conditioned fea- ture alignment, 2023
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5fdfe32c-e0ae-4d09-97d9-3009ec7483f9 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Efficient long- text understanding with short-text models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 859ddc87-c4f6-4c55-a492-ae3650642b4b · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Diffusionret: Generative text-video retrieval with diffusion model, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7e84c5c3-f984-47f6-8364-31ec71538ba3 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31af705-9116-4cde-8097-d2707e6b1939 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Large Language Models Must Be Taught to Know What They Don't Know
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f91482b-c689-4763-a06a-5d7a9054805b · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Uncertainty-Aware Evaluation for Vision-Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1793bfdc-bdeb-44ae-980b-102e7a7189f6 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Detecting mo- ments and highlights in videos via natural language queries
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 46ec3526-134e-46ce-bf8d-cdbd49dce8be · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76d82182-d59a-49bd-8510-a33aeb440e24 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos VideoChat: Chat-Centric Video Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb60991-0dce-4779-9f93-9019603535f1 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Ground- inggpt: Language enhanced multi-modal grounding model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cd4f3baa-467a-4ef6-bce8-b42ab4540ad7 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05cbee80-bb5a-4306-b1f2-fb7b8f0f0578 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Univtg: Towards unified video- language temporal grounding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 39bb5fd4-231b-40aa-a575-3eb309801b08 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Visual Instruction Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c89be5-11c2-4918-a3c3-3aba6a6ddf88 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96bdaf0f-68b8-4ce2-b281-74fd7486bc94 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Lost in the Middle: How Language Models Use Long Contexts
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 353f11b3-9fc8-4847-b641-021352ca4cbd · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ab253223-2440-4d2e-9dcf-5f41221f69d6 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Decoupled weight decay regularization, 2019
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 37251822-9afd-464b-9f2f-b3981631e3be · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ff04d7-30bc-4f41-b283-d45d25fb7c41 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Query-dependent video representa- tion for moment retrieval and highlight detection
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ebb83402-3994-44ec-9a7b-18b8f6713bf9 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Snag: Scalable and accurate video grounding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 68f00f37-6ea9-412c-aede-02cffb35ab00 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Towards Calibrated Robust Fine-Tuning of Vision-Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d85814d-1246-4674-b31f-fe6f5d6d28f6 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Obtaining well calibrated probabilities using bayesian binning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8c9c6ddd-4e66-4dc0-bd29-dfb737e37993 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f894d7f6-1da3-4fff-8285-c0f3687dff1e · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Momen- tor: Advancing video large language model with fine-grained temporal reasoning, 2024
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 385ffb20-aa04-4efe-9481-fc8712c9a1ea · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Chatvtg: Video temporal grounding via chat with video dialogue large language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 369d7e0f-0aba-4b96-8864-7d0504029881 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learn- ing transferable visual models from natural language super- vision
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bd3e9ee6-103f-49aa-92b8-11e2a1e79dad · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fa1786fc-b54f-462c-9a64-d3e8d146a890 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vlg-net: Video-language graph matching network for video grounding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3cb809b6-898d-4661-918b-d83729b0dc55 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Mad: A scalable dataset for language grounding in videos from movie audio descriptions
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961da660-c1f6-4d3b-9626-7d6cb1786d99 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Moviechat: From dense token to sparse memory for long video understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7322f6a-c423-4ba0-b032-eb1cf4246819 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Avicuna: Audio-visual llm with interleaver and context- boundary alignment for temporal referential dialogue
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec48ea8d-2526-42e0-8254-dd4948a38cd6 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos LLaMA: Open and Efficient Foundation Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b3fd40-def7-460e-bfd8-1db356291c92 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4943ab4-0d48-4b0a-9490-36a41f56a9a5 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Omnivid: A generative framework for universal video understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f34862bb-cd24-414a-b9a5-ddde6d1215ee · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Text is mass: Modeling as stochastic embedding for text-video retrieval, 2024
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 09b2e3e7-88b7-464a-aabb-3bfc16f0a067 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos HawkEye: Training Video-Text LLMs for Grounding Text in Videos
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b96e8328-abb2-4c62-89d9-c9ad92851365 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Five factors that guide attention in visual search
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 33b0a0e2-eeea-45ab-ba58-bc38164d8fb1 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Msr-vtt: A large video description dataset for bridging video and language
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1a570ca5-d499-4696-9f06-b73cd1f2e861 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d11ef86-3c09-4298-adee-355fb136229e · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e40798d-db98-4886-944f-82aedba8fb24 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Clip-vip: Adapting pre- trained image-text model to video-language representation alignment
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6290a9e5-877f-4046-b121-d32a102f6e91 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Clip-vip: Adapting pre- trained image-text model to video-language representation alignment, 2023
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 39f0e8a5-21c9-470e-b485-db4f92b1617e · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vidchapters-7m: Video chapters at scale,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c9001ed8-d3d1-4cf2-89d8-19583cb0605e · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe54dfd1-c547-4831-b9f9-96fa5c24a3a1 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Self-chained image-language model for video localization and question answering
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc5565c-50d5-4e59-9bab-48ea7cf2f19e · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos A joint se- quence fusion model for video question answering and re- trieval
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db7b359-bc0c-425b-8507-1cbba8219f49 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos A sim- ple llm framework for long-range video question-answering,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11d10cb2-8c7c-4929-a6b0-fef46608327a · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Span-based Localizing Network for Natural Language Video Localization
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b79c7ce4-cf4f-45e0-be7b-fba95f2d2af5 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed8e90c1-123a-42a6-8552-8f642f137f6f · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learning 2d temporal adjacent networks for moment local- ization with natural language
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4458d16d-b41f-4419-8d00-ddaebcc8bdbc · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos Learning video representations from large lan- guage models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6de24dff-be19-4a52-86a4-6078f8679cf5 · outbound
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos <video> Does the <event> happen in the video? Answer yes or no
Reference 4096
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fbe52eca-4221-4173-a058-749e81744064 · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.