Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2411.04923.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:29:28.478864Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T07:16:45.136013Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d0e403d2-1654-421a-baa1-3aa765ea77ea · inbound
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c42c50f-28d6-4f2a-a0c1-5d2179bed05c · inbound
InterRVOS: Interaction-aware Referring Video Object Segmentation VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b072b2-aea8-4318-be73-26ac1948c2af · inbound
VideoMolmo: Spatio-Temporal Grounding Meets Pointing VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 388bf383-4714-45a5-bdcd-8e3fbec35f84 · inbound
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f13312fa-4a4e-4185-93fd-866c791bac9f · inbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc1d753-aa5c-4bb7-b320-61d974c04886 · inbound
Promptception: How Sensitive Are Large Multimodal Models to Prompts? VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe14149-175c-46ca-a48e-28dd91c9803d · inbound
GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 79b286d7-2b67-4b29-9d0a-09514d3b4227 · inbound
VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.