Pith. sign in

Paper Citation Record · LEDGER

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2411.04923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04923 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:29:28.478864Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:16:45.136013Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d0e403d2-1654-421a-baa1-3aa765ea77ea · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.504390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:aba3c05a1d71ac1e523553f2f58669c864b9ae02bd8d29a8b31724dfe4ae37b1

Observation 3c42c50f-28d6-4f2a-a0c1-5d2179bed05c · inbound

InterRVOS: Interaction-aware Referring Video Object Segmentation cites this paper.

InterRVOS: Interaction-aware Referring Video Object Segmentation VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.478864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.478864Z digest=sha256:8d37c5112fcd5a2cca9b69923b5b7b33f9b38fa0dc9d9b4d1c20b43368f16d35

Observation 23b072b2-aea8-4318-be73-26ac1948c2af · inbound

VideoMolmo: Spatio-Temporal Grounding Meets Pointing cites this paper.

VideoMolmo: Spatio-Temporal Grounding Meets Pointing VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:28.947430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:28.947430Z digest=sha256:80d8d06e792b0af73c52479d17b06dd954686ba666feb0c4b91437cb91df66f1

Observation 388bf383-4714-45a5-bdcd-8e3fbec35f84 · inbound

MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning cites this paper.

MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:59:10.935802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:59:10.935802Z digest=sha256:28f983bf76de48522f6419f896e2bb94488abda114a6cc222f600776c12e22ce

Observation f13312fa-4a4e-4185-93fd-866c791bac9f · inbound

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? cites this paper.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.102050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.102050Z digest=sha256:fa870b8fc8cfc6ed1ee05c73f881bfcc57c5edb8d15711d451844618b6819a38

Observation 4fc1d753-aa5c-4bb7-b320-61d974c04886 · inbound

Promptception: How Sensitive Are Large Multimodal Models to Prompts? cites this paper.

Promptception: How Sensitive Are Large Multimodal Models to Prompts? VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:34:18.347942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:34:18.347942Z digest=sha256:0c94d8c7b7695e704bc5d2a99823da6bb5967d1920c208fa5b7d82ed82d22488

Observation ffe14149-175c-46ca-a48e-28dd91c9803d · inbound

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning cites this paper.

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:24:43.023557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T07:23:53.811794Z digest=sha256:ba494777e37b663ebdc69312e438d57a84c3dc389fed913788eff6d8b322bf2c

Observation 79b286d7-2b67-4b29-9d0a-09514d3b4227 · inbound

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding cites this paper.

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:16:45.138116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T07:00:21.192082Z digest=sha256:cc1f55104952fc9e20b65bd0b29f4940a3027bddf41f4cb4480bb2161da290bb