Pith. sign in

Paper Citation Record · LEDGER

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2509.02807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02807 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:27:47.145067Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4018e7bd-4846-4503-be68-928453b8ea2f · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.077290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.077290Z digest=sha256:d659715278326880444f3c593644fabeaac86490ae7c8c1a1cb42634d2646f39

Observation 2c9655e2-ebaa-4cb8-84bb-095876763e1c · outbound

This paper cites MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-05T11:27:47.523163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T11:27:47.082104Z digest=sha256:158868c52fa217092c2a7fe7145f1baf93e675a8b4dbd66361805e833d3c7861

Observation 004ce299-6190-41c0-9871-e2cded83df37 · outbound

This paper cites Visual instruction tuning.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Visual instruction tuning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T11:27:47.602290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T11:27:47.092352Z digest=sha256:f0083b0633d9e7044370127ec250775810eef97a0b90c67a452381f5a298eb0e

Observation eba6ded5-0332-4278-a22e-9cfa58f8dd31 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.097393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.097393Z digest=sha256:8db65990bde40eb7d960a05ca7820ed12e1fbcc189d022b2b44a7d6d3e040c42

Observation cbaa8710-72bb-4e3c-b023-83dabb5b3358 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? The 2017 DAVIS Challenge on Video Object Segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.106806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.106806Z digest=sha256:cc8bedea9bc265eb2b9c248c36b1948b2e4a6d5b8823c3a7e9287f10c234a1ae

Observation b6705038-e826-46a5-93ff-615f14bea5d4 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? SAM 2: Segment Anything in Images and Videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.111735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.111735Z digest=sha256:e105fc39d0f4fa65ec662e613761ce8dd05e8db8a0d613c80eef4fd360f6435f

Observation 85c6c5a5-706f-4501-b376-a90ea40fade8 · outbound

This paper cites Seeing the arrow of time in large multimodal models.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Seeing the arrow of time in large multimodal models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.126396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.126396Z digest=sha256:430cc611bbd927177b39be90e468e239fb4ab06bad0bd15bb7a07f5a786fdf1b

Observation 06762684-0978-4228-8d19-64855a174207 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.135264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.135264Z digest=sha256:7cdccba9a28560442a6557ec5dbd93028053c7f987c70ffeca9c5944df1e90b3

Observation 48401b11-edf3-420b-9e40-32acd7052281 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.140036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.140036Z digest=sha256:7ebe06d3159d6c5cb1b755ad1e86fcee3fe0b0c44281a48b0fd24bfed109d6d5

Observation 4ebc1d82-7cb1-49f5-9264-9aaf78b8426d · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.145067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.145067Z digest=sha256:a60e3282346421d964af56b777e52a26f18ddf1bd89a33234b29ee8476e0f22a

Observation 6e33aa3b-2862-416a-ab2a-c549c15742d1 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.130766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.130766Z digest=sha256:0de44d1bc4ce77aa2c2ad7689e18d854950973d55371f7f7a6cade2ae7fe41d3

Observation 95e3f717-51cf-4246-8c63-32409f313d6f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.086884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.086884Z digest=sha256:1837dc840c206d6079240906a4fe8a885b9b729ca55a367613f2d62054524fa4

Observation 517b939d-d764-4c18-9746-ecb1f30d78b0 · outbound

This paper cites Pixfoundation: Are we heading in the right direction with pixel-level vision foundation models? arXiv preprint arXiv:2502.04192,.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Pixfoundation: Are we heading in the right direction with pixel-level vision foundation models? arXiv preprint arXiv:2502.04192,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.117264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.117264Z digest=sha256:edf626d930ac2ea171a55f4704aac2cc461dd8c7d65d1ccb6b1def25e547bebc

Observation f13312fa-4a4e-4185-93fd-866c791bac9f · outbound

This paper cites VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.102050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.102050Z digest=sha256:8af3b6c9b6773a2425c147f0389e874157d14ef48ac7c6a824dbd8683fa2623f

Observation bc7b336e-7c28-4247-84c1-c97ccebf12f9 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.072265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.072265Z digest=sha256:9c92c96f2363fac8ccc97e415f3e941a1fb15bacf28ebc8b2f004914ced0ac68

Observation 21ce8e8d-0df4-4b25-bdb4-0b5dcea8cd19 · outbound

This paper cites Qwen2.5-VL Technical Report.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Qwen2.5-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.067394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.067394Z digest=sha256:e83a4e39963beda9ad02cb7ba2343280c2339011b7cc420dc7840b4973ffee70

Observation 0f7d0fa9-84e4-4575-b4cc-093f85176bcc · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.062248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.062248Z digest=sha256:588f383d5e13edd04a1e166d5c659fe35183e0394e00939194bc0c3e9001ffec

Observation 6bafd136-ad1a-413c-96a7-1c0255ab1620 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T11:27:47.121911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:27:47.121911Z digest=sha256:b26cd89d5c058449d4ea9182be2020405c9f214bd7d6b7a66c530aa8ee096745

Pith citing papers

No inbound Pith citation observations are available.