Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T11:27:47.145067Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2509.02807.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T11:27:47.145067Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4018e7bd-4846-4503-be68-928453b8ea2f · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c9655e2-ebaa-4cb8-84bb-095876763e1c · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 004ce299-6190-41c0-9871-e2cded83df37 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Visual instruction tuning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eba6ded5-0332-4278-a22e-9cfa58f8dd31 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbaa8710-72bb-4e3c-b023-83dabb5b3358 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? The 2017 DAVIS Challenge on Video Object Segmentation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6705038-e826-46a5-93ff-615f14bea5d4 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? SAM 2: Segment Anything in Images and Videos
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c6c5a5-706f-4501-b376-a90ea40fade8 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Seeing the arrow of time in large multimodal models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06762684-0978-4228-8d19-64855a174207 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48401b11-edf3-420b-9e40-32acd7052281 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ebc1d82-7cb1-49f5-9264-9aaf78b8426d · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Apollo: An Exploration of Video Understanding in Large Multimodal Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e33aa3b-2862-416a-ab2a-c549c15742d1 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e3f717-51cf-4246-8c63-32409f313d6f · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517b939d-d764-4c18-9746-ecb1f30d78b0 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Pixfoundation: Are we heading in the right direction with pixel-level vision foundation models? arXiv preprint arXiv:2502.04192,
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f13312fa-4a4e-4185-93fd-866c791bac9f · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7b336e-7c28-4247-84c1-c97ccebf12f9 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ce8e8d-0df4-4b25-bdb4-0b5dcea8cd19 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Qwen2.5-VL Technical Report
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f7d0fa9-84e4-4575-b4cc-093f85176bcc · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bafd136-ad1a-413c-96a7-1c0255ab1620 · outbound
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.