Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T11:31:14.532101Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2607.02551.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T11:31:14.532101Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 78780a60-8ca1-4eac-915f-ef6af007fe08 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1bc4a72-369b-4ff6-ae07-21f5c2886aac · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a08d8246-c3bc-4b65-8ae7-3caee423e81d · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Qwen2.5-vl technical report,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef088c35-77af-4ac4-a359-2f001e15ed29 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d29c4c6f-a8c5-41cf-a8fd-174fa1a651a3 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video Action Differencing
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca87c0c8-03f9-4cdb-a11c-ef424b73eaf3 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0e3e16-a9a6-48c6-9d3e-7f2d6d77f769 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b97f4ddb-c151-4608-bef0-c1a52b5e3a8d · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bbc422a-7778-46a2-99b4-7bcf84415e5e · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a106b022-c5b1-47c9-b6d4-09c55db838dd · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5dbbe02-ea69-45f3-b2dc-7a409d0be62b · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Videozoomer: Reinforcement-learned temporal focusing for long video reasoning, 2025
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d352d82a-df1a-436c-ae80-26a90ce30ce4 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53cd208a-dde5-4aa1-bb57-b645cbb90d2a · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4fd2406-8441-4541-b8ed-3e64849db6a8 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04019fa-cc50-425a-abd6-4c7329306b0a · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a01f791-16b4-4d15-adc3-ebf0ed53f88b · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5516c8da-5110-4983-be32-5fbdc10b7e64 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ae4074-eea7-41a1-9e60-87eb220447d9 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb887c39-69da-4c78-8382-a63b3c42305a · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a7288c-ee21-4334-bbbc-cb25f96b79de · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Learning to describe differences between pairs of similar images
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aef92ea5-f003-4b6d-b31a-5aeceb38e04e · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69bf5080-dc00-4ba0-84d5-5a843e09ea05 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Llava-onevision: Easy visual task transfer,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69696830-e34f-4184-8acd-9860b8ae4ae9 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 971a6c9f-fbd1-41c7-b3ac-93c310bbd9d9 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d3bd96-7f8b-4075-beb0-c72b98c8b360 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences VideoChat: Chat-Centric Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49ae5b0-bf28-473c-95ef-70c91b5a6fef · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Llama-vid: An image is worth 2 tokens in large language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 653e810a-99a8-4e3e-bae9-96973078926a · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Evaluating object hallucination in large vision-language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c89af71-099c-452f-acd6-5ba2254169c5 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-LLaV A: Learning united visual representation by alignment before projection
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb37579-304c-48e9-8eab-0547dacea3e9 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences doi: 10.18653/v1/2024.emnlp-main.342
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b06c5fe-346d-4cc1-94ea-82f67970a17f · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Visual Spatial Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b06e165-4617-4f74-95fa-4fce090b5004 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences World model on million-length video and language with blockwise ringattention, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92382b7b-d5b8-4f02-9224-a9d46f203cbf · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Visual instruction tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd71d87-d6b2-4c33-944d-7956f6f3ff8a · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c889a346-046d-41b0-ab11-1c6bba6027ac · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences ISBN 978-3-031-72658-3
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9262eca8-8c79-4650-833d-aa859f92909e · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35905029-b4a9-4e8c-a2d4-1d81f6886e68 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-ChatGPT: Towards detailed video understanding via large vision and language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3998862a-bbb5-4c75-beea-34963e447f6d · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 147703cb-67cb-40e8-bc8c-067e8b137984 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences GPT-4o System Card
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8e74dd-013f-4b03-bd32-3d06e8a564f1 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Training language models to follow instructions with human feedback
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64ffb21e-4e02-4042-ab36-d6ce7d9f58e5 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Direct preference optimization: Your language model is secretly a reward model
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e22faf1-5308-4869-a1b5-4ca00032fe8c · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a14b41-c897-4be9-80c9-0050ab0517fe · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Proximal Policy Optimization Algorithms
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6077b994-5f1b-4684-87c4-5d5d9fa18aa7 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b204ed7-b544-4d00-b49a-40961db8e10f · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 280dd3ea-ad13-4f53-a439-aadb88274f05 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Reason-rft: Reinforcement fine-tuning for visual reasoning of vision language models, 2025
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a02848cd-f8d5-4f17-a108-e71239d64011 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94aae5d9-0598-4d49-b775-a0ba5c644feb · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c4639ee-9939-43cf-8533-2a3508e43a68 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2992964d-a258-4a76-b9fc-2f5902d895f2 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences LVBench: An Extreme Long Video Understanding Benchmark
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3d4dad6-963e-43b0-a461-ffe4a4862636 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a18514-9a7c-40d2-8ab9-2ef60add99fa · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Internvideo2: Scaling foundation models for multimodal video understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 905c145c-42cf-47e9-82ca-ac7190c80433 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Blaschko
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a41d2246-b9ea-4d8b-a2e1-83d53d033fdb · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef1e0dc8-d15d-4d2e-a27b-366e4d2c0d4b · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Vidic: Video difference captioning, 2026
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce54db9b-5438-4540-9a1b-cbffd719bea6 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe90e30b-33b8-4a4d-97d6-38e4e744237e · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35d974e9-fbad-4f29-9edf-19989a11d0e1 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Perception-R1: Pioneering Perception Policy with Reinforcement Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68362232-eef0-45a0-8d7d-117dbab3cd1e · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences CoCa: Contrastive Captioners are Image-Text Foundation Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202374ef-509e-477e-b184-11783a55c022 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Mm-vet: Evaluating large multimodal models for integrated capabilities,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07445521-f978-4da4-aab1-3523623cc1f3 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e07ea42-60f6-4020-a755-7c31ab44f2d8 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-LLaMA: An instruction-tuned audio-visual language model for video understanding
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df954dac-3780-4aa4-a134-d70c850f9352 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c170f5-5c23-4ce4-8034-1b5a9c945dfb · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f85be1a-23a9-4c37-ab92-71f14c27d7e5 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365b1ffa-f52a-4c04-97fa-2d4b92d5cb56 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MLVU: Benchmarking Multi-task Long Video Understanding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae38da21-5a3f-4ea7-91e8-ec12a6aa5048 · outbound
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.