Pith. sign in

Paper Citation Record · LEDGER

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

As of 7 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2607.02551.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.02551 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T11:31:14.532101Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved64
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78780a60-8ca1-4eac-915f-ef6af007fe08 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:3a2db5a22f1219f450ee1785e2cf783f7654d688b2ce95a4018923ecd2cc35cb

Observation b1bc4a72-369b-4ff6-ae07-21f5c2886aac · outbound

This paper cites Qwen3-VL Technical Report.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:b54df71c0952d6729f9649648943a492d2c19ca96550e6869ba45f6809bc5675

Observation a08d8246-c3bc-4b65-8ae7-3caee423e81d · outbound

This paper cites Qwen2.5-vl technical report,.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Qwen2.5-vl technical report,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:a4a175cd06bad3c04e759c5e0538a567e4cf65cc0dcb18796cee44732cd7c071

Observation ef088c35-77af-4ac4-a359-2f001e15ed29 · outbound

This paper cites Qwen2.5-VL Technical Report.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:edc97e96b248aa1631a5e3c1c4a1d2cc6ec1c501ca4daea2b590dc8901cba0f1

Observation d29c4c6f-a8c5-41cf-a8fd-174fa1a651a3 · outbound

This paper cites Video Action Differencing.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video Action Differencing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:97557393c9be25235af86235ff3772096fd5609d54b41ce87d34258624684801

Observation ca87c0c8-03f9-4cdb-a11c-ef424b73eaf3 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:92b2bfdd4037e1207300e535744299af0d7ca22d49e6f07aee9cc96b323ff68e

Observation ef0e3e16-a9a6-48c6-9d3e-7f2d6d77f769 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:d5a347fd562e661923d5d015e451e770f767a2fb40cb7bd659dd8fc6145e25f8

Observation b97f4ddb-c151-4608-bef0-c1a52b5e3a8d · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:fe0f040eb8c9cb4cfc97e375ac03bab7bf4966ee55be16875eba058f3eb87eea

Observation 5bbc422a-7778-46a2-99b4-7bcf84415e5e · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:2dda4d80bfb6a6815a33a9bd53547e6325f2798b628b8e8c0b804156d5a42362

Observation a106b022-c5b1-47c9-b6d4-09c55db838dd · outbound

This paper cites an unresolved cited work.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:eee29d5076567bce2ca72af409c93d4f266f93c0b1b00ed1d63d58ba175809f9

Observation e5dbbe02-ea69-45f3-b2dc-7a409d0be62b · outbound

This paper cites Videozoomer: Reinforcement-learned temporal focusing for long video reasoning, 2025.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Videozoomer: Reinforcement-learned temporal focusing for long video reasoning, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:3e009208b27ee11e45a8160c8670c98ae1e0980fecbd34c46fec4c38dd910e2f

Observation d352d82a-df1a-436c-ae80-26a90ce30ce4 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:7787fabc6109abc20a7e04c5bde3190b6c6cdaadb9a4d0849315672c236a4d4b

Observation 53cd208a-dde5-4aa1-bb57-b645cbb90d2a · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:029292ca80e780cd5578dba82fc7da7624f0a9a38915d3196079a3eca45ff51b

Observation f4fd2406-8441-4541-b8ed-3e64849db6a8 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:45a5c7979fc3e2a86200a6a7d2cbbc1c46cfb2635f84e6973cac81ab270c04ef

Observation e04019fa-cc50-425a-abd6-4c7329306b0a · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:c9be9b4d67d319b70f65c1c74ce3235bd1dc71528776d25da0f44e4aca2128b9

Observation 5a01f791-16b4-4d15-adc3-ebf0ed53f88b · outbound

This paper cites an unresolved cited work.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:ee248d490021d149ef3f58b92b24911a79b9bcb807669ccdd8564275bc011575

Observation 5516c8da-5110-4983-be32-5fbdc10b7e64 · outbound

This paper cites Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos,.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-mmmu: Evaluating knowledge acquisition from multi-discipline professional videos,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:53a266a060a24bacfe16383ca476b3f03c9beab3c738f88f47c4dc997862c360

Observation 56ae4074-eea7-41a1-9e60-87eb220447d9 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:47bc83b11d4a1dfeeedc3f9e4757031961fd5e27a0de6d1322d5a367bf3f4c05

Observation bb887c39-69da-4c78-8382-a63b3c42305a · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:b1a90ae19ba5c74d60748a0912a008463b4f3df3e636c5d3d8b0e1054da1a654

Observation d4a7288c-ee21-4334-bbbc-cb25f96b79de · outbound

This paper cites Learning to describe differences between pairs of similar images.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Learning to describe differences between pairs of similar images

Reference 20

Resolution
verified exact
doi, observed 2026-07-12T11:38:47.321100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:25b5dfc3897ae77856281b9c6a43c73932fbf82a0e7838d9f322f2d92ae840e0

Observation aef92ea5-f003-4b6d-b31a-5aeceb38e04e · outbound

This paper cites CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:ea6d850e47b1a3f00d48a4ace294f857a071575bdf6ea90826e08b9e20bba07d

Observation 69bf5080-dc00-4ba0-84d5-5a843e09ea05 · outbound

This paper cites Llava-onevision: Easy visual task transfer,.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Llava-onevision: Easy visual task transfer,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:1ee816f48557142640926de548d2ee909ce8eeb7752affeba3230b3c8316b3a1

Observation 69696830-e34f-4184-8acd-9860b8ae4ae9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:6e5964fff711957ea11999a7281d28d3b5c489e6972071537635f1a64c3eb369

Observation 971a6c9f-fbd1-41c7-b3ac-93c310bbd9d9 · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:36ca8029a8963ec253c5098362ea4be271d3e98f1f8cb95f83d0f02b275bfff8

Observation c2d3bd96-7f8b-4075-beb0-c72b98c8b360 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences VideoChat: Chat-Centric Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:dff174eca14a59c665d70184ee4b973788d7753ce1a84a92ca76d5488262c644

Observation c49ae5b0-bf28-473c-95ef-70c91b5a6fef · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Llama-vid: An image is worth 2 tokens in large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:1c37a056bfc23c9e8a0f7ef1fcf0d2049a62a4f25aaac9a9cbfd89a0fe1ca241

Observation 653e810a-99a8-4e3e-bae9-96973078926a · outbound

This paper cites Evaluating object hallucination in large vision-language models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Evaluating object hallucination in large vision-language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:394d6751169a704e5086a3993d073c4ad415b5a73914002c551a265ff068a69b

Observation 7c89af71-099c-452f-acd6-5ba2254169c5 · outbound

This paper cites Video-LLaV A: Learning united visual representation by alignment before projection.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-LLaV A: Learning united visual representation by alignment before projection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:c0891fc2827ae463af0815c59f86c216dbc256627512afacb657e4fbfb4e7259

Observation 7cb37579-304c-48e9-8eab-0547dacea3e9 · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.342.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences doi: 10.18653/v1/2024.emnlp-main.342

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:fe20fc3d30919cd5548a63837b8da1a962c24bf930849211dac1482758545755

Observation 5b06c5fe-346d-4cc1-94ea-82f67970a17f · outbound

This paper cites Visual Spatial Reasoning.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Visual Spatial Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:e040cc99e24e9170e35dd7d0aa8da42bd9f7dffb7283b39f536c1f7ae0400430

Observation 6b06e165-4617-4f74-95fa-4fce090b5004 · outbound

This paper cites World model on million-length video and language with blockwise ringattention, 2025.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences World model on million-length video and language with blockwise ringattention, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:8be5d15dd1fcfcf28bc00fa5ec34210bc11259a9692f5a60bb36279bc5ecfa1c

Observation 92382b7b-d5b8-4f02-9224-a9d46f203cbf · outbound

This paper cites Visual instruction tuning.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:b996660aeeaff819a785bb9f85217e65a052e218dd9eeb5f7e6790e11678d463

Observation 2bd71d87-d6b2-4c33-944d-7956f6f3ff8a · outbound

This paper cites an unresolved cited work.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:4bc9547b37274a74830416b6eeef5c9bc1f8faabaa533a11706d71d51468bc8c

Observation c889a346-046d-41b0-ab11-1c6bba6027ac · outbound

This paper cites ISBN 978-3-031-72658-3.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences ISBN 978-3-031-72658-3

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:562c1cac99185f2ae121de6ac0e0c7a76dd7d8f0eb6ef73ed8af640bbf6bf6db

Observation 9262eca8-8c79-4650-833d-aa859f92909e · outbound

This paper cites an unresolved cited work.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:b89ad8c0974b61a4374ea2a90271f7fa7371af31bca4dd844ae5206752b7f8f3

Observation 35905029-b4a9-4e8c-a2d4-1d81f6886e68 · outbound

This paper cites Video-ChatGPT: Towards detailed video understanding via large vision and language models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-ChatGPT: Towards detailed video understanding via large vision and language models

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:58635b3c5710df324f78394e07be4f3ac08e1751d75007331067fbdaabddb71e

Observation 3998862a-bbb5-4c75-beea-34963e447f6d · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:de51cc99b857768c1c1c4e4d5f303fcede25b5b7b818443d1c38a8191093b4e0

Observation 147703cb-67cb-40e8-bc8c-067e8b137984 · outbound

This paper cites GPT-4o System Card.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences GPT-4o System Card

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:aa835e41bb54be9d361fa2f0781116ad87937c5b5b15ba60e265171f73058f1a

Observation 8a8e74dd-013f-4b03-bd32-3d06e8a564f1 · outbound

This paper cites Training language models to follow instructions with human feedback.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Training language models to follow instructions with human feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:f3259e83426f72c907d50d93c78b4bd6f77a4aaf478ef24efb975f2d8f1d3693

Observation 64ffb21e-4e02-4042-ab36-d6ce7d9f58e5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Direct preference optimization: Your language model is secretly a reward model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:ba8f0535bc50131e10a32745990c8d1a21bd24e2159e16c2d204fd9badbc725d

Observation 0e22faf1-5308-4869-a1b5-4ca00032fe8c · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:39930770a98982e357e887112f914e18e3f46f238037e4186d04b901cc56a8ed

Observation 41a14b41-c897-4be9-80c9-0050ab0517fe · outbound

This paper cites Proximal Policy Optimization Algorithms.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Proximal Policy Optimization Algorithms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:88c50fe16979e659f0c2c561ada96c91effedb704eaea62e7a40fcd6eb8e5b74

Observation 6077b994-5f1b-4684-87c4-5d5d9fa18aa7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:f3425572b84a460f43a3427ae7250a8690a35cd28079f9568d53aec2c5b0fdc2

Observation 1b204ed7-b544-4d00-b49a-40961db8e10f · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:086e14ed56691f56b592ad7dbd5bc6aa57980e900999e8c6c1cec448aa2dd904

Observation 280dd3ea-ad13-4f53-a439-aadb88274f05 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning of vision language models, 2025.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Reason-rft: Reinforcement fine-tuning for visual reasoning of vision language models, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:a78c6bb4b82f1403f3d4b7ff82cadbb58abab4207bea8871c5304c335dc8d66f

Observation a02848cd-f8d5-4f17-a108-e71239d64011 · outbound

This paper cites an unresolved cited work.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:d09d8f189389afd548e170cac6bf7750dae22f57b9ae5b4978d705957cb2ae18

Observation 94aae5d9-0598-4d49-b775-a0ba5c644feb · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:e6a261759d44240639e6777d7b85840c51792ea11ae342b671917d437838e713

Observation 1c4639ee-9939-43cf-8533-2a3508e43a68 · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:350a09eb30474def985451023757b00c6f1ee318074b180ec5488db15cbd7629

Observation 2992964d-a258-4a76-b9fc-2f5902d895f2 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences LVBench: An Extreme Long Video Understanding Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:677d22cfef27b65cd7f1abf3ac6062b5be967d0e723eec73eb050389eb3bd5c1

Observation f3d4dad6-963e-43b0-a461-ffe4a4862636 · outbound

This paper cites ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:ac07ff430509350130347ed1fcdd9ce00e3451bb550b8754f169b838e86ac784

Observation 92a18514-9a7c-40d2-8ab9-2ef60add99fa · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Internvideo2: Scaling foundation models for multimodal video understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:2f06ecac254a82a330259ea9a51928b9daf52cf5bd3572d754581242962060fe

Observation 905c145c-42cf-47e9-82ca-ac7190c80433 · outbound

This paper cites Blaschko.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Blaschko

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:8755f70ea7c05c13eeacbb91524a9b878dbe04bf5422a736c7cdf8c16767bbbc

Observation a41d2246-b9ea-4d8b-a2e1-83d53d033fdb · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:8e472ff96bd7b297be02317984889d0243a2f8e3f5efd1f20dcaea42e995c1ac

Observation ef1e0dc8-d15d-4d2e-a27b-366e4d2c0d4b · outbound

This paper cites Vidic: Video difference captioning, 2026.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Vidic: Video difference captioning, 2026

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:fa3650c90e208c644886672867b6f135faec413c1dea1d28bd632c8dfe360707

Observation ce54db9b-5438-4540-9a1b-cbffd719bea6 · outbound

This paper cites TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:256c76515bd62a4aa1c7f2f74565d82363c2c2e14ac7442c136fee0ee13f365c

Observation fe90e30b-33b8-4a4d-97d6-38e4e744237e · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:a4354c0ac280326fa957fa7a8a8c6ee00d8437d2efc6970f59e30dbeb6475e83

Observation 35d974e9-fbad-4f29-9edf-19989a11d0e1 · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:211057a6a3f1aa218cc3f20635b4f475cc7f5bf6229d6e711d1eefbdac5c7815

Observation 68362232-eef0-45a0-8d7d-117dbab3cd1e · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:6bf78694ea02cdd2a98d3b27e66c58730090063090297a89b98845f2a99dbb3c

Observation 202374ef-509e-477e-b184-11783a55c022 · outbound

This paper cites Mm-vet: Evaluating large multimodal models for integrated capabilities,.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Mm-vet: Evaluating large multimodal models for integrated capabilities,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:2759d6979fc827b7fc8a719dc1453468d37af465cc687702b84d547288745f48

Observation 07445521-f978-4da4-aab1-3523623cc1f3 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:1681e98ea70669621caef73815aa9c8c08ecf84ad854e35f4cfc0d96909ee44e

Observation 8e07ea42-60f6-4020-a755-7c31ab44f2d8 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Video-LLaMA: An instruction-tuned audio-visual language model for video understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:f8e25f8547c1edf827cfe7cf00bf2ce17932476ea170c32c5ec216a4870778b0

Observation df954dac-3780-4aa4-a134-d70c850f9352 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:0acb663cf9137eb3dc4b2e1898b611871be3318963c53cac6f072fb29e4746fd

Observation b2c170f5-5c23-4ce4-8034-1b5a9c945dfb · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:45672d47674adf53dd3ab1469a2807184a7419252944daf5b44e49142dd00e00

Observation 7f85be1a-23a9-4c37-ab92-71f14c27d7e5 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:a7fc2d021b3456be7be6f961330142d38e4fc611ddd603abe627fc380b2fccb8

Observation 365b1ffa-f52a-4c04-97fa-2d4b92d5cb56 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MLVU: Benchmarking Multi-task Long Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:05107853023bee7e210bef4d764e1d27350e83cb9179bb30dfbcc17b96ac7895

Observation ae38da21-5a3f-4ea7-91e8-ec12a6aa5048 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:ba317e5de2f024e7fa16907469241589938acf777816e24cd9c65a6b1fdfd100

Pith citing papers

No inbound Pith citation observations are available.