Pith. sign in

Paper Citation Record · LEDGER

SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2505.12589.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12589 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:06:23.701939Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T22:44:01.321727Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4b115999-15e5-4575-8c98-d5ceb4a12853 · inbound

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models cites this paper.

STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:23.701939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:23.701939Z digest=sha256:1ca400fd280c89b197d639a2a1fd13d60e4302a6c9352df6ca06a780b9115a62

Observation 239e5431-98ec-4e19-9a00-164b5ca3ca63 · inbound

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos cites this paper.

ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:08:56.600848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T03:05:42.773193Z digest=sha256:33ee5a3d65094d21b529873f7622faa1c3ef9b95f7aa58aee23fb5ddbc7dd902

Observation b1c66789-5806-49e3-8774-61e562d6ebc8 · inbound

CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning cites this paper.

CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:57.908140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:09:57.040135Z digest=sha256:3a41ffa744a0757a484dfe105776bf9f6062192247bb4678daa5364c0e1857a6

Observation 3eb130d2-a3de-47cf-ae0d-9389acf53a4b · inbound

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks cites this paper.

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:36:13.989290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T07:36:09.562362Z digest=sha256:c20fd203502b5f57d5fced4457697a4f7f70a06bc1d23854cd45a3639dc06d98

Observation 29ab977d-45d4-4862-a0a2-88ef6011c0f5 · inbound

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks cites this paper.

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T16:17:21.074502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:17:21.074502Z digest=sha256:bd315b4f7b55f2e467c958907ed5c42d36ea473c811d54a65ad8424d8a632f70

Observation d59e0931-11dc-4425-8c61-dece5328e698 · inbound

MetaphorVU: Towards Metaphorical Video Understanding cites this paper.

MetaphorVU: Towards Metaphorical Video Understanding SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.323382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:43:20.101830Z digest=sha256:bde7fd18f4811e2aa5bb216f587d0e64d1922ba1b095ea63e93582b0716f2d8d