Pith. sign in

Paper Citation Record · LEDGER

Fine-Grained Captioning of Long Videos through Scene Graph Consolidation

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2502.16427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.16427 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T23:09:01.340759Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:11.312748Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e9841d58-d9f5-4f29-93ec-bd548fc09fd3 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Fine-Grained Captioning of Long Videos through Scene Graph Consolidation

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:27:15.765503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:fbeb8271a391eda524585d858599eabbd470b7502c10cf0c52d8396422dc777d

Observation 4a3c1ceb-e825-480c-903f-222887dc8ae7 · inbound

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales cites this paper.

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales Fine-Grained Captioning of Long Videos through Scene Graph Consolidation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:09:41.275122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T12:09:01.026544Z digest=sha256:e385bec45f0420445f6d1fc373691dd7cb36c817d72633934e5900e71a5149a1

Observation 0ed5d67b-c00e-460b-b136-a6212230d909 · inbound

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs cites this paper.

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs Fine-Grained Captioning of Long Videos through Scene Graph Consolidation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:11.314032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-25T20:58:01.925342Z digest=sha256:3d26b2c1e8293f61cc10cff307dd6f47c3bcb5a667b0c75b1839476f53308fbb

Observation 85cc3e89-454d-48f7-9c21-45b5ea608467 · inbound

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs cites this paper.

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs Fine-Grained Captioning of Long Videos through Scene Graph Consolidation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.846566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T23:09:01.340759Z digest=sha256:d11850be3d7b617437b471e6dfd035bc464c4463ba70e00ec24955d33d81ac32

Observation 869836ed-d635-412a-8e18-26e31c525685 · inbound

Learning to Evolve Scenes: Reasoning about Human Activities with Scene Graphs cites this paper.

Learning to Evolve Scenes: Reasoning about Human Activities with Scene Graphs Fine-Grained Captioning of Long Videos through Scene Graph Consolidation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:32.383264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T15:06:06.430083Z digest=sha256:d12d66e1425721f9e30fc322cac6ab18fed70051cb7964eb2461a6f4575929e9