Pith. sign in

Paper Citation Record · LEDGER

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

As of 22 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 4 inbound Pith citation observations for arXiv:2602.08711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.08711 v3

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:15:11.601779Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:02:03.266407Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T17:40:00.898211Z

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c776e3f-1171-4350-aae9-9bc360072212 · outbound

This paper cites • Each minute usually contains4–5 segments, but prioritize the video’s logic over strict numbers.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions • Each minute usually contains4–5 segments, but prioritize the video’s logic over strict numbers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.405612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.405612Z digest=sha256:b0391bfb5eed17f5a95a190b9760cafcd5e308e833c951c71c40ff984adb7267

Observation c7a12329-78df-4c02-a3f0-2b912ae77545 · outbound

This paper cites • Includecharacters, actions, objects, emotions, and scene detailswhere relevant.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions • Includecharacters, actions, objects, emotions, and scene detailswhere relevant

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.419923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.419923Z digest=sha256:7dcf433179adf83ed85c77ce4c6fb806d883cddae62e14ef76688b4059767324

Observation ebf37857-5121-4598-a9ae-5277186fafd9 · outbound

This paper cites • Timestamp format:minutes:seconds(e.g.,0:00 - 0:11).

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions • Timestamp format:minutes:seconds(e.g.,0:00 - 0:11)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.451377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.451377Z digest=sha256:bceb1e532322d1a96b3c25488fbc33ac2d3d3a2a061fbded33d53d47b9dc9ef1

Observation ed901d18-4d82-40c8-a04b-510d88aaa265 · outbound

This paper cites Each GT caption defines the rough boundaries of a segment.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Each GT caption defines the rough boundaries of a segment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.478148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.478148Z digest=sha256:9dad2c492d9b821443e8f3ae75c9f5adfd85bf3c1d70dc0214af334d6216e6ba

Observation c71922d9-2cfc-4bca-b50e-9b56dc80e979 · outbound

This paper cites Use the video itself toexpand with details: •Characters: actions, gestures, facial expressions, emotions.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Use the video itself toexpand with details: •Characters: actions, gestures, facial expressions, emotions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.509777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.509777Z digest=sha256:2c00767429a71a466a8ec246e64572f070b5ee6bcaffe302a272c93d9e82806b

Observation 163a77b1-b06a-42c1-9fe1-1480feb004d2 · outbound

This paper cites timestamp.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions timestamp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.541291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.541291Z digest=sha256:f7668d79528751fe5e78488d8e76ee4fa30692be4776cb777aed246c3d3a7e12

Observation a576c31c-d18f-463f-b3fd-044dad84c7e8 · outbound

This paper cites an unresolved cited work.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.569122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.569122Z digest=sha256:db03cdcd7466e39549d8fbf5e3d2e62e1dbee87c2addeda07181bd978843d2f5

Observation 8bbc8a29-9465-4f13-b4cf-9b98fa517a21 · outbound

This paper cites by_dim": {.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions by_dim": {

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.601779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.601779Z digest=sha256:0d87c67919d4c024bde628bd55371477be165b9ba1c01d4f91b63956b7df0a2c

Observation 95218483-f9a3-4ebb-b2e0-62ea836414ee · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.377213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.377213Z digest=sha256:14b87005b6a0d12f553edc9facfdfaf88a45aa0b61d1c9085c61711850b71e3a

Pith citing papers

Observation 52056ccb-249e-40b0-a81d-75779896f93b · inbound

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video cites this paper.

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:17:09.422974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:35:23.842691Z digest=sha256:3e6e0d8b6054486574cf591094587c2bc8cd8619add4a9c24f9bd790d823fe0d

Observation 0a8d3301-f527-4e13-bd04-03d87c9e37d8 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:17:09.422974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:459ec65c75c021014f65afeb5169196d2f7dfa5f1104815bf39d61b35a974df4

Observation ae5e7b81-915d-4acf-ac1f-cc23263dbccd · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:40:00.899555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:d740a71cad0453b63bde7d6eb73829a113f20a650ff262a9d82bcc832a4b35dc

Observation 7deff422-86ba-4f42-a199-15919089ae58 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.266407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.266407Z digest=sha256:913a2f3fdc0215a497cd544c889e6738643ac8f8c35861b65d8f3bbede3f6ff8