Pith. sign in

Paper Citation Record · LEDGER

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2412.15220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15220 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:05:47.713763Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:40.931162Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dea83d22-1d9d-49da-ac40-e2eabc4cca0e · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.713763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.713763Z digest=sha256:7c911ce878d103dfb3d58a4dfe1e9d8370b757f87e53ec006ab87549b8301b9d

Observation 42b6baa9-358b-4bba-97f5-d90724595e5d · inbound

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction cites this paper.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.383496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.383496Z digest=sha256:d078e5a06ba797678e5d45a97e6461a6a5798ae61d8e6bbc49f1b8dbec5e3800

Observation d7a23a1c-972f-4725-8ad4-3706843219ac · inbound

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? cites this paper.

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.544967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T21:46:43.353305Z digest=sha256:2fc4c845323829efd0c2d14fea7df398d9f59cd59ba8f51f91315685191c8063

Observation cfb1d7bd-7733-41fe-aa49-4ab177adeeae · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.933593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:93072103727c463e8512f28b7de818d48ea67870d268921c3e11894599990db2

Observation 2d635f72-789b-474c-9e74-2db868550dc6 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:14:15.274194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:949c899f604aab5c1d382aa82eaaaa2fa13fd7cc0630f940baba441f10d2b15c

Observation e0f11733-31b9-4529-adee-38832a8e15ee · inbound

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling cites this paper.

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:06:14.373792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:44:36.000353Z digest=sha256:7c5a94e01448e4593c1c0e35800c46bd284257f4ead3868909b4cb8b8be33d76

Observation 1574171a-d19c-4fb8-a268-7320017c7ec7 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:45.368384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T02:28:14.734682Z digest=sha256:3a4da4f8384b110550b9c33b354030ff312018e7e3e8e5ccb2e40b09d947bc71

Observation f4d24b73-9292-4f2d-87f7-8cf8abc57f52 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:35:07.852450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T23:26:46.077894Z digest=sha256:85a3bb19e06ddefbf9887fa5f93452e5bb61725b545239e11e6902f304831756

Observation 01c6ea4a-f099-458b-8219-85773421ecd2 · inbound

Inference-Time Scaling for Joint Audio-Video Generation cites this paper.

Inference-Time Scaling for Joint Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:40.932702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T07:49:20.194990Z digest=sha256:cbb1cf174e1b19160d9f30d7b04b3024c5f985a5b0c0ad89ea9efa49f8652efd