{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:EDWKWCOKHWYGUWAQAHDG6NHSRE","short_pith_number":"pith:EDWKWCOK","schema_version":"1.0","canonical_sha256":"20ecab09ca3db06a581001c66f34f2893d5230252449fd1c2f7e200ff7e1871c","source":{"kind":"arxiv","id":"2412.15220","version":1},"attestation_state":"computed","paper":{"title":"SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.SD","eess.AS"],"primary_cat":"cs.MM","authors_text":"Anurag Kumar, Gael Le Lan, Haohe Liu, Mark D. Plumbley, Varun Nagaraja, Vikas Chandra, Wenwu Wang, Xinhao Mei, Yangyang Shi, Zhaoheng Ni","submitted_at":"2024-12-03T21:48:08Z","abstract_excerpt":"Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on either a cascaded process or multi-modal contrastive encoders. These approaches, however, often lead to suboptimal results due to inherent information losses during inference and conditioning. In this paper, we introduce SyncFlow, a system that is capable of simultaneously generating temporally synchronized audio and video from text. The core of SyncFlow is"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2412.15220","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.MM","submitted_at":"2024-12-03T21:48:08Z","cross_cats_sorted":["cs.SD","eess.AS"],"title_canon_sha256":"a8b7a55e2065eabd16d7d858ff42c8af10ac2d34dc614f856427da7262a2c41e","abstract_canon_sha256":"0f840d875fe2387f63cc770421bd298e143dd4134b86c0aabe8bf082dbb3f0ff"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:51:59.657618Z","signature_b64":"W7QlFpNMcxTM2fe5zcg5wEEm+0MlgY6048x19Yxvz4SYM5p/Wh5Npq6F88Ysl9kXe/Bhx/fKoK709GEtse8oAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"20ecab09ca3db06a581001c66f34f2893d5230252449fd1c2f7e200ff7e1871c","last_reissued_at":"2026-07-05T09:51:59.657108Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:51:59.657108Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.SD","eess.AS"],"primary_cat":"cs.MM","authors_text":"Anurag Kumar, Gael Le Lan, Haohe Liu, Mark D. Plumbley, Varun Nagaraja, Vikas Chandra, Wenwu Wang, Xinhao Mei, Yangyang Shi, Zhaoheng Ni","submitted_at":"2024-12-03T21:48:08Z","abstract_excerpt":"Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on either a cascaded process or multi-modal contrastive encoders. These approaches, however, often lead to suboptimal results due to inherent information losses during inference and conditioning. In this paper, we introduce SyncFlow, a system that is capable of simultaneously generating temporally synchronized audio and video from text. The core of SyncFlow is"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2412.15220","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2412.15220/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2412.15220","created_at":"2026-07-05T09:51:59.657172+00:00"},{"alias_kind":"arxiv_version","alias_value":"2412.15220v1","created_at":"2026-07-05T09:51:59.657172+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2412.15220","created_at":"2026-07-05T09:51:59.657172+00:00"},{"alias_kind":"pith_short_12","alias_value":"EDWKWCOKHWYG","created_at":"2026-07-05T09:51:59.657172+00:00"},{"alias_kind":"pith_short_16","alias_value":"EDWKWCOKHWYGUWAQ","created_at":"2026-07-05T09:51:59.657172+00:00"},{"alias_kind":"pith_short_8","alias_value":"EDWKWCOK","created_at":"2026-07-05T09:51:59.657172+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.03183","citing_title":"Inference-Time Scaling for Joint Audio-Video Generation","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08729","citing_title":"Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2512.23994","citing_title":"PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2512.13281","citing_title":"VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2512.23994","citing_title":"PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08729","citing_title":"Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2604.23586","citing_title":"Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling","ref_index":11,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/EDWKWCOKHWYGUWAQAHDG6NHSRE","json":"https://pith.science/pith/EDWKWCOKHWYGUWAQAHDG6NHSRE.json","graph_json":"https://pith.science/api/pith-number/EDWKWCOKHWYGUWAQAHDG6NHSRE/graph.json","events_json":"https://pith.science/api/pith-number/EDWKWCOKHWYGUWAQAHDG6NHSRE/events.json","paper":"https://pith.science/paper/EDWKWCOK"},"agent_actions":{"view_html":"https://pith.science/pith/EDWKWCOKHWYGUWAQAHDG6NHSRE","download_json":"https://pith.science/pith/EDWKWCOKHWYGUWAQAHDG6NHSRE.json","view_paper":"https://pith.science/paper/EDWKWCOK","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2412.15220&json=true","fetch_graph":"https://pith.science/api/pith-number/EDWKWCOKHWYGUWAQAHDG6NHSRE/graph.json","fetch_events":"https://pith.science/api/pith-number/EDWKWCOKHWYGUWAQAHDG6NHSRE/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/EDWKWCOKHWYGUWAQAHDG6NHSRE/action/timestamp_anchor","attest_storage":"https://pith.science/pith/EDWKWCOKHWYGUWAQAHDG6NHSRE/action/storage_attestation","attest_author":"https://pith.science/pith/EDWKWCOKHWYGUWAQAHDG6NHSRE/action/author_attestation","sign_citation":"https://pith.science/pith/EDWKWCOKHWYGUWAQAHDG6NHSRE/action/citation_signature","submit_replication":"https://pith.science/pith/EDWKWCOKHWYGUWAQAHDG6NHSRE/action/replication_record"}},"created_at":"2026-07-05T09:51:59.657172+00:00","updated_at":"2026-07-05T09:51:59.657172+00:00"}