Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:46:19.243232Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2507.10109.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:46:19.243232Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T06:04:29.936928Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T06:04:30.077501Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 37ec9f19-0267-400f-8dc4-186cffcba3fa · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33e3d267-b0bd-4db1-9e4a-d1cc222ff02c · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 279008d1-a795-4170-97b3-afdac7bdc411 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6801440-a98a-4f92-acfd-75a1da97f3bc · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b2da09-6e34-4360-a3a7-76032610e109 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb1a958-27a4-49d4-a92e-1e2e09f2bb59 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 713fd661-c934-40de-b955-c4023f195efc · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d79e06-cee1-4956-9569-cf243ea2cd07 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af46798e-68a2-4bfc-a683-538e14f5b39d · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3539a8e8-b6b9-422e-b7e8-535a6d171926 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd1e16fc-e7a3-459a-8aa2-8b710ccffd68 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05516ffb-ffcb-411c-b448-4b524f94e2d8 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Russell, and Andrew Owens
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7500dc1f-d9f1-4e22-816f-8eff8d1f2d58 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f60c6bcb-e0e9-4e14-a6af-32311c28ccdc · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da33c1e8-ba39-4f81-b34e-d7a3e86c388c · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93b9eb47-acfb-442e-98db-12b9101fda90 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e0bb405-fcd7-4792-a562-d92b225e889c · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c03ef0c9-698a-47ef-b11b-dc059e450910 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Hawley, and Jordi Pons
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abba1354-4e34-4433-88c0-93456f3e50af · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Contrastive Audio-Visual Masked Autoencoder
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b38237-e27d-4dac-af82-ec1e839fb9f0 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ff0220a8-2b66-47d3-abd6-7c0f316c06f6 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eda6f2e-526f-4acb-b809-f56b05765310 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ee185fd-e353-4680-b406-11bc1b508eea · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce3f8763-5543-4aa8-a484-95c08a948cd2 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7780ee1-ce41-48bc-aa46-3f2993a269f1 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2292ea90-dbba-4505-9b79-ede17c98baab · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03647407-1f64-4c07-840f-723e7258cd1f · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1183abe4-53c7-4251-8fd3-81917221eb3f · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb3d7ec8-f2d5-4e92-acaf-e0b4150495b1 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Salik, Rajiv Ratn Shah, Yifang Yin, and Roger Zimmermann
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfe93e57-9358-4e6b-ae3c-1fd69c32cc39 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b4caf5-dafc-4d0a-bb99-b0f39eb64b0c · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f9c274d-f627-426c-829c-1842c25a0dfc · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f69a0db-57dc-429a-8de0-b4d50dccc736 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mandic, Wenwu Wang, and Mark D
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9cd59548-046e-4908-977a-3d14562cecff · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Raghavan, Gavin Mischler, and Nima Mesgarani
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3194c8a-04cd-421e-adef-e2557f37f3b6 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Autoregressive Speech Synthesis without Vector Quantization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99e73ac7-28ec-4a05-8240-968939682a8b · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3af9a247-6149-4362-b3ec-3b5024433dbf · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42533c02-709b-444c-ba82-f108d0b1d43e · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c3fb29b-10e6-4788-9da2-c122d56c57c6 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf519b0d-1df2-45f7-a72c-fb11017f577f · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d151fd1-b8cb-4da5-b6dd-dafea2042c69 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5aff840-67a9-4dbe-861c-bc0ed3ecf6f6 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c778f8fa-b4e8-4720-a550-e1c19f05d7d0 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13942b3b-21ae-497b-8103-3869f8f17fe3 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis LLaMA: Open and Efficient Foundation Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aadd354-21aa-434d-b3e4-e651d848f1c5 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mel-Band RoFormer for Music Source Separation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d1a0b5-080a-4f04-8e92-8ef224a181b4 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25de76f6-9bd3-4faa-b791-113c7b55626b · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3869cc5b-f98e-484f-8fbd-91cb5e62436a · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91a24ba9-2e29-43ef-827b-95867ffd5fa3 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6e79082-fc53-44c9-8f1e-6ec750618d11 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5232a114-2d96-40a1-b022-b42e42fff0c2 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3fbe57a-f389-445e-a3c7-8911c34f56f5 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Qwen2 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eda949d-c7f0-4718-9603-63ce8d097b60 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66285bcf-7fed-49f0-8900-4f4d5a0d744f · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7417fdab-75af-4ed8-918e-d769dee767a7 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f08d860c-4432-4423-a6b6-64609e10eb18 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e39c3107-8a78-4cd1-b18b-1e2b4f16793d · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b4e0a7c-e4c5-4cd3-bead-1dee93540884 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 952abe25-feca-40f1-a9b9-985be6455714 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bed8f7c-9191-4d2a-b1a7-59f1dd70c6e5 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 946540d6-4f50-4a0a-bf43-337744e816ff · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ecb8ab1-8022-41c7-bd52-9c10e0998921 · outbound
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9f73405d-b2d7-4205-bea6-b06ea22e0311 · inbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.