Pith. sign in

Paper Citation Record · LEDGER

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2507.10109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10109 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:46:19.243232Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T06:04:29.936928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T06:04:30.077501Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 37ec9f19-0267-400f-8dc4-186cffcba3fa · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.620782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:12.972404Z digest=sha256:d01c12028484dbaa946e06a9fa3bf6a81d65cdd3ecf064d8c024c7bbecf48503

Observation 33e3d267-b0bd-4db1-9e4a-d1cc222ff02c · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.409135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:13.063080Z digest=sha256:45d6ce1b769ac2fb5a31a3f3a8d02a173a6c2cb34c8d03efa457f9fc83692cb9

Observation 279008d1-a795-4170-97b3-afdac7bdc411 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.181674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:13.181371Z digest=sha256:092f4ae61e45f468814b9405ddd28960350dde4a820e18e971ef73b191246797

Observation e6801440-a98a-4f92-acfd-75a1da97f3bc · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.348025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.348025Z digest=sha256:725dd82c5365d5c9cb06f0ced412201c5daf45e592a864f42717d9dbc12598a5

Observation 20b2da09-6e34-4360-a3a7-76032610e109 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.509047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.509047Z digest=sha256:476c758c31f7486903acd81d96f1a5221d9669212a55b8745df05fe80382762c

Observation 5fb1a958-27a4-49d4-a92e-1e2e09f2bb59 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.008575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:13.584990Z digest=sha256:7a9676c80d679b661b760f2f7c4fb9016bfed23cd16c205c5d7f6bd8ff08991e

Observation 713fd661-c934-40de-b955-c4023f195efc · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.686519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.686519Z digest=sha256:d64e6c13e54e16e2a80a77760f50ff8d1bea1ed25927a87b0b1df85b245433ee

Observation 96d79e06-cee1-4956-9569-cf243ea2cd07 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:25.846602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:13.792235Z digest=sha256:16e6404440e8cf3493c83defd58ac4c2fd3edbc6aefa41c50678a87b23fda9dc

Observation af46798e-68a2-4bfc-a683-538e14f5b39d · outbound

This paper cites V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:46:19.957973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:13.885497Z digest=sha256:54b246465ddf06a1b8c5b95e4d68b9abbd6268bac534fbe20e374fb50145b031

Observation 3539a8e8-b6b9-422e-b7e8-535a6d171926 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:25.693597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:13.957998Z digest=sha256:6fece83906ec6bac55993c4dcd99e9561942a12246deb8bcb4f9568becd054eb

Observation bd1e16fc-e7a3-459a-8aa2-8b710ccffd68 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:25.478203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:14.028538Z digest=sha256:11dfc368adf6e83aea7f3da366ff674e3d056272d0c37eeceef33791f1f6eb2c

Observation 05516ffb-ffcb-411c-b448-4b524f94e2d8 · outbound

This paper cites Russell, and Andrew Owens.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Russell, and Andrew Owens

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:25.270345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:14.097517Z digest=sha256:ac6db4837626e2773386d6daaa82f9d9e14278cbde332ca11834a0c40551dd09

Observation 7500dc1f-d9f1-4e22-816f-8eff8d1f2d58 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.237637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.237637Z digest=sha256:a41d1bbea80695450fe1a51782dacf15fe55e8dce526cd9bdfe7f3cf7ce3fb4c

Observation f60c6bcb-e0e9-4e14-a6af-32311c28ccdc · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.325703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.325703Z digest=sha256:d58786f4862cb2cf47b85ba42509280b5035f05a4dfaa1dab88a07d11c4a9bcc

Observation da33c1e8-ba39-4f81-b34e-d7a3e86c388c · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:24.620552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:14.470174Z digest=sha256:b7e7f941fcc619905cd3284006ffa65b45732808105159065f94f5c9d1d5b5a6

Observation 93b9eb47-acfb-442e-98db-12b9101fda90 · outbound

This paper cites In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:24.828783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:14.399881Z digest=sha256:7612c0bd0ae55ab1aa0a03d3ae4c4127178d5c92bcc73afc06d787e0d92c1a1e

Observation 3e0bb405-fcd7-4792-a562-d92b225e889c · outbound

This paper cites MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.616924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.616924Z digest=sha256:ca92d9b39b4f6940d863106499318e2f1147f913d64e96adc59addc335fe382b

Observation c03ef0c9-698a-47ef-b11b-dc059e450910 · outbound

This paper cites Hawley, and Jordi Pons.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Hawley, and Jordi Pons

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:24.392742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:14.542109Z digest=sha256:9fcf7df5db39a7e1b5f9149be89a771986f28f79ac7b674737edc0c4908b723d

Observation abba1354-4e34-4433-88c0-93456f3e50af · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Contrastive Audio-Visual Masked Autoencoder

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.781335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.781335Z digest=sha256:0dd66d3ab93509e0a2c2713a9e0787641eb907d661fb4b09393019823d295053

Observation b9b38237-e27d-4dac-af82-ec1e839fb9f0 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:24.175716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:14.711057Z digest=sha256:2c9b9ecf72e7a363407676e904dbdd99f0def21ec9aa40ff17dcb3dff7f7647a

Observation ff0220a8-2b66-47d3-abd6-7c0f316c06f6 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.936094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.936094Z digest=sha256:044400d039d38014e96d46f69b906bd1970703658eed453c852853cd8f8fd951

Observation 1eda6f2e-526f-4acb-b809-f56b05765310 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.960830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:14.853327Z digest=sha256:4b088c159e3574797bd368261f62801172ca50b26976d943922c297d462d4907

Observation 1ee185fd-e353-4680-b406-11bc1b508eea · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.412143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:15.132879Z digest=sha256:cbf18c7af36b2c55ed8a99b3b6b820031dbda2875a318a9966c80c749e5bffc4

Observation ce3f8763-5543-4aa8-a484-95c08a948cd2 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.702173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:15.031709Z digest=sha256:4c32a6f82118b4d85b9745ed5f2630e620cf00708481292f1e4d4d02a8f21a4c

Observation a7780ee1-ce41-48bc-aa46-3f2993a269f1 · outbound

This paper cites Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.334498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.334498Z digest=sha256:77f28bd65b9251b07cf72a6ea85dcfbae86337aedef59b70bd343d85a720d482

Observation 2292ea90-dbba-4505-9b79-ede17c98baab · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.230611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.230611Z digest=sha256:41db07937ca982047880457d2946751fd36041135abb5de96eef30c7bf57bdaf

Observation 03647407-1f64-4c07-840f-723e7258cd1f · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.533022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.533022Z digest=sha256:d5a60457cd63b9f0ba2953220d15adc3c8a868643550d2711de7206b4298f938

Observation 1183abe4-53c7-4251-8fd3-81917221eb3f · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.177628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:15.431630Z digest=sha256:097a279b6c94a8f11cc4808b0eb07ed24d8e1f87b05f6fd14ac510337738c88c

Observation cb3d7ec8-f2d5-4e92-acaf-e0b4150495b1 · outbound

This paper cites Salik, Rajiv Ratn Shah, Yifang Yin, and Roger Zimmermann.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Salik, Rajiv Ratn Shah, Yifang Yin, and Roger Zimmermann

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:22.955376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:15.732760Z digest=sha256:9e2beb429004c4971e68672475e22ca4b0dc5cc86b2fdb784b04d361ec895a34

Observation cfe93e57-9358-4e6b-ae3c-1fd69c32cc39 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.631928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.631928Z digest=sha256:a556bdc5fdbc72decfb9130d325a3c4661998f061123ff96c179ead9ce637de5

Observation c1b4caf5-dafc-4d0a-bb99-b0f39eb64b0c · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:22.591730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:15.965671Z digest=sha256:66b4fe75c1241cd2337f447f6049cb65fb0aca8edf2bd98f41aa243be6b6eaa6

Observation 6f9c274d-f627-426c-829c-1842c25a0dfc · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:22.788610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:15.862042Z digest=sha256:4883000771aaf1194555e6176aaf7079a7ad57b1692d0d8b00927dcfbbaeaa01

Observation 3f69a0db-57dc-429a-8de0-b4d50dccc736 · outbound

This paper cites Mandic, Wenwu Wang, and Mark D.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mandic, Wenwu Wang, and Mark D

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:22.193356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:16.165396Z digest=sha256:d3b2ed369f4915f36a5a1987b224b450f75df8afc1ac311411a3d8fa83dc4c9f

Observation 9cd59548-046e-4908-977a-3d14562cecff · outbound

This paper cites Raghavan, Gavin Mischler, and Nima Mesgarani.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Raghavan, Gavin Mischler, and Nima Mesgarani

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:22.365402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:16.059867Z digest=sha256:ad29476f91c7dc259ee39128ea995f1a317d28a8ce4bb9511b6934a6c36e6983

Observation e3194c8a-04cd-421e-adef-e2557f37f3b6 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Autoregressive Speech Synthesis without Vector Quantization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:16.401709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:16.401709Z digest=sha256:247414711316fcf08ee9da253960afac6d9dfd79c7816afd9564ea71383d5960

Observation 99e73ac7-28ec-4a05-8240-968939682a8b · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.977075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:16.281334Z digest=sha256:44c997899939bea6156584216c0bca1d190413b77972ee2f4d9a0aecf0d9bb6e

Observation 3af9a247-6149-4362-b3ec-3b5024433dbf · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:16.612710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:16.612710Z digest=sha256:c4e52efc6edbe4b896f3497bdc14d41e642dcee62c7fd29fe130004ed4a91e4b

Observation 42533c02-709b-444c-ba82-f108d0b1d43e · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.844207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:16.513575Z digest=sha256:1bffe228a701ac7c5ed5707f82139d0ed9d06fad6a3f5df3b3325332dc35cfad

Observation 0c3fb29b-10e6-4788-9da2-c122d56c57c6 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.699975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:16.856077Z digest=sha256:8e8ade9d622ba8c7b4cf7b496bde4576f359a19e78964d54967e419c1e3ae5eb

Observation bf519b0d-1df2-45f7-a72c-fb11017f577f · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:16.729243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:16.729243Z digest=sha256:f6882a80ea5712234b0eb7f3f81de43f67283f941944c458f663b4f325d8d251

Observation 6d151fd1-b8cb-4da5-b6dd-dafea2042c69 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.376640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:17.074451Z digest=sha256:09b88f5264b9a408863a84cab63140ca23fd0dd0470c38841145ce7e37c7ceb5

Observation d5aff840-67a9-4dbe-861c-bc0ed3ecf6f6 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.517553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:16.972078Z digest=sha256:1e0deac63263549eeeb361e62610e5418f90732803556d0478db47a2aab81a76

Observation c778f8fa-b4e8-4720-a550-e1c19f05d7d0 · outbound

This paper cites FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:17.295613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:17.295613Z digest=sha256:0f2283c5a19f36aa8c2c58398b9252bdba9273a28234dd270e53e344ed7573f7

Observation 13942b3b-21ae-497b-8103-3869f8f17fe3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:17.194390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:17.194390Z digest=sha256:488a3585ef57a5557a484a297c26b85f3a5b578dfcadb99cf8829beb130c2391

Observation 5aadd354-21aa-434d-b3e4-e651d848f1c5 · outbound

This paper cites Mel-Band RoFormer for Music Source Separation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mel-Band RoFormer for Music Source Separation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:17.731914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:17.731914Z digest=sha256:8504d0d1223c73d10166848a47eaf2b4f2647a907bc770c8fb036495633eae66

Observation 80d1a0b5-080a-4f04-8e92-8ef224a181b4 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.242636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:17.454006Z digest=sha256:a1072f0b0f0e252265fcc1c7786fb539d9439f6d74bcb4fa472019baa589a298

Observation 25de76f6-9bd3-4faa-b791-113c7b55626b · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.768252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:18.070098Z digest=sha256:a526ed291af3d74961bc99d83246afca2a26962a6a11033606fbd0fa23252094

Observation 3869cc5b-f98e-484f-8fbd-91cb5e62436a · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.618868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:18.195413Z digest=sha256:84b88b82f89e989e16cef95ead49495517f3339734d3ea969443fc614d35f6eb

Observation 91a24ba9-2e29-43ef-827b-95867ffd5fa3 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.915025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:17.855498Z digest=sha256:abd2febec5f37bd644b2ed79977e8e1614a78c0bf8730d86fdc16671b6472945

Observation c6e79082-fc53-44c9-8f1e-6ec750618d11 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.500277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.500277Z digest=sha256:a6ee01e219fcc9db40fad3b87b04d228a7ca76a18978590db3fba1a7dad65bf9

Observation 5232a114-2d96-40a1-b022-b42e42fff0c2 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.489871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:18.587362Z digest=sha256:f579346a3f1c2247bafe9dc96ad588051a112c45e3b527174d988f780aaea7cc

Observation b3fbe57a-f389-445e-a3c7-8911c34f56f5 · outbound

This paper cites Qwen2 Technical Report.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.341711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.341711Z digest=sha256:08c91a8644bc15fca388931ac89b1a169813cae4a24b1c767fb096b62df901f3

Observation 7eda949d-c7f0-4718-9603-63ce8d097b60 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.798620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.798620Z digest=sha256:e7d743d72b0e2c2b32a94f63814c3e6020cdc74e3e32315a5c3ff9ab99dd6107

Observation 66285bcf-7fed-49f0-8900-4f4d5a0d744f · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.900460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.900460Z digest=sha256:5f6a2fe382b2d7c4580ad722be10ee20fd9b729e2120652155f13ca4056c5c31

Observation 7417fdab-75af-4ed8-918e-d769dee767a7 · outbound

This paper cites Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.671928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.671928Z digest=sha256:4ca96d9203d94d395c1ad481672385bb8e08a8f34978b65db460d4d9e5ee24a9

Observation f08d860c-4432-4423-a6b6-64609e10eb18 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.709302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.709302Z digest=sha256:75b06b4133b7767bee412bb2eb3531b8308736a332c7a050c261f22e8b23a05a

Observation e39c3107-8a78-4cd1-b18b-1e2b4f16793d · outbound

This paper cites CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:46:19.565364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:19.155991Z digest=sha256:f9d449f0f7abf074d52c0088292400015ba2c6fce4f2ecc24ddb8aeccce99e6c

Observation 5b4e0a7c-e4c5-4cd3-bead-1dee93540884 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:19.243232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:19.243232Z digest=sha256:1c0746d2fd49d9965b0346e0658dc030f4e69c1732f193f6ed61d3eda2c25102

Observation 952abe25-feca-40f1-a9b9-985be6455714 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.344490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:18.985234Z digest=sha256:a3edf32acaa443a2949f9a2c7b009cf7f36b279eafb4f85bff61810ae917e16c

Observation 6bed8f7c-9191-4d2a-b1a7-59f1dd70c6e5 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.150751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:19.066327Z digest=sha256:05ea1e77a1a32ebee3041307a5799c1749dba8fc2b6c9e069ca49252459c75f5

Observation 946540d6-4f50-4a0a-bf43-337744e816ff · outbound

This paper cites In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:25.054650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:14.166030Z digest=sha256:06d6ea64b92313b0dd5cdca1d6dbd6a17362a5aa7c981c952acffc6e0fb2d0eb

Observation 6ecb8ab1-8022-41c7-bd52-9c10e0998921 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.053627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:46:17.612626Z digest=sha256:3ef5b7c877f02bb096d958428dc2e41035b803b290e7930ae38e31d4488ab551

Pith citing papers

Observation 9f73405d-b2d7-4205-bea6-b06ea22e0311 · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T06:04:30.083418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T06:04:29.936928Z digest=sha256:a629843157d309ae65cd404877ee041069c6ec0b2c9220296fccd4455fe5c5ac