Pith. sign in

Paper Citation Record · LEDGER

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

As of 4 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 1 inbound Pith citation observation for arXiv:2603.14267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.14267 v4

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T11:56:13.914121Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:46:58.010112Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-11T09:51:00.769273Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact5
  • verified fuzzy56
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d5f1bba-507c-4995-b77c-a2db99d50867 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.318239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:fe919e731f8408f1a61d6babeb2c4977ea0b40f7722e1aee8e32387b6b1cf6a6

Observation 1d678fc8-8c2e-4127-bdd9-d8e39793d0f3 · outbound

This paper cites V2c: Visual voice cloning.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization V2c: Visual voice cloning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.263338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:67144a16b315b9630292e91003ece69080e2910c94ab5e9c103d4e1a1f131ffd

Observation 8a06e516-a4b9-40c4-ac2b-67d1c6f02cfb · outbound

This paper cites Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.315683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:8bdf419d3e8eb21562442b84b11e097ced775fb98f8a35a8d7008abeed982696

Observation 0b9d66b4-0c1b-478f-8411-b3b0341ad95a · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Neural codec language models are zero-shot text to speech synthesizers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.320855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:970f131d194a39226b4316260781a152965555ea59af7e9916c9e751f33905fb

Observation 9a249c86-d129-439a-99a1-9e4e73fb4754 · outbound

This paper cites F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.288345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:f556c6eedfc53a65812ec965112a0fa646934ca02acc237d1fa1ebe826b64a98

Observation 78267e3a-602d-4195-a58e-580949862523 · outbound

This paper cites Intelligible lip-to-speech synthesis with speech units.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Intelligible lip-to-speech synthesis with speech units

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.328390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:0ff174c0e42f6ffb85c43c1107a14dc4929d36739fbdf664fcea85debb4ad57d

Observation 24c5ffae-93be-4c96-ac7c-116228609ef0 · outbound

This paper cites Aligndit: Multimodal aligned dif- fusion transformer for synchronized speech generation.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Aligndit: Multimodal aligned dif- fusion transformer for synchronized speech generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.274474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:6322b5bfcb3943b6b1fcfd821e70812691799f9d4466e17f8eaa4f9fa39d1333

Observation 3ff8c2c0-a1eb-4df4-95af-a7471d34c2dc · outbound

This paper cites Accelerating Diffusion- based Text-to-Speech Model Training with Dual Modality Alignment.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Accelerating Diffusion- based Text-to-Speech Model Training with Dual Modality Alignment

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.267812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:0c7c354801d3b66c94b7721474dc34b137a78a96d32c32660467dce0a1315d9a

Observation aa7c9b87-bf90-4dc2-848d-dc3346e30468 · outbound

This paper cites Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an un- voiced/voiced classification frontend.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an un- voiced/voiced classification frontend

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.270998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:a747eac3d8a506c9f51bd49b0775d0b842704b0ac896b18dd1b16ffa637792c1

Observation 6113a4be-2e74-42ec-b59e-4a8f894d0367 · outbound

This paper cites Out of time: au- tomated lip sync in the wild.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Out of time: au- tomated lip sync in the wild

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.319183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:aae436fa726410549965a8362d3084c4ce06f7285990a023bbb3cf662ab17490

Observation 65516fd5-6a0e-44db-99b9-b14de03b9c53 · outbound

This paper cites Learning to dub movies via hierarchical prosody models.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Learning to dub movies via hierarchical prosody models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.341942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:4907b7cf735572f897a8f20ec6e21a08728ce1668fd211a9300cbd68d25eddef

Observation ecbf7438-009c-4c50-9b91-65490977c615 · outbound

This paper cites StyleDubber: Towards multi- scale style learning for movie dubbing.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization StyleDubber: Towards multi- scale style learning for movie dubbing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.334665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:252530530151f0c5764ad2dde66d3e38c8dd9e0b482575740969d7d191341d80

Observation d0dd85ca-9e14-46ab-b4a9-62df08bcc87a · outbound

This paper cites Emodubber: Towards high quality and emotion con- trollable movie dubbing.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Emodubber: Towards high quality and emotion con- trollable movie dubbing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.364485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:77372a86234e95e2658fd5f9637ba61be19084868b63da9c6f06dd8eef05b0a6

Observation 0b3d3851-40f9-40ce-b885-75240137e53c · outbound

This paper cites An audio-visual corpus for speech perception and automatic speech recognition.The Journal of the Acoustical Society of America, 120(5):2421–2424.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization An audio-visual corpus for speech perception and automatic speech recognition.The Journal of the Acoustical Society of America, 120(5):2421–2424

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.229695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:7ef4454e238a406c1628203bdaeef047bee0ac19ca1702c4389e66253d5be6b3

Observation 610db2db-cf00-4f9a-8785-5b95df970a71 · outbound

This paper cites Sigmoid- weighted linear units for neural network function approx- imation in reinforcement learning.Neural networks, 107: 3–11.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Sigmoid- weighted linear units for neural network function approx- imation in reinforcement learning.Neural networks, 107: 3–11

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.341579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:8d2627a1ceeaf5de4768aeec0d85d0c1887c18f16ed5bc4d67ccbeb33173756a

Observation eded40d4-19d1-4aa7-8c13-607fbaed4df8 · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.345188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:7d97000babf2156296bea6449133667afc6657798a96bab9b001e5e402a5acf2

Observation 05cec627-2750-4cf7-b4ec-a65fbff14f5e · outbound

This paper cites LLaMA-omni: Seamless speech interaction with large language models.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization LLaMA-omni: Seamless speech interaction with large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.225863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:e1141538fa56315d5a9d908bb1d2b0f3252d66de933e79c96a2ac997637a0874

Observation 7af4c8a5-b4c5-44e7-8cc5-2634a55c9f33 · outbound

This paper cites an unresolved cited work.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:00:00.237377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:58eb4f7cef4ff8e8624c7f420bc0b42d06f8ae3f21fc96888b23b0dfd98ad399

Observation 9230ce63-01fe-45c3-a14b-00f7964073b3 · outbound

This paper cites V oiceflow: Efficient text-to-speech with rectified flow match- ing.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization V oiceflow: Efficient text-to-speech with rectified flow match- ing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.370565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:02d74f696a4e9d9ab101966271fb3f717cb8e2d29b0233ca476211df60d27d74

Observation 972ee4d0-1303-4922-bdb3-2f0576d74bae · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:59:59.585731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:908348a4610d49d62459f5ba7b2c058cdfb7923b3cf4331582a4e9de07e950c4

Observation be3df557-aabd-4dfb-aa11-22010fe593a4 · outbound

This paper cites Boosting large language model for speech synthesis: An empirical study.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Boosting large language model for speech synthesis: An empirical study

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.331174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:bd4b14f145d28d0a1f1098f90ed5821bf514507354c0add8f31f9408b091d36c

Observation ce3bc253-7c51-46bf-b025-60570bd53054 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Gaussian Error Linear Units (GELUs)

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:59:59.605778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:5812cd885badc6e566fd04e526a1c59f4fd35059591f19a9dce13ca77a38be4d

Observation eda4591f-69d0-4a8a-a2a1-d51027a49c27 · outbound

This paper cites OZSpeech: One-step zero-shot speech synthesis with learned-prior-conditioned flow matching.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization OZSpeech: One-step zero-shot speech synthesis with learned-prior-conditioned flow matching

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.409904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:9015f59768999cce3de0a3a794ca7a766b23f28196f44660864f01ba9d5d3379

Observation 83c8e916-1ff7-41bc-a4bf-56e67d60473b · outbound

This paper cites Neural dubber: Dubbing for videos according to scripts.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Neural dubber: Dubbing for videos according to scripts

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.419240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:a32d7496252171d998249dc5434f2b4995e8421107f29de370cd6b437e65c913

Observation 8d86afa4-e82d-4cd7-8dd9-fe2d248850cf · outbound

This paper cites Hunt and A.W.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Hunt and A.W

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.378570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:d11d6d752827d9ccb9c49cc90abc2b288589af8b5de629c86547c5c7f24d2619

Observation b13465eb-c358-47ce-93d0-ffc7f18cfc66 · outbound

This paper cites MobileSpeech: A fast and high-fidelity frame- work for mobile zero-shot text-to-speech.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization MobileSpeech: A fast and high-fidelity frame- work for mobile zero-shot text-to-speech

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.338292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:2750bb2acae40b2a9ec4e530f7f75419d17f47a4a4c97bddb4b9afbc17114734

Observation 33523617-7809-41ba-9b6a-61afe36c93b1 · outbound

This paper cites NaturalSpeech 3: Zero-shot speech synthesis with factor- ized codec and diffusion models.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization NaturalSpeech 3: Zero-shot speech synthesis with factor- ized codec and diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.249282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:3b42d06b2e9ee1e46abc7d3f787def946c8f05a15f555d600f8198c5ff198b57

Observation 0707335d-1c50-4199-bbf3-b7666a8f69f4 · outbound

This paper cites Zet-speech: Zero-shot adaptive emotion-controllable text-to- speech synthesis with diffusion and style-based models.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Zet-speech: Zero-shot adaptive emotion-controllable text-to- speech synthesis with diffusion and style-based models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.347000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:84f564664d47aa27fa70b8039ee2c585f818f6c4f905795dbd5837d824ec05f0

Observation 91dd46d8-7084-42e2-888b-aab5cad29b4a · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Glow-tts: A generative flow for text-to-speech via monotonic alignment search

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.278006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:c4fe5a48708ca91a83a165f92245fac520ef2b025a7c075db83f209e3ad6bd57

Observation 013720c2-a86e-4ca4-94dd-56a97aeed683 · outbound

This paper cites Shih, Rohan Badlani, Joao Felipe Santos, Evelina Bakhturina, Mikyas T.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Shih, Rohan Badlani, Joao Felipe Santos, Evelina Bakhturina, Mikyas T

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.284703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:f773f1310938fe2044379761f49d5ae71de290ca9a01b823a91af016cfc0f25a

Observation 92891c00-3e5c-433a-aee5-477f5062b300 · outbound

This paper cites V oicebox: Text- guided multilingual universal speech generation at scale.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization V oicebox: Text- guided multilingual universal speech generation at scale

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.244587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:6ee57e4cf37679a74b5ddf3a7354b4ce4a82ae6a254f92429a39869bc534f14d

Observation fbd66383-48ca-4d87-bbb1-ee6403c2eca6 · outbound

This paper cites DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization DiTTo-TTS: Diffusion transformers for scalable text-to-speech without domain-specific factors

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.288898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:0dd6c5df29b7bd1835273e6ccb11e9151cb705d1254193003b7c5f27a52cb528

Observation 5c43cf4b-cb38-4815-9f3b-e93c03285115 · outbound

This paper cites Decoupled weight de- cay regularization.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Decoupled weight de- cay regularization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.361770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:3ee852bcfdfce910b03640f35e1c4a1765c1207b40bfc44be94534574f3751b3

Observation 0db26fd0-7396-452f-b1a8-3d1895eed41a · outbound

This paper cites Monotonic multihead attention.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Monotonic multihead attention

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.335784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:847f41278f71826d4031c49d938ec4d2a11a3b9cbc117f7d828c938d7245c005

Observation 72858e84-e149-4e9f-a9f1-931307c2d965 · outbound

This paper cites Montreal forced aligner: Trainable text-speech alignment using kaldi.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Montreal forced aligner: Trainable text-speech alignment using kaldi

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.260249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:0fa147d4aa9abeb07c340a42dd2c10e9acb5ee1c4c0a6f83bdbb5e540bc5dc64

Observation a251bfc1-96a7-4048-8339-e99a18161037 · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Matcha-tts: A fast tts architecture with conditional flow matching

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.321390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:8486bd0e7fc43c410030b7a7fd65034219068e42565e8a06a3e541c9a17ec70f

Observation 5ae73557-32a8-4416-8004-af9aa69e2c3f · outbound

This paper cites Meng, and Furu Wei.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Meng, and Furu Wei

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.323268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:823c3a37738772586b30a25b7b43c402eb22112b7add720c0a901b5ae227a661

Observation 3c2b4ae7-6690-47b1-a932-130221a83ab8 · outbound

This paper cites Mish: A Self Regularized Non-Monotonic Activation Function.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Mish: A Self Regularized Non-Monotonic Activation Function

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:59:59.601393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:c2b7304c118dbc3abcf3ccd9d282bdb9c7985324eb2ebc77bf1197e365e94241

Observation 7d0661db-97be-4de2-8a1f-81cc01666265 · outbound

This paper cites an unresolved cited work.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:00:00.373497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:ddc2d8110e534817c60634f8ffe3ffbbbe1ce4662d55eebf03761fc1e46ce939

Observation 9e85afdf-2bbd-4be9-91e7-3fbad0df55e5 · outbound

This paper cites g2pe.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization g2pe

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.376057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:f7fc15043104e7f4b5b9f7fd0f5511babdfc15632b44ce97895920ea4793a5a8

Observation a51c1974-851f-4233-808c-5b6c7f5fc9d6 · outbound

This paper cites Scalable diffusion models with transformers.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Scalable diffusion models with transformers

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.382071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:051824018ab9f89cac09810db609485105d2e850f9de5e476b964856477d7e21

Observation 0aa3e1f0-945f-4ba2-873f-fcac30605741 · outbound

This paper cites VoiceCraft: Zero-shot speech editing and text-to-speech in the wild.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization VoiceCraft: Zero-shot speech editing and text-to-speech in the wild

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.388183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:6f14fb3f41929021e3a1bdc6bd7e9bc975466334967b80c9d95bfd644f3f34e7

Observation 22c663d5-c5e6-4bbf-aec6-6b26cf43b5e9 · outbound

This paper cites an unresolved cited work.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:00:00.391889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:1dae33d761045b98abbf3268eb17e07e28b658c0332a21f69f6cbb8216127889

Observation b76406a8-6833-4157-8000-65cd8ea8ab7c · outbound

This paper cites Speech resynthesis from discrete disentangled self-supervised representations.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Speech resynthesis from discrete disentangled self-supervised representations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.416611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:76c1f2bc2b941e4c73ee97c54ed8cd28111f227e4e1171d91b16e946a7776b12

Observation 269ba7b0-f5e4-4e7d-bc8a-63d08b6f3fc6 · outbound

This paper cites Robust speech recog- nition via large-scale weak supervision.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Robust speech recog- nition via large-scale weak supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.422015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:cdb323fa034aaba48b2962ab65e4574b4554e7179186a8cd94fbdef1565eed0e

Observation 1a56d666-ccce-4957-adde-0f7f44932a82 · outbound

This paper cites Fastspeech: Fast, robust and con- trollable text to speech.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Fastspeech: Fast, robust and con- trollable text to speech

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.404267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:1aa39e2e3bc9ed40689c5ecd78111cb626a554b0e38352b651046dd385fb2a99

Observation feda58f0-07e5-433c-98eb-3c4618056172 · outbound

This paper cites Fastspeech 2: Fast and high-quality 10 end-to-end text to speech.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Fastspeech 2: Fast and high-quality 10 end-to-end text to speech

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.398848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:7223b745f07de1b719ddd8a66f2d1e1de0652a21d3b5dd2dc7b9059e0b25625a

Observation df84d74b-44f9-4720-93f7-a92cce382c1c · outbound

This paper cites Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, Rif A.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, Rif A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.401814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:10589b13a05ce72fee7aa35107d021017777bc5dae5a193de5fd5d2b80ee66bc

Observation 040dcdfc-c036-4e61-8f75-a0c83d71980a · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.407103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:650ba301e51f4c354bbc5bc2774ac65396f690a89ae96e2a5ca109aecc79b76c

Observation be7bc2df-9edf-4a90-9cfa-a6f0aff4ea13 · outbound

This paper cites Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.413747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:bd0702f5715af5328b75e6125b10ef283f92cbf064c448ea79fb7869bb5271af

Observation 6e2656ce-a6d0-493f-b076-b597206a1946 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.385031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:6c039d80b172b8342f0eae2292f5ce146b9118341e6eae1b6a05921a9d4eeabf

Observation a6c6e768-8556-4d06-b5e7-30e87bdc5d7a · outbound

This paper cites UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.395139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:d737b9f6289c85f97ed8b2496a9d115b63bee2e9bcda7d8c959d52a2562b0308

Observation 28390b1a-4208-49ef-befb-ea2cd143fcc5 · outbound

This paper cites an unresolved cited work.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:00:00.367463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:ce0f8db73e044d038d61b943b5b1266f3606ea905da771efb0ff5550f8dc4778

Observation d5fc9a45-8b02-43b6-939b-541e33fa06e6 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:39:51.552824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:78486fbf1c9cfbe35768b816b47e3230ef6464909ddd9ed5bd266ad2cb51e481

Observation def93211-0a50-454b-a80e-72cc5994cdcc · outbound

This paper cites Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.313086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:43fe885b400b5f531b9b4bd6c71c273af7425f7336e8007e773b93380f6010e1

Observation 20a97d3e-33e2-4976-84e0-e16c9c3195ea · outbound

This paper cites MaskGCT: Zero-shot text- to-speech with masked generative codec transformer.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization MaskGCT: Zero-shot text- to-speech with masked generative codec transformer

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.325986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:36471149f43a4b43cae8f61226565f500e3d2ad4315b4cda186c4f6a794dede2

Observation afa13a78-7126-4280-825c-a4ea4cffc492 · outbound

This paper cites Con- vnext v2: Co-designing and scaling convnets with masked autoencoders.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Con- vnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.312127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:a974e70afa32f96e00b6b67c811064b2af8761760d39cceffb0fc7c67cb2f59b

Observation 5d465546-025b-4237-8d89-9df19158282e · outbound

This paper cites Lipvoicer: Generating speech from silent videos guided by lip reading.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Lipvoicer: Generating speech from silent videos guided by lip reading

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.333858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:d250cc6ae8b3cbdfe16c0c351b50f87874932f561bed909e3678330d9cb7a5dc

Observation cd555d55-37e6-4687-9cce-3c7e2055ecb8 · outbound

This paper cites Simultaneous mod- eling of spectrum, pitch and duration in hmm-based speech synthesis.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Simultaneous mod- eling of spectrum, pitch and duration in hmm-based speech synthesis

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.331620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:b4639c475aa8cde87e64472544111d7fe2d4d530aa6e66b7b41aad4fcec5354b

Observation 63454965-20ff-4690-b961-e1a8fdbdebb8 · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.297574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:c3967359c24573c202bffcf6a7993cd728fd12f7db21a3b1cd0250f0791afc09

Observation 84b0e2ae-172b-42a0-a79f-bcfdf1f1c229 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:59:59.597146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:3eeae65933556a0aa7701ac27fd21449043049ffb62b79aa2cc30ba2fa48f25e

Observation 4d7cd245-f93b-4952-a2d6-0d65447059cd · outbound

This paper cites From speaker to dubber: Movie dubbing with prosody and duration consistency learning.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization From speaker to dubber: Movie dubbing with prosody and duration consistency learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.300906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:9d8db7b49c96cfeeda3b87b8457c8d27f53535f587f0c3eb81d6c26a01118084

Observation 0899a3c2-894a-44ea-8591-f54b72650c29 · outbound

This paper cites Prosody- enhanced acoustic pre-training and acoustic-disentangled prosody adapting for movie dubbing.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Prosody- enhanced acoustic pre-training and acoustic-disentangled prosody adapting for movie dubbing

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.323140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:55ed1cc2cb5caa6be6f23de87a08f92a01b15b699f35c8b8b87582df28e81fd6

Observation 271bd3ac-dcb3-4798-ad9f-81fdbb636d86 · outbound

This paper cites DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchro- nization.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchro- nization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.355602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:2f41fece5c97146da856335b6a219fb6bbb18e9d92050e97cf1bad6c49e31c87

Observation 648d5ba8-dafb-47f4-ab5b-a6358189169e · outbound

This paper cites an unresolved cited work.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-05-15T12:00:00.291487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:af88216f57016399405db0c515503b357c16fd0d60598cee108b9a10743bc2ae

Observation 85d3807e-bfaa-495c-8b10-c15935c3ce5c · outbound

This paper cites The light fusion network is implemented as a linear projection layer.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization The light fusion network is implemented as a linear projection layer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:00:00.359070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:76a4e8b19540f4092a7fa84d11c8aa8fb9d5e3f7362f149389dc464a90b41313

Pith citing papers

Observation 886aca88-63d5-4480-afe1-cc6036298bfe · inbound

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing cites this paper.

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T09:51:00.772682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:46:58.010112Z digest=sha256:533bc4cfd475e46d2c1628cb01fdc4643ca799f0e5bc3203f54293bc464fcf02