Pith. sign in

Paper Citation Record · LEDGER

Audio-Sync Video Generation with Multi-Stream Temporal Control

As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2506.08003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08003 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:24:30.637501Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T11:34:32.558440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:38:14.727310Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59ac458d-94e1-44f3-8a0d-867e9db24c5e · outbound

This paper cites Diverse and aligned audio-to- video generation via text-to-video model adaptation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Diverse and aligned audio-to- video generation via text-to-video model adaptation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.824720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.255357Z digest=sha256:8f5e14c1792389bba3c81a9d1fb7c342ed43504530938b23d08c566193214a3e

Observation 6e1fd37e-ec24-4972-92c5-1cf29c734613 · outbound

This paper cites Long video generation with time-agnostic vqgan and time-sensitive transformer,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Long video generation with time-agnostic vqgan and time-sensitive transformer,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.809045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.260509Z digest=sha256:c07687818971a4eb150f54ed9a50ddb20fa909a0209f971eda16d2d8de932c4b

Observation 6aa8241b-dfda-4759-b83b-3a53ba12a598 · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.793603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.265481Z digest=sha256:b306a5f0677311314c3a03981e6300418c39df09c7e805087935b442ab8fca69

Observation 173c70ac-33a1-4244-a17b-a8251bd92aaa · outbound

This paper cites Sound-guided semantic video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Sound-guided semantic video generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.778883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.270005Z digest=sha256:8a89c7d3d24390736321382042f82a4d1016955214d3f2e830ea9207aac55517

Observation 104fdeee-48a2-4e1b-8463-d7fcf09cdd8f · outbound

This paper cites MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.764037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.274528Z digest=sha256:9161d9801e0e9ed7d233977e39e56dac9ec7f62e2e2fe168988b86074498a522

Observation eb9bff16-e028-4a01-bc6a-7395298b8204 · outbound

This paper cites TA2V: Text-audio guided video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control TA2V: Text-audio guided video generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.749032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.279142Z digest=sha256:9da9a0d19ab6fe947327579008e6689efcc27d569d7d0061943eeab71aca524c

Observation 02f5e58b-3e74-41af-9b4c-8f011ee5de61 · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.732415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.284477Z digest=sha256:6382490c4f6ba3d805658f65ee84744651cfc2cc63a3bebb9af0340a2b482150

Observation 5b0f35db-ba58-4d9d-ba5f-624f0791daca · outbound

This paper cites Speech drives templates: Co-speech gesture synthesis with learned templates,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Speech drives templates: Co-speech gesture synthesis with learned templates,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.716351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.289145Z digest=sha256:19fdc5764e51e7ddb4cf48d307638d7a56005b9b39016e7ce50f62d44c5c2dca

Observation 9db135e9-97a3-4b6e-9623-a31610f879cf · outbound

This paper cites Visualize music using generative arts,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Visualize music using generative arts,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.700762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.293448Z digest=sha256:6d3b42fcf395fe2d8baf0da7c78ed83f4ac5ab2e3d1bc95633810ea163653e88

Observation b6fa7cb1-0213-4c01-84e5-2fb459e1f6a5 · outbound

This paper cites CogVideox: Text-to-video diffusion models with an expert transformer,.

Audio-Sync Video Generation with Multi-Stream Temporal Control CogVideox: Text-to-video diffusion models with an expert transformer,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.684858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.298059Z digest=sha256:d7f319a039168aa4eee05747bf6ec22406e1e90573fd995a328d9c2534e5402e

Observation 53b57413-2458-448f-9c31-6ddfcba84fd7 · outbound

This paper cites Structure and content- guided video synthesis with diffusion models,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Structure and content- guided video synthesis with diffusion models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.669493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.302949Z digest=sha256:fd4f1db151fd57a29f7cf53ec2925b9a1d6375fa956cd122d125d4d2976ef037

Observation 44cb9eee-3cfe-4393-bca8-532eafaafac4 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Audio-Sync Video Generation with Multi-Stream Temporal Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.307892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.307892Z digest=sha256:e080d950509f98db111b1d704619fdad5751523a1ca264898a105e0f8608744f

Observation 90957916-2290-4432-99d0-5096cb5e939f · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.313733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.313733Z digest=sha256:9b7b273aabb2a8d3f2c8b754c69a1cd1a6b2c893679f06260878ad7fa50d501f

Observation e1db0a67-136b-46e0-98a4-1d8f3ddb1577 · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Audio-Sync Video Generation with Multi-Stream Temporal Control High-resolution image synthesis with latent diffusion models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.318712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.318712Z digest=sha256:86f500c3981e995064d56c24a6de5c22874e48ccaf3c5fd54b816819c45ac450

Observation 130049ee-dfac-44ae-9cf8-72b93248efbc · outbound

This paper cites Interpretable 3D human action analysis with temporal convolutional networks,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Interpretable 3D human action analysis with temporal convolutional networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.641049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.324986Z digest=sha256:2db63b05bad89432511cdfd9b7b7beb1bed960ce67c93a8cdad83201c500b155

Observation 9fcda4e2-2151-4ea0-bb10-0a8fbb1ac67c · outbound

This paper cites Is space-time attention all you need for video understanding?,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Is space-time attention all you need for video understanding?,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.625063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.330694Z digest=sha256:c7d88cb4804d17e2c98aa8da67b441e3af1e031a9e24051c3d483fe92eeef18f

Observation 5086d922-64e6-440e-be42-6440bf04d4c8 · outbound

This paper cites U-Net: Convolutional networks for biomedical image segmentation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control U-Net: Convolutional networks for biomedical image segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.606277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.335384Z digest=sha256:69d97d01248cdf30672d97e77b0fbf64b17a1f090eae441585d8ac4df4ebdd08

Observation af590379-9528-40cd-a146-62c41fa36b46 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.341073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.341073Z digest=sha256:ac61157af0303cfc4f3bf8d85f7d8a101a8e9f26bbefc9155a345308a56c023c

Observation 4cb69949-eeaa-4b83-976e-1315e82e3419 · outbound

This paper cites Scalable diffusion models with transformers,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Scalable diffusion models with transformers,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.349078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.349078Z digest=sha256:7c74b422edc2fd0e903ef867d95fc8d5dae63bfafc3bf63cd0591bb960900e4c

Observation f185d2ab-9d88-4072-bcac-028ccd651926 · outbound

This paper cites Language model beats diffusion – tokenizer is key to visual generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Language model beats diffusion – tokenizer is key to visual generation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.579163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.354412Z digest=sha256:c97962ed9237d80109e097094e0b75f016b8fd6803d9451933013426bfa8b21d

Observation 74cacca0-99d8-497b-a584-3fae50b1c88e · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control Wan: Open and Advanced Large-Scale Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.363858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.363858Z digest=sha256:8cc8ea7c02a4494dd02e7916bca9390145d7afa7c7f7e3f3b06de4d329d945f4

Observation c90e47a0-d7d3-4a13-b202-f45f984da477 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.370136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.370136Z digest=sha256:07eabb1f9e7dc528f42072584c040e783374a768f38af4cb6cd66e7710d51db8

Observation 1c9c5c64-36ce-4a6d-9ade-88715a693c0a · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

Audio-Sync Video Generation with Multi-Stream Temporal Control Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.377023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.377023Z digest=sha256:51b7a375544d7b33a161a5e8eb22e357f0bd7a567b52dd28a2cec0fb00b477ce

Observation 9386c427-f05d-4b6c-9596-cf0abfce3c4f · outbound

This paper cites Sound2Sight: Generating visual dynamics from sound and context,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Sound2Sight: Generating visual dynamics from sound and context,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.564156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.382367Z digest=sha256:4a5e510099ef96f9090569baf8cc82884434bfb8048c0d7adb4652e354587f73

Observation 482fc39c-24fe-4b38-896d-706d3e01cbd0 · outbound

This paper cites CCVS: Context-aware controllable video synthesis,.

Audio-Sync Video Generation with Multi-Stream Temporal Control CCVS: Context-aware controllable video synthesis,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.549939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.387065Z digest=sha256:c2289c05932f95e0444688d5dead840aae5481a0026223823809d99f896353b6

Observation 31c2da8d-43a4-4bf0-a56a-2187c03388e3 · outbound

This paper cites The power of sound (TPoS): Audio reactive video generation with stable diffusion,.

Audio-Sync Video Generation with Multi-Stream Temporal Control The power of sound (TPoS): Audio reactive video generation with stable diffusion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.535252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.392168Z digest=sha256:221368187967438e7f61ad5b49dc2ecbec68aed98c84b2d719ffeb97da70197d

Observation 0da452e7-443d-444a-b16d-2020746276fd · outbound

This paper cites Audio-synchronized visual animation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Audio-synchronized visual animation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.520802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.397425Z digest=sha256:30ab3da8cfbca2666755f0d5016e613a69c5511c67479af23a47ce8e18a66496

Observation d141dff9-2d35-4719-a13c-b94753f9f633 · outbound

This paper cites Loopy: Taming audio-driven portrait avatar with long-term motion dependency,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Loopy: Taming audio-driven portrait avatar with long-term motion dependency,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.505797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.402130Z digest=sha256:d5f780b317aca9c15c715133b0ded5d0e82b6440c4ae4f51c25c287eeb3e251a

Observation c3a1929b-d64c-4a7e-b2a7-33342cbf61a2 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

Audio-Sync Video Generation with Multi-Stream Temporal Control AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.408110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.408110Z digest=sha256:763f0243d1c78ad09d413b5ca2c1d0f52ccf07b501c537dc0db7d32f4d155fbb

Observation 7fe47a68-6801-4dd5-bcc5-7f8468080216 · outbound

This paper cites CyberHost: A one-stage diffusion framework for audio-driven talking body generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control CyberHost: A one-stage diffusion framework for audio-driven talking body generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.476293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.419495Z digest=sha256:4c754b6ee12a0c0b5f3d068299aa5c20e9a2e929b8730cd844a0c5698f03fa32

Observation 03919a94-3a27-4efb-87cc-cd6bf72d9b33 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.424440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.424440Z digest=sha256:f707395408af9f6a5443b15bdd44c34e2a66b78b23b252c1b396f5dea2e025dc

Observation 97c70973-2b49-4825-8c6e-fa7e72169ac0 · outbound

This paper cites Dance any beat: Blending beats with visuals in dance video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Dance any beat: Blending beats with visuals in dance video generation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.461031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.429761Z digest=sha256:d3878887a2334fca26ebe70dcb10a4988f6be35e720a0be6e010305fc66f7101

Observation c4d9c21c-b91b-4105-94ef-c6ca6d463b0f · outbound

This paper cites X-Dancer: Expressive Music to Human Dance Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control X-Dancer: Expressive Music to Human Dance Video Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.435331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.435331Z digest=sha256:ce0742a5a99552ec574c865543d4fc9c4927cbfa1898ea8f101bd599e0b115cd

Observation 7a10a7d8-b6b4-48a6-9fbc-5f009fdd31da · outbound

This paper cites Taming transformers for high-resolution image synthesis,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Taming transformers for high-resolution image synthesis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.444999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.442078Z digest=sha256:fe9f5dba1f5bf9a69f1afc3cd1829bc877aa37995a705717342c812a1caf26b4

Observation dd210d7b-cf20-478e-bc0d-459bbdea9649 · outbound

This paper cites A style-based generator architecture for generative adversarial networks,.

Audio-Sync Video Generation with Multi-Stream Temporal Control A style-based generator architecture for generative adversarial networks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.428879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.447252Z digest=sha256:e5010314995aedff3fc694ae68a9273539bd5040262d7e513d81696ef5c10bf0

Observation 0292bced-28bf-4495-94c8-8a5709d65e6f · outbound

This paper cites Tr\"aumerAI: Dreaming Music with StyleGAN.

Audio-Sync Video Generation with Multi-Stream Temporal Control Tr\"aumerAI: Dreaming Music with StyleGAN

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.452897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.452897Z digest=sha256:8d179a39e8e887a7f84db45628d6915733e835d36faabd36414aa99e0bc45b88

Observation 0a69a0f0-b5f1-4d9e-8971-7ff74ac49f2d · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Audio-Sync Video Generation with Multi-Stream Temporal Control UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.458514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.458514Z digest=sha256:60b54cefc175e759c218e48945dd0c5aa3b33a4e2ee469461a84612c4833cb90

Observation c80b6ac1-61f6-41c0-b13c-5c9b8137ce49 · outbound

This paper cites Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint.

Audio-Sync Video Generation with Multi-Stream Temporal Control Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.463937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.463937Z digest=sha256:6c49c9083536f70e7c82a1d52f63ad45c2aece7e9dc9675b8705b16fe60c7d4f

Observation a8e9d9d0-2234-436d-b21d-f7333d88753a · outbound

This paper cites AudioSet: An ontology and human-labeled dataset for audio events,.

Audio-Sync Video Generation with Multi-Stream Temporal Control AudioSet: An ontology and human-labeled dataset for audio events,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.413510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.469091Z digest=sha256:572a88eaab4acf360453e99f877eb1d1af57b5cfef40a2a55fb5e1423fb3a204

Observation d1a3ecbc-0034-46f2-bdcb-4a9efbaef075 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

Audio-Sync Video Generation with Multi-Stream Temporal Control VoxCeleb2: Deep Speaker Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.474142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.474142Z digest=sha256:32baaa8e3be38d9b3b093c8f15a72a04990e903a242add1c989a7dc9fafbf192

Observation 7b68e6ec-ae47-4859-815e-e0e7d579afca · outbound

This paper cites VggSound: A large-scale audio-visual dataset,.

Audio-Sync Video Generation with Multi-Stream Temporal Control VggSound: A large-scale audio-visual dataset,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.394335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.479131Z digest=sha256:b9f76983a2786a21661f04e0722e6ddf65b0a00c0658b5bf8afa899c4d8208e2

Observation 6bfa8e3d-ce63-4a43-983c-962631ef68ff · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.378605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.484063Z digest=sha256:a860fd85a8b1d983b38e0b3edb1292b3b46f120841605c3120a159f5db5d645e

Observation 32fdc992-0e22-406b-ab5b-a58d40724d58 · outbound

This paper cites InternVid: A large-scale video-text dataset for multimodal understanding and generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control InternVid: A large-scale video-text dataset for multimodal understanding and generation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.363011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.488419Z digest=sha256:b5497e915d2bebe3b1cd1177e4ee16735837896256865487bcf892d8b5240f6f

Observation a7f64bdf-cd4f-43d4-848f-6e33548caf24 · outbound

This paper cites CelebV-HQ: A large-scale video facial attributes dataset,.

Audio-Sync Video Generation with Multi-Stream Temporal Control CelebV-HQ: A large-scale video facial attributes dataset,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.492770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.492770Z digest=sha256:96923131a4fa5bfe3953e636faa519aebda9eff0c60164ee121316d47d663b14

Observation 02f33fb4-2f3b-4c8c-8b03-06a0ea14503c · outbound

This paper cites MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.497996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.497996Z digest=sha256:865e6bfaae288ddba000c7bc94ecb51e79c38a1ce0ee9b1dcdf3bddc927e6b34

Observation 87cd5216-68e7-4058-b722-039abb8eaf52 · outbound

This paper cites Condensed movies: Story based retrieval with contextual embeddings,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Condensed movies: Story based retrieval with contextual embeddings,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.336609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.503180Z digest=sha256:568feae201b25bac289c732af123fffb3f22b70007f8095bfe62656a2b68d6bc

Observation 8e73c803-64ed-40bb-a691-b58d78a20ef3 · outbound

This paper cites Long Story Short: Story-level Video Understanding from 20K Short Films.

Audio-Sync Video Generation with Multi-Stream Temporal Control Long Story Short: Story-level Video Understanding from 20K Short Films

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.507687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.507687Z digest=sha256:639660b6f61833b80c167631f8e57ddcd46d816b18221c2024739eeb4070420a

Observation dcc8e9f4-e5a2-46f9-8abb-7876849b57ea · outbound

This paper cites VideoCrafter2: Over- coming data limitations for high-quality video diffusion models,.

Audio-Sync Video Generation with Multi-Stream Temporal Control VideoCrafter2: Over- coming data limitations for high-quality video diffusion models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.319998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.512852Z digest=sha256:c673e36d908efd50a01a594e954f3f23c1f97f4779a49711f0c94fc61a0d2c57

Observation d62e1bd4-120d-4c6d-9a71-f055e92dacbb · outbound

This paper cites Video cut detection and analysis tool.

Audio-Sync Video Generation with Multi-Stream Temporal Control Video cut detection and analysis tool

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.304988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.517286Z digest=sha256:f226c32aaa1162191a9716506364ffa6487e31274dcc7a939066388dd6a414e7

Observation b8fdad6b-7583-4436-a9d0-30ad09f57a16 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

Audio-Sync Video Generation with Multi-Stream Temporal Control Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.522443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.522443Z digest=sha256:846b8ef6a7da361cac1f77d8114ca80cc1fa3047ed80fcbcb1c859179cc56ab0

Observation 5a25c88f-a587-46d3-8cd0-b7f0467bacc4 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Audio-Sync Video Generation with Multi-Stream Temporal Control LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.527178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.527178Z digest=sha256:92fd442594017caa7bfe5cf1549043b9898e8e7139f6edca5ab72eb4abb20cdc

Observation 2fea16d8-7016-45c3-b221-8c19d508a5a2 · outbound

This paper cites Cinematic sound demixing.

Audio-Sync Video Generation with Multi-Stream Temporal Control Cinematic sound demixing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.289568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.532116Z digest=sha256:573f62e390a0a8a2af77d99c2c8394d0d60539798fbb74942a7868b50f7072fa

Observation d017bcb6-ff74-45fb-8a8a-c1960fa909e5 · outbound

This paper cites Spleeter: a fast and efficient music source separation tool with pre-trained models,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Spleeter: a fast and efficient music source separation tool with pre-trained models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.273725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.536937Z digest=sha256:37688bdd7016d12167a3d8b617d8b33885af43a64de9a8a134abfd8c1ac179a9

Observation 11c5ce5e-efa2-41ea-8d26-e62ee7767c38 · outbound

This paper cites Ultralytics YOLO.

Audio-Sync Video Generation with Multi-Stream Temporal Control Ultralytics YOLO

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.257250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.541280Z digest=sha256:df189641300a2f72d2c6c1f120117bbb257c7a5b5a57e1f21421b7a8d63b744b

Observation fe301372-756c-4a2f-905c-8cc24c4c07ce · outbound

This paper cites Meet scribe.

Audio-Sync Video Generation with Multi-Stream Temporal Control Meet scribe

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.242082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.546019Z digest=sha256:e90eec694d79f9377b2e9510bad6261f0ab441fcf8f4de6063c10b17e626b190

Observation c24d23f4-aca3-4d3f-a2ef-5d519b718316 · outbound

This paper cites TalkNet 2: Non-Autoregressive Depth-Wise Separable Convolutional Model for Speech Synthesis with Explicit Pitch and Duration Prediction.

Audio-Sync Video Generation with Multi-Stream Temporal Control TalkNet 2: Non-Autoregressive Depth-Wise Separable Convolutional Model for Speech Synthesis with Explicit Pitch and Duration Prediction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.551203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.551203Z digest=sha256:1c46821da63c8d2e59806b7849caa1a02721c26f947559c78c95d081a9247959

Observation 0505b90b-5a1f-4c55-9af1-34f06ae77fe3 · outbound

This paper cites wav2vec 2.0: a framework for self-supervised learning of speech representations,.

Audio-Sync Video Generation with Multi-Stream Temporal Control wav2vec 2.0: a framework for self-supervised learning of speech representations,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.225434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.556119Z digest=sha256:6464d3730e1d9dd91ebebc061df26329d7eafea6a07d0a6286e9a6e0a5fcbcfb

Observation de1ebc8c-a82f-4075-9e3f-21ffa472914a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Audio-Sync Video Generation with Multi-Stream Temporal Control Adam: A Method for Stochastic Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.561313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.561313Z digest=sha256:e49ae349140e300258cd89696a3b22543228f79aca3cb6316a6c6ee3d1803379

Observation 7d5800e6-a8a9-4933-8e39-85d99e5040c8 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Audio-Sync Video Generation with Multi-Stream Temporal Control Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.566199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.566199Z digest=sha256:2a119ecc5faa037f5ca570ac2a6a098b87b5c5f4fb1d434b3cb5e3f60c905fba

Observation 861757ed-d16a-4659-8865-6d48a251e24f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Learning transferable visual models from natural language supervi- sion,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.207263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.571380Z digest=sha256:b21a14f502eddd4465a1ec418becb1c285447af2c34b4541446ca54f6c3055db

Observation ccb49950-47e0-4796-a667-e079589500d6 · outbound

This paper cites VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.577220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.577220Z digest=sha256:f0418a11f16f387f6b4cc7b47c532388b5fe5ea3dab860ce7e9884b3bc4196ef

Observation fbd91222-ad01-4fd6-9b32-3c82239582fe · outbound

This paper cites ImageBind: One embedding space to bind them all,.

Audio-Sync Video Generation with Multi-Stream Temporal Control ImageBind: One embedding space to bind them all,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.190770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.582083Z digest=sha256:3212a3fc9bab450d98fdd7d3c08b1a82bcd2556d91e12dae7f1b3b9369e18199

Observation 646ecfe8-29ef-4aba-af15-a3ad89219233 · outbound

This paper cites Out of time: automated lip sync in the wild,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Out of time: automated lip sync in the wild,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.175016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.587157Z digest=sha256:141ca21c07be2729e7325204aceedef3c17c229f036b79c59fee6f728a502f32

Observation b9b7134a-a243-48a1-95e5-66c7ac1cf334 · outbound

This paper cites DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.598574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.598574Z digest=sha256:8972c6f8c222ba04d1003960efba30d1eb6030e756a6765733f84e7f480980e2

Observation 2f578fe6-ae65-46dc-8359-c8063755c4de · outbound

This paper cites TA VGBench: Benchmarking text to audible-video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control TA VGBench: Benchmarking text to audible-video generation,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.157131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.604212Z digest=sha256:9a47d04d9ed326b8ef89830b8ed978189537ceeddd3e6c1e65c6b29ac9d2d6e6

Observation 1e644474-6768-4be7-b191-5aef89235a09 · outbound

This paper cites MMDisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control MMDisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.137563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.610563Z digest=sha256:f97f52382857c556439d30246c4cbf2b1ef9a16e8cf60418c529274ae20a00d5

Observation d3aa9b78-8bee-4c6f-a667-ef4b871127dc · outbound

This paper cites AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.615859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.615859Z digest=sha256:11ef8661bb25ec67aea32175f4bd7d261711067111434ecde4dc45b6e947a007

Observation c59b8c40-6eb5-4122-870f-29532280bded · outbound

This paper cites Hallo2: Long-duration and high-resolution audio-driven portrait image animation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Hallo2: Long-duration and high-resolution audio-driven portrait image animation,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.621496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.621496Z digest=sha256:60e07329ab88e77cf40b6adeee57f4d6eda8e9ec5046fb734b7ae66dd0d6ac9e

Observation 21c9b3e6-2ec4-4247-8fa4-8220abff7b6f · outbound

This paper cites SadTalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control SadTalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.490873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.626811Z digest=sha256:4d2f63451e81aeb1c2a1675182694ddeaeebae0e90bd19177c02391cc1ab91cd

Observation e04b7e44-6576-4e27-a303-72d88c1d9e66 · outbound

This paper cites Maximum filter vibrato suppression for onset detection,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Maximum filter vibrato suppression for onset detection,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.108216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.632022Z digest=sha256:62d946e910b9a1f5f28ee63c72ce5d5f7b649d0c8a6a63a38cb9ef97fdd3f898

Observation 533130be-c7b5-4938-a734-f57842d3dd34 · outbound

This paper cites Determining optical flow,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Determining optical flow,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.087789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:24:30.637501Z digest=sha256:92cd2f2bbdc80113090c1a664b2c65d29dcab9d331d5ab156973f22241d3409a

Pith citing papers

Observation dea592b8-9aa9-4846-b81b-926220787e57 · inbound

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing cites this paper.

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing Audio-Sync Video Generation with Multi-Stream Temporal Control

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.728832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:34:32.558440Z digest=sha256:30a6aa0e8e48f0bd12655f881b936fc6d9d1626e87b325a734622d8dd5bca9fa