Pith. sign in

Paper Citation Record · LEDGER

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion

As of 16 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2608.11913.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11913 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:27:56.984255Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e7819bb-46a5-4095-844b-9aa85c87df2c · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion AudioGen: Textually Guided Audio Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.712678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.712678Z digest=sha256:86c427e69a409c6d5c725d596c43a1958bb9464f2195f0fe75ca7cd28675f95a

Observation 091b67d6-4056-4fb5-b0de-355b858815dc · outbound

This paper cites In: International Conference on Machine Learning, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: International Conference on Machine Learning, pp

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.913775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.718048Z digest=sha256:1e867a54af11393083d32b6cbcd4651b8b2bacb2f840e42d91b74bb18736adfe

Observation 5f8384a3-22bf-42f0-bfa5-ca7b7925f6e5 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.722383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.722383Z digest=sha256:f983c0c7556deb8781e1f38a69f26f7b6e83cbc5d64bfa4edd644864583f7ef8

Observation 732a2beb-af5e-4de3-ae60-8c4583503fb1 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.900236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.726676Z digest=sha256:2581eaf5bae08a4b54e4bb92795bf81eb769ea930122b1dfb7f44c510d0afb4f

Observation f2814413-27d7-4f0f-ad66-1b8ce175a5f4 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.887526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.731597Z digest=sha256:1585bc5cf68d4c0017b0afed5583bcd143d557b65bf87f3f042ba359f9850d0d

Observation 52ae80db-ecd6-4720-bd3e-77e755cb7060 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.735765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.735765Z digest=sha256:c017ffddfcc28e93570d3bc2397173d14b314ebe9aa14092bed27a023a6a2ca0

Observation 8252999e-aeb5-4a3c-bbc1-d58235d76c56 · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems36(2024)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.873825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.740329Z digest=sha256:5b09c2e278a79f96109aa29d6e97806cceb7dc7d72eda453ecc9b3babe3927c1

Observation c7850795-4826-465d-92fa-e522197828a4 · outbound

This paper cites In: 25 Proceedings of the AAAI Conference on Artificial Intelligence, vol.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: 25 Proceedings of the AAAI Conference on Artificial Intelligence, vol

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.858771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.744066Z digest=sha256:f50b6cba967003c5b14df357893b33643ba4bc447f174dc7e83e53afdf0bda99

Observation d56a2bbe-43c4-4577-8b00-6f5b14d534d2 · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.747595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.747595Z digest=sha256:93bafa06a1c93b5eb4d20cf38d547c06afe25f4d64f91c750e4a2fdeeb7e0c4a

Observation 8c6298df-201d-4fae-a2c8-03a640fb03c0 · outbound

This paper cites Jukebox: A Generative Model for Music.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Jukebox: A Generative Model for Music

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.752114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.752114Z digest=sha256:b2a2d487d0272f777f29baa8bee8c9066cd40c30661257f68cd97ccadc7abbf6

Observation a87b5197-e653-482e-b3a6-0de78d8b2111 · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems36(2024)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.844657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.756554Z digest=sha256:c2ae0f8332ab5c74d0a7af066fbb1a1576ecc76c9de3567a5cefb41fd8b99504

Observation 375c3569-f328-43d4-824d-aace6b6e536e · outbound

This paper cites Advances in neural information processing systems36, 14005–14034 (2023).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems36, 14005–14034 (2023)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.830642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.760900Z digest=sha256:257f53a8f50e102e0b90a74bf9ce701e0dc81d5917cf489ef06e498187749fd4

Observation 99b3b07d-b06a-4258-8fc5-763fce20dff8 · outbound

This paper cites Masked Audio Generation using a Single Non-Autoregressive Transformer.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.764846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.764846Z digest=sha256:1fcd499d300f9c93006950459f3c8a6ac6e858e8cc54caaeacfdafd04f059070

Observation 6323bf76-1628-4056-9d74-21c204b9d7ef · outbound

This paper cites In: NAACL-HLT (Findings) (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: NAACL-HLT (Findings) (2024)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.816176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.769736Z digest=sha256:ebf0dd18f24808d7a572785854a24da4d17d60350dc895d745e41cdd3a0ad992

Observation cd8c9d4c-7258-4381-913b-ba42fc6b457d · outbound

This paper cites V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.773970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.773970Z digest=sha256:71bc5478e87e46cd909d35fe4a6758f1a82d1a2348337b7ede92598660504572

Observation 81112b9e-3a70-40e2-b223-891ac4bcf10d · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing31, 1720–1733 (2023).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion IEEE/ACM Transactions on Audio, Speech, and Language Processing31, 1720–1733 (2023)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.801617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.778338Z digest=sha256:bd1dcc8d012964fcffc7ed2fef607da0b3abf8a6f37151dcff3b1ccb7b29dbdf

Observation 4a26f552-1407-462e-b59e-f41d73cb198d · outbound

This paper cites In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.786691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.782497Z digest=sha256:6865b94992f1024d1b34c1f27787fb8d664f9b0f5b4a582b5915b2f9e17e84f1

Observation a5a43a36-4f0f-4c93-b81a-e3bc9e2bf2b8 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion IEEE/ACM Transactions on Audio, Speech, and Language Processing (2024)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.772134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.786388Z digest=sha256:f595e4c153ec3ccecf011f6015400945f410cb86ae099979ae183d9f2fec5165

Observation b15cb50b-ba18-4504-adfd-2d9bd3e953b6 · outbound

This paper cites In: Forty-first International Conference on Machine Learning (2024) 26.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Forty-first International Conference on Machine Learning (2024) 26

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.757133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.790367Z digest=sha256:a047bf41519cca0e0d3656504692d38d36eda8e9a71d76f3f0997a555cc05bc7

Observation d6e21cea-612a-485f-ab37-1bb8f76c625b · outbound

This paper cites In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.742362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.794702Z digest=sha256:9d8e9bfb9e87fdde7fbaf23ee5ab23ee7092be34e189860a9268197f5f202605

Observation 65c58431-bd3e-4ba4-8909-31fb2fe49cbe · outbound

This paper cites In: British Machine Vision Conference (BMVC) (2021).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: British Machine Vision Conference (BMVC) (2021)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.727619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.799087Z digest=sha256:9ab1e6bfd78ef752cc2031f529f4aa264c0748ae90bd4929f91f765930f48ad0

Observation 80fbd0f6-4583-4055-90ef-a5eae7d5d9d6 · outbound

This paper cites In: European Conference on Computer Vision, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: European Conference on Computer Vision, pp

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.712813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.803267Z digest=sha256:67a8281418a6b533e3bbfa4d2db9fb42f58b659f3c2d2d5886d4f1db056a2c27

Observation f0a71da5-dc07-4b08-8159-5abc0f61ece8 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.807382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.807382Z digest=sha256:30a5a7dab4934b481f705ada71a933ef63127e95bafff54b4a1f115c2b94de98

Observation bee83027-0efa-4327-a425-230eaa6c2862 · outbound

This paper cites Advances in Neural Information Processing Systems37, 128118–128138 (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems37, 128118–128138 (2024)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.812139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.812139Z digest=sha256:31e69bdaf3aed40c4d4875869d674027d2cef7a42d74a042a2b6832e40e0d8b2

Observation 10418fc3-bb32-4f0a-9106-8eb9122011e9 · outbound

This paper cites In: European Conference on Computer Vision, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: European Conference on Computer Vision, pp

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.688460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.816312Z digest=sha256:9f0b5215aa1f069b1476bbecea5d6046c369496f57f22cf63a38489f0bfe61ad

Observation 463212b0-d45c-4c69-91d8-1a837e07dc4a · outbound

This paper cites In: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.674109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.820327Z digest=sha256:7960a9c8af0a7ac2dcdb94197f9a41bb84eafe4629ee5f97e3bd3bd3ad6cbda5

Observation 293e2e7d-31bd-492c-a703-3dc06404301a · outbound

This paper cites Expert Systems with Applications249, 123640 (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Expert Systems with Applications249, 123640 (2024)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.660151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.824643Z digest=sha256:9b8349516031e20064ff47af18b16df984d63f7044d350b07f983b539d6fba9c

Observation 57fd110b-6e5c-4d40-a358-8c2ba84d2b09 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the Computer Vision and Pattern Recognition Conference, pp

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.646119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.829173Z digest=sha256:4ed2a648d4a01fe096031c1a59a9fe3d2d54b6d85f68f2598ef6b83912f705b1

Observation 1cd0ce97-2a41-4f14-b4d2-d1623e23a208 · outbound

This paper cites In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.630402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.833475Z digest=sha256:f3d0b8db384161594a33def56ef8de5e0126ca2dab2b8edc30b7b67eaaa2d08e

Observation b7d643b2-c701-4277-adfb-3cbfe477e85a · outbound

This paper cites In: Proceedings of the 32nd ACM International Conference on Multimedia, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the 32nd ACM International Conference on Multimedia, pp

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.616375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.837480Z digest=sha256:58f8552fc332ef6be6a9b0a3bea5a28933a7eb563499f3665bc6314e55dc3725

Observation f07b40b7-08f8-4344-aebf-3158027ea702 · outbound

This paper cites In: Pro- ceedings of the AAAI Conference on Artificial Intelligence, vol.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Pro- ceedings of the AAAI Conference on Artificial Intelligence, vol

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.601398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.841715Z digest=sha256:178739cda109a4cf48ee4a101d85eb23455ac776e809ef57837fad912f922174

Observation 57d58cc0-26c3-4354-937f-d7c441c7c267 · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence, vol.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the AAAI Conference on Artificial Intelligence, vol

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.587214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.845812Z digest=sha256:2386ebec56d13d15cb3fcbb3da3bfdc7e891d5268e02149c0291d30659859734

Observation a1f59108-06d7-4902-80fb-7dab094bf2af · outbound

This paper cites In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.573885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.850012Z digest=sha256:42e4ed7cb1b7a60f2cbda611b1dd5cf3d2b4208582426a7cb7264fd4c60cf19a

Observation 7dbec8da-ac2c-4f17-9efa-eb93f6857390 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.559378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.854403Z digest=sha256:412437a3c13dc8bbda4a03c564a4e84c4bc1388565f81913bdf2969a048fb0aa

Observation 06a8869e-afd8-497b-a31d-23ee28793e30 · outbound

This paper cites MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.858527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.858527Z digest=sha256:62b25d663fc17b32cfb5bb7bcc5caf9da8f62f87934520a3fccb0411d3d38640

Observation 6a83718f-d1e1-417e-9475-4ce5a1ed377c · outbound

This paper cites AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.863037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.863037Z digest=sha256:501e8520bae93324c5b13a556aa1c52c025ccab2065ee41897fe5f27214c17e3

Observation 394b22ae-1db9-415b-90bc-547dc7b013a3 · outbound

This paper cites an unresolved cited work.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:27:57.542854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.867299Z digest=sha256:2c496fa8e22bee005afc648de4195d298e7c849f8cd0dc6b3b1cc7cb8bcb64a1

Observation 7d27d25d-af74-4969-8c6c-92baafdd4d82 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.528185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.871072Z digest=sha256:838bb8c44f79964bedae8501739720fc019664a0c09221b3eb0ec560c2b9cf1b

Observation 509ca0ea-9b3f-4d54-bdc9-be1ffefaa7f0 · outbound

This paper cites Advances in neural information processing systems30(2017).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems30(2017)

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.874666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.874666Z digest=sha256:9705b2f8f4bcdba5e5382cf0a08335172010ef91f35a5863f1aae693bc6f5670

Observation b497b9f1-69dc-41ee-8ec8-eb70a53a3c0a · outbound

This paper cites Advances in neural information processing systems33, 3008–3021 (2020).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems33, 3008–3021 (2020)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.504639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.878323Z digest=sha256:ecb6d9af82660641e5e791c275f50b94492020032cd7c57a16baf10e90a64eaf

Observation f88d134d-575b-42b4-ace0-cfe37d0431ee · outbound

This paper cites Advances in neural information processing systems35, 27730–27744 (2022).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems35, 27730–27744 (2022)

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.881893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.881893Z digest=sha256:135f8a23a3867b966754a2976f22ac025e5f9d4366599e4aa7803083b007e83f

Observation 627e4055-c6d2-452d-b298-755be48300a0 · outbound

This paper cites Advances in Neural Information Processing Systems36, 53728–53741 (2023).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems36, 53728–53741 (2023)

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.480657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.885400Z digest=sha256:b9ef9e0ff2d47a4700fccd8988306d38eff81c7e7cb04506d17d03c4b521f872

Observation 017b35da-32a1-4171-bc5a-c986da2b7c18 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.889364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.889364Z digest=sha256:1e8d6c3670de5bd68905ff054d27eacb09cae894d37712dd76772ec85912ba82

Observation fa21302e-f917-4716-a010-b2d2a414ddef · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.893803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.893803Z digest=sha256:1184d367b3e8671b62803170e1b41bd13c14e41ea103955ed84de39a3372e845

Observation c22c0a52-147b-49a1-94dd-c87ecf787fdb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Constitutional AI: Harmlessness from AI Feedback

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.897942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.897942Z digest=sha256:2e341e13c5cd54be8b8b070be3ecf6b7f524d2bf9dd7a0e46233019574fed85c

Observation 1c5d24ae-298d-4fa4-9ae8-de5340b82c53 · outbound

This paper cites Advances in Neural Information Processing Systems36, 15903–15935 (2023).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in Neural Information Processing Systems36, 15903–15935 (2023)

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.902231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.902231Z digest=sha256:617d6565a2163e26784475eb955d7b6609519f0723bc9d85459aeb849350c6e3

Observation c8a9da8e-d582-4f0f-8579-e218e17ecd70 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.453843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.906308Z digest=sha256:be6ee54c257ba754d8a92b70f0347daf53edf0da4b617e927012d5f5766e663f

Observation 31449035-8014-4085-8416-fb6394ee92f0 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.910237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.910237Z digest=sha256:d627cd24a5090f56ce5d719d534d896d26b60bbdc6805acdd81e09642f283707

Observation b27ff933-b8de-4cb5-b04e-88ee0e7455fb · outbound

This paper cites Advances in neural information processing systems33, 6840–6851 (2020).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems33, 6840–6851 (2020)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.914138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.914138Z digest=sha256:7f39bf249d93772c7acb40d01a9795bcc03721ce73c99c88344e00ff6dc6788c

Observation c30dd17c-72f5-4c70-99ce-fb504e5ce0c1 · outbound

This paper cites In: Proceedings of the 32nd ACM International Conference on Multimedia, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the 32nd ACM International Conference on Multimedia, pp

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.421466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.918107Z digest=sha256:7dc923882d93ea9835ade57bddd6852f4a5f1b47cd13909309fbb0c83983fb31

Observation a3a0e9c7-dc9c-49e1-b335-482581e42b40 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.921927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.921927Z digest=sha256:d0384ed4873ff8d01403f42a6f649fcaf2e07738996374860ac8361ea9e6fd40

Observation da1d74bb-9cb2-4016-be8b-872fd52240e2 · outbound

This paper cites Neurocomputing568, 127063 (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Neurocomputing568, 127063 (2024)

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.927264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.927264Z digest=sha256:f0742de5992bb1dc45a8810141cfb32b62d6221e009f0de17c5d69c7b3abd1ff

Observation 291410dc-31fd-44ac-8e07-8a66cb894816 · outbound

This paper cites Journal of Machine Learning Research25(70), 1–53 (2024).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Journal of Machine Learning Research25(70), 1–53 (2024)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.397921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.931337Z digest=sha256:10ea87edb2cae163c349c21f227967f3f8fe898b3beece5014f4cd40a7fe9618

Observation 82135de5-b739-4df1-b12a-9d2478d93f9c · outbound

This paper cites Self-Taught Evaluators.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Self-Taught Evaluators

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.935166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.935166Z digest=sha256:2b5278e7c643def7dec453d231222045fe199a8bdfb09187c62eb61fa151c4b2

Observation 9a3c7be1-1c5c-46b4-a022-e8d7256a9cfa · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Contrastive Audio-Visual Masked Autoencoder

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.939307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.939307Z digest=sha256:272c2ae927af0aab40a821334c6323a95fffe6bb5e4d6ab43f634c623af3d85e

Observation 39b4b44d-c61a-4f15-a3e5-7edb5a2953c5 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.943597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.943597Z digest=sha256:dc545d6a117b5770adf1a9d8d30aa90a12db26ad133f0db4e30c1681595a31f2

Observation 3c9305fc-45c0-4fd9-915c-d21bf107b95d · outbound

This paper cites In: ICASSP 2022-2022 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: ICASSP 2022-2022 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.383585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.947797Z digest=sha256:9dd247192abf6ce721a4f5a328b72f036393c83052462ee2f21b681055c0730c

Observation 2ef3e406-8604-47d4-8e3c-55c6c4367650 · outbound

This paper cites Audio-Synchronized Visual Animation.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Audio-Synchronized Visual Animation

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:27:57.056509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.951652Z digest=sha256:98aaeddcfcdc611aa9e91a10fefddd29a4d63181c53ba39cdd52ec66d04c9172

Observation 1e167d94-415c-4193-8688-668b854a9410 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.955809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.955809Z digest=sha256:6dfd8b6a25659541a8a7312ba883ba6d0deeb8b88bbbb62052e04029d389008a

Observation c67033e6-21a6-441d-a757-4f5aacc88296 · outbound

This paper cites Taming Visually Guided Sound Generation.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Taming Visually Guided Sound Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.960058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.960058Z digest=sha256:ce6a91d6daf7ab4107da58d13c80afb28a6914a7cb20ef7f4402a1597cb3ce51

Observation 8e8120bc-b658-4ade-81b8-d6507c21881d · outbound

This paper cites Advances in neural information processing systems30(2017).

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Advances in neural information processing systems30(2017)

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.964446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.964446Z digest=sha256:551e89c134c046ff4701ce62bc317dc7a3bc24009f62f978a998bdaf71a0c713

Observation 8e43efb9-e006-433e-a0df-087fe1b17fc8 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.968322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.968322Z digest=sha256:a46bd7e96b78f3b4b3a0803f175a346b753e2b8e508ebee0909598ca99524379

Observation f85598bf-480d-45c5-bd31-b50260cd735e · outbound

This paper cites In: 2017 Ieee International Conference on Acoustics, Speech and Signal Processing (icassp), pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: 2017 Ieee International Conference on Acoustics, Speech and Signal Processing (icassp), pp

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.352105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.972384Z digest=sha256:92eb746b1ddce9c8e21ea4f1a437d53c17b8fc9ab60a5721b299535ff3ec93a3

Observation 75138c27-91a7-4047-8f9c-bab2aa7b1110 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.338428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.976488Z digest=sha256:32e75eec76c3a3ec7f95d690df9d3928a136c7645fed4e0637db5368c7212d12

Observation 4d3be700-3182-49c0-9d76-244937e34b80 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.324496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.980300Z digest=sha256:8afc16d4d7a98172e4db0eded52d7cfc399e192d7ed58cf6f1a6f6c6cc224566

Observation 33e1ab7a-c8cb-4dfa-af12-52064fc4aeb7 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:27:57.310492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:27:56.984255Z digest=sha256:dcf3823e91c6bda8c24dd2900b931645e1ca59ef5ea53e3b6f3d1cccf7c05982

Pith citing papers

No inbound Pith citation observations are available.