Pith. sign in

Paper Citation Record · LEDGER

Conditional Flow Matching for Visually-Guided Acoustic Highlighting

As of 4 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2602.03762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.03762 v4

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:55:17.332539Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dcb018c5-9232-4a81-a9c8-8b5c8182fbc9 · outbound

This paper cites Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.097919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.097919Z digest=sha256:9ad72d6681bf09126b1197a993b8b391e28290992ae71d50a0c4ddbbdd57b534

Observation 08ea8b95-f93b-4af3-9784-02b2e185a856 · outbound

This paper cites Diffusion-based unsupervised audio-visual speech enhancement.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Diffusion-based unsupervised audio-visual speech enhancement

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.131508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.131508Z digest=sha256:63b24c3d416dbd8404633f03d4b5f86aaa0a2b33cd28a64eb18db75b08ad368a

Observation 19ca2430-ded3-471b-89c0-f7a67c50220d · outbound

This paper cites D-Flow: Differentiating through Flows for Controlled Generation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting D-Flow: Differentiating through Flows for Controlled Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.149156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.149156Z digest=sha256:93bc03de94acc87d4e56b044ccaee36bd568085ff8ab48d0d86ab3a73e6ee8d7

Observation f89b7e4e-0178-499c-b6f5-b0f3ed87aa35 · outbound

This paper cites LBM: Latent Bridge Matching for Fast Image-to-Image Translation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting LBM: Latent Bridge Matching for Fast Image-to-Image Translation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.168179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.168179Z digest=sha256:a721c365153b136d318ef71dffd798b0fb67f784586b1d4ce1b1cbe4cf797126

Observation a4567355-30c9-4c85-b6b2-9ce257e9c62d · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101,.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.189245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.189245Z digest=sha256:9c493b70fe2687a38a8ca67905e6f592f9ca2d3b4bdf4f5cfc4227d637dfd06e

Observation 83a732a9-6a78-4526-803b-3781a7c6fdaa · outbound

This paper cites Mmaudio: Taming multimodal joint training for high-quality video-to- audio synthesis.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Mmaudio: Taming multimodal joint training for high-quality video-to- audio synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.262685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.262685Z digest=sha256:2a88343f97a409b2a45e0ca6c88ecd3857708ed7ad5f70e713129aa7724a876b

Observation c99690fe-1d2d-460e-b6c9-a5130cd37165 · outbound

This paper cites Samwise: Infusing wisdom in sam2 for text-driven video segmentation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Samwise: Infusing wisdom in sam2 for text-driven video segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.354766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.354766Z digest=sha256:c5a942afec36a5c37c0c28802c4afc9a55c5ab4ad36a04c78e1bf5401adf6086

Observation 3ee101ae-a0ae-4677-839c-aa2058ef089c · outbound

This paper cites CROME: Cross-Modal Adapters for Efficient Multimodal LLM.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting CROME: Cross-Modal Adapters for Efficient Multimodal LLM

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.422928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.422928Z digest=sha256:b7d9e7ffe0138d137a2bcb629e7baeeed0e5611046ffc5e6e5ce3fc0f08cd917

Observation 1ad81d4e-55b3-46b5-a4ee-483eba98a371 · outbound

This paper cites Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.603924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.603924Z digest=sha256:e9fbdf70a7c20d4a27101dbb9b84aed2a7eb0d3143f2b91683b63e0c3517fbe8

Observation a78143f0-78d2-4b52-9ba6-bb9c1bffff07 · outbound

This paper cites Visualvoice: Audio- visual speech separation with cross-modal consistency.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Visualvoice: Audio- visual speech separation with cross-modal consistency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.728994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.728994Z digest=sha256:fb788567576f7b8bc1f5e22e114b6cc486208e5baac49489b3efade0c5b9411a

Observation 143aba14-314d-47ca-baef-0edf86a03602 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Imagebind: One embedding space to bind them all

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.774161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.774161Z digest=sha256:cf258d6e96a35c7608ac0d1d64d0ca21608ab1fe3910515c25bb22b628cdf863

Observation 25ef3e7e-e350-43d7-9d91-fc5da5c7b93a · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.863400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.863400Z digest=sha256:4411296e413a00daf475707f18b1b88d82fa6fd899c8a9e944366abfaaae830a

Observation f25a9529-024e-4e1c-a241-30dc4b7b91b0 · outbound

This paper cites Davis: High-quality audio-visual separation with generative diffusion models.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Davis: High-quality audio-visual separation with generative diffusion models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.940550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.940550Z digest=sha256:3fb646ebfb686fc9f01b1a8553ec910c4f86c772252b8fc7cd82eff120b9db2f

Observation 7a27e098-c6e5-498e-8e75-4cb3cdfe25eb · outbound

This paper cites Learning to highlight audio by watching movies.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Learning to highlight audio by watching movies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:12.958618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:12.958618Z digest=sha256:f6c887e2d27564aa76e2c48b37dc7e39eb4d55b59820d0949276a622824a3d16

Observation 11cc946d-4341-43f8-a33e-dc0b241f76d7 · outbound

This paper cites High-quality sound separation across diverse categories via visually-guided generative modeling.arXiv preprint arXiv:2509.22063, 2025.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting High-quality sound separation across diverse categories via visually-guided generative modeling.arXiv preprint arXiv:2509.22063, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.017673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.017673Z digest=sha256:3fc101ac60565a02ef0be2809e38d9ab998f486b4307c7eda965f48f70a46c68

Observation 7e4996ba-7a04-43d0-aeaf-d7ec6b020c74 · outbound

This paper cites Mu- sic mixing style transfer: A contrastive learning approach to disentangle audio effects.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Mu- sic mixing style transfer: A contrastive learning approach to disentangle audio effects

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.071397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.071397Z digest=sha256:6d5639244b1b8c869dfef66190b6adecdbcf47067403f4528e62309947dc630e

Observation 4bb54ab7-7168-4529-8112-4002e8e33a81 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Efficient Training of Audio Transformers with Patchout

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.142515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.142515Z digest=sha256:a798e8eeda1593f1637ce4cbdc3a988598d11d8eaf1a164e010145a35e69e4d9

Observation 7c76bbe8-888b-41d7-9408-841c3c8ccc79 · outbound

This paper cites Seeing through the conversation: Audio-visual speech separation based on diffusion model.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Seeing through the conversation: Audio-visual speech separation based on diffusion model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.251521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.251521Z digest=sha256:7a2887e5944c40bdea766856c6c5eb03e0c7b70846f1beb6a90f74d4e79e6beb

Observation e2498842-f216-4347-84a1-d231f5a3a3c8 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Detecting mo- ments and highlights in videos via natural language queries

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.292769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.292769Z digest=sha256:50fd2f5c1b1f740455b97f96ad29dabc261396467b45c97b8bd008cef2dbceac

Observation 4032d820-1c23-4e1b-b88f-adc2b083a3a4 · outbound

This paper cites Storm: A diffusion-based stochastic regen- eration model for speech enhancement and dereverberation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Storm: A diffusion-based stochastic regen- eration model for speech enhancement and dereverberation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.431879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.431879Z digest=sha256:2d0b28c2577c45eeb2332cd9505634098d02b17240b877740315658785976314

Observation 1325f259-c34c-4a88-902a-89b905692995 · outbound

This paper cites Av-nerf: Learning neural fields for real-world audio-visual scene synthesis.Advances in Neural Informa- tion Processing Systems, 36:37472–37490, 2023.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Av-nerf: Learning neural fields for real-world audio-visual scene synthesis.Advances in Neural Informa- tion Processing Systems, 36:37472–37490, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.509646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.509646Z digest=sha256:ddaa599c6bf2a6bbd80d1805f027233d6fe58a3622aec89dcdd96aa2ab562ba3

Observation 5af4acf7-3558-434d-b64b-16e6e8c26f89 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Univtg: Towards unified video- language temporal grounding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.587719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.587719Z digest=sha256:51601ea041c3b0c586aa71e340d3de26482f58ba3eebc77bd17e468632f525df

Observation bde8db0e-c462-420b-a5e4-81f68314e2d4 · outbound

This paper cites Flow matching for generative mod- eling.ICLR, 2022.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Flow matching for generative mod- eling.ICLR, 2022

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.690488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.690488Z digest=sha256:3225e2759f4eef817e4fbc27555615169eef4dab1a1638c44fcf75373f060e0f

Observation ec5615f8-feae-45b2-9ded-c61ff83b4433 · outbound

This paper cites Let us Build Bridges: Understanding and Extending Diffusion Generative Models.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Let us Build Bridges: Understanding and Extending Diffusion Generative Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.895260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.895260Z digest=sha256:e36d183fc47863bbe4fe473146baec7cc93e8d982163354bd7dec6fcbfeee55c

Observation 15152251-b812-44af-86a0-8634980cf528 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:13.974182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:13.974182Z digest=sha256:5b745074d27c72dbb245218915df3f7e4c7fbba6868ec80b24629990299f6d4a

Observation c5eb4f62-7979-4937-bd29-6a3a0031149f · outbound

This paper cites Multi-Source Diffusion Models for Simultaneous Music Generation and Separation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Multi-Source Diffusion Models for Simultaneous Music Generation and Separation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.055110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.055110Z digest=sha256:b26c95ef2052dce86e2ab579494ff4067831f009a964f0e3193627c3f66c2f52

Observation 2aaf0f95-2f27-453f-97de-f7c52ae69e47 · outbound

This paper cites Deep learning for black-box modeling of audio effects.Applied Sciences, 10(2):638, 2020.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Deep learning for black-box modeling of audio effects.Applied Sciences, 10(2):638, 2020

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.108968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.108968Z digest=sha256:9044fd45c3d22f4d41b2ad092d22920259e63faf76ebaf0f7b0e3422fc589cdd

Observation 27f79caa-f928-4cbe-936e-a2e2eef9de04 · outbound

This paper cites On discriminative vs.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting On discriminative vs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.181056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.181056Z digest=sha256:2c87c5d57fca8033745f77092725aa22d71a14176e247a0f0ee6fff7ac570dc6

Observation a6282e1a-01ce-4ad4-baba-6c29edad9be9 · outbound

This paper cites Elucidating the Exposure Bias in Diffusion Models.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Elucidating the Exposure Bias in Diffusion Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.323365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.323365Z digest=sha256:83c70a09cbc8fac66acb650273b0fa86b021ad9baf499fe20abde97f8776d579

Observation 9ddd07c8-a515-4eaf-b561-b5ec4802fcd3 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Movie Gen: A Cast of Media Foundation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.391930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.391930Z digest=sha256:31bbd78d769b0976a47f6c635fb3ca920d08da4faa48cf88e0f377d2412b7279

Observation dd8812b6-b87c-4afe-817e-5aef23f19540 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.470045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.470045Z digest=sha256:61e0212f571e213e394cff14cebe9ba3305252b961643cbf6c251ada821b0e59

Observation 592b5884-4994-487d-ae83-9357c62d0de2 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.524917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.524917Z digest=sha256:85d357d1156238cfb4a767b33697490f20a3b27651e3b6fa72d9df244d4b5ec7

Observation 2e9a38d3-88af-41b0-a582-42c43f24a9f0 · outbound

This paper cites Model- ing nonlinear audio effects with end-to-end deep neural net- works.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Model- ing nonlinear audio effects with end-to-end deep neural net- works

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.623311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.623311Z digest=sha256:d43d97ebd9b3340fbd9595c61b49007e916d2aa5e4776e57ef312ecc928b11c3

Observation 288dfbfc-eaa4-4799-aec9-1c05ecc2283f · outbound

This paper cites Sequence level training with recurrent neural networks.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Sequence level training with recurrent neural networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.694302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.694302Z digest=sha256:319a683554a56ecb6ef06e98a8ff031691520b313c638d5f06b5d06e2023e911

Observation 93ab2f1e-aa85-41d0-a44b-e80d6cfb9217 · outbound

This paper cites Hybrid transformers for music source separation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Hybrid transformers for music source separation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.809643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.809643Z digest=sha256:75b768c5399815129031477f04a3de6eaf230a8a36007a3b108937d25a041c5e

Observation 197db680-9a1e-48f8-93be-37fa82b8cf1a · outbound

This paper cites Source Separation by Flow Matching.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Source Separation by Flow Matching

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:14.983041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:14.983041Z digest=sha256:2d885ae1aae846a52429b6c135d711acf6abde9ad90134a1c0697ed185dbe8fc

Observation fa012a8f-be05-4833-bb6f-42ea517e43be · outbound

This paper cites Physical modeling using recurrent neu- ral networks with fast convolutional layers.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Physical modeling using recurrent neu- ral networks with fast convolutional layers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:15.237244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:15.237244Z digest=sha256:2d9217a6496fdf12f8f252a8e44cd1cf2eb60caa6b74cf0be473d051fad61eb2

Observation c630fc7d-8219-4c77-ac26-309a4281b438 · outbound

This paper cites Improved techniques for training consistency models.ICLR, 2024.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Improved techniques for training consistency models.ICLR, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:15.305094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:15.305094Z digest=sha256:d303143fbf58dfcef04e291738112c4c24b87322ed5e3414fb5e212d41a3ef51

Observation ebdabe01-dbd0-4a2e-b301-fceaabf8fcd4 · outbound

This paper cites Consistency models.ICLR, 2023.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Consistency models.ICLR, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:15.406955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:15.406955Z digest=sha256:6b2e5589fae44b721a9be0588e32014498e42650b75f82c0560ea11ed5c63607

Observation d891c011-1753-437d-b09c-2cda18d979d6 · outbound

This paper cites Style Transfer of Audio Effects with Differentiable Signal Processing.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Style Transfer of Audio Effects with Differentiable Signal Processing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:15.462383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:15.462383Z digest=sha256:ef9939f44b8487f8e9f3d44e48d0cae072bbaea0deb706fb40694650a49ab873

Observation f18f7a01-4fde-4a93-9624-9b5b3f9bc50b · outbound

This paper cites AudioX: A Unified Framework for Anything-to-Audio Generation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting AudioX: A Unified Framework for Anything-to-Audio Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:15.558525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:15.558525Z digest=sha256:169986ecf7db89f8ea434360633893ed6899d497a5f7f34a37e5555deeeddb61

Observation f6ab1385-76bb-4a28-ac5f-a7c752010573 · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:15.774393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:15.774393Z digest=sha256:6b15e83a39c56e2c6edfc0db3d5180928615815aa9a5b6d9ebc0f44b9d8a2c65

Observation 12509606-6a23-48f7-9f45-9512df7e4ff0 · outbound

This paper cites Diff-MST: Differentiable Mixing Style Transfer.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Diff-MST: Differentiable Mixing Style Transfer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:15.879850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:15.879850Z digest=sha256:ab2408e4a20f706879c38a8b4118e07957d14e08cfdab2fc19f05b66a2c45500

Observation aea663de-c192-4579-b61b-aa59a341e6c3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:15.963640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:15.963640Z digest=sha256:0ddce5e449064379a658ecd0bb9a5247c7351a58f75ac00d548578b833577e7a

Observation ef11026f-595f-4d30-ae2f-30c670372a06 · outbound

This paper cites A generalized band- split neural network for cinematic audio source separation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting A generalized band- split neural network for cinematic audio source separation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:16.048112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:16.048112Z digest=sha256:23b7e88b54b3bb407b0db6a9281cdbb0ba602b5eecd7a86a6cac128849c4c7f1

Observation b9dd051e-3f3e-48d6-bbcb-ccb401d3e7fb · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Chain-of-thought prompting elicits reasoning in large lan- guage models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:16.166636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:16.166636Z digest=sha256:2bdfe981f778894033d8b3982400dfd2e04254a4477d383b0691a17736a8792c

Observation 05f88f90-8aa9-4455-be9e-d7dc700f3c89 · outbound

This paper cites FlowDec: A flow-based full-band general audio codec with high perceptual quality.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting FlowDec: A flow-based full-band general audio codec with high perceptual quality

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:16.382839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:16.382839Z digest=sha256:10f65dc78f2a7f6d2c9057889ec55a8e2f0b74fb9b1f9d07e1545a4fa74f90ce

Observation afc3ee98-8ab6-4c6c-bf36-a2c85a489682 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:16.447053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:16.447053Z digest=sha256:98ab3b50ac83d34d4b3974f42a008fade21e4da69e24b31c4002e880e38a34ad

Observation 6721cc05-3066-4c0e-88c3-2c24cb29a106 · outbound

This paper cites Visually informed binaural au- dio generation without binaural audios.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Visually informed binaural au- dio generation without binaural audios

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:16.506970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:16.506970Z digest=sha256:5f23c7e1005c4bcf53c5e3689f3bf665c0caff116cb2415eeea7f593531f55cd

Observation ec389530-180e-4dca-bab1-8d74577351a9 · outbound

This paper cites The sound of pixels.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting The sound of pixels

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:16.656865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:16.656865Z digest=sha256:eee131a480d1a1bdfb4e6a799de58e776c8cd607b927bf35070aad3baedb00e5

Observation 2caee42c-1691-4b03-9001-0b39b5981904 · outbound

This paper cites Table 5 presents the results obtained on this variant of the dataset.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Table 5 presents the results obtained on this variant of the dataset

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-03T04:55:16.766369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:16.766369Z digest=sha256:b6c27d4c919895433b3d8c8e5fe56d5619e836a44d64f01855a563f2c3754f44

Observation f330bd86-2e44-4b8f-8974-bc84a5b27041 · outbound

This paper cites an unresolved cited work.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Unresolved cited work

Reference 52

Resolution
malformed identifier
no resolver link, observed 2026-08-03T04:55:16.958526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:16.958526Z digest=sha256:154e2b1f67d5a52a8ea0ec83b3ec4552fa7dfecb6cfdaffef486634428e70c3c

Observation 16b6620d-24dc-48b5-86e1-864cdc738dcc · outbound

This paper cites an unresolved cited work.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Unresolved cited work

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-03T04:55:17.027290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:17.027290Z digest=sha256:100bda08f785253769052ca9f22f8e4bf1e76b2fea723bba828e6bc620fa1b46

Observation a1b604da-29fa-4b7a-8bd0-9fb6722788e3 · outbound

This paper cites However, when its contribution becomes too dominant, the trajectories exhibit non-linear behavior again.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting However, when its contribution becomes too dominant, the trajectories exhibit non-linear behavior again

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:17.132619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:17.132619Z digest=sha256:891f963d43c669831cf92b94e3b9f304d7558c0bca1482b6cbb0ba275ed868fd

Observation d74ba8d7-358e-42fc-90e9-6285000803a9 · outbound

This paper cites The rollout loss allows more consistent predictions across steps, resulting in more highlighted sources.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting The rollout loss allows more consistent predictions across steps, resulting in more highlighted sources

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:17.204233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:17.204233Z digest=sha256:9d4368241bc123f9e210f48fd0b7759d6e6b99e5632224ee6d93c572e45e3e64

Observation c82e4946-497c-4455-b200-dfdeacfc13d7 · outbound

This paper cites Future work should evaluate the model on real-world data once such datasets become available.

Conditional Flow Matching for Visually-Guided Acoustic Highlighting Future work should evaluate the model on real-world data once such datasets become available

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:17.332539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:17.332539Z digest=sha256:ec268a60d8b3fb8d81fae7bf960b39ebeab289dc6892253e924d96dd157e940a

Pith citing papers

No inbound Pith citation observations are available.