Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:44.257683Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2506.16020.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:44.257683Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:39.652813Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T23:53:44.519215Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 52b214e2-5f4f-4332-a795-61f44f5b65f1 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge As the continuous development of diffusion models [6–12], the naturalness and fluency of syn- thesized singing speech have now approached those of human performances
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1e4c4461-0d32-4f73-83e0-d8b00da30345 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 506c1bba-3090-49ed-9ec5-2322f40c60c0 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge test- seen
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 277be0cc-4485-447a-a6c9-487e60bc5453 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer consists of a modal interaction network, a decoder based on consistency Schr ¨odinger bridge and a spatially-aware feature enhancement module
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ab92c3dd-d228-4d7b-8c13-724f6c3a1018 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d4bddf8b-9472-435b-b305-b47b5ffe890b · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge 7) Sep- Stereo [35], a model that converts mono audio to binaural audio using scene images
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 52f1e380-921e-4605-b131-3406651a4f90 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Video diffusion models,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0f9af91-738f-4989-9de4-300298b27e02 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Diffsinger: Singing voice synthesis via shallow diffusion mechanism,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d98e1918-bcd3-461b-ba9e-42dd8f65eb67 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visinger2: High-fidelity end-to-end singing voice synthesis en- hanced by digital signal processing synthesizer,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6020805f-aa77-4bcd-80b9-29d618d53d66 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Audiogpt: Understanding and generating speech, music, sound, and talking head,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 09c9318b-5542-4054-b79f-2687f059b699 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge An End-to-End Approach for Chord-Conditioned Song Generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c088f08b-dec3-49b5-b3b5-ef3c8cc19328 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unisyn: an end-to-end unified model for text-to-speech and singing voice synthesis,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 773e026a-481e-4ded-9e08-7b799704bf1b · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Elucidating the design space of diffusion-based generative models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e481f1-25ae-4764-bc6f-92bfbadd4674 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Score-based generative modeling through stochas- tic differential equations,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c524afbb-bf39-4a08-a685-cf7d4cd61a08 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Learning the beauty in songs: Neural singing voice beautifier,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1d54c7ef-8809-4841-9688-b83bfd164986 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Align your latents: High-resolution video synthesis with latent diffusion models,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 99d8ab57-2439-4cf3-a06a-dcfcecd73e90 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge An image is worth one word: Personalizing text-to-image generation using textual inversion,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 199a1995-ba54-4999-876c-43f6e67dbf4d · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1f057377-e4ef-48be-9bb6-1bff537e9148 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Scalable diffusion models with transform- ers,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c9d916d6-6dcf-47b5-962b-70de28f610e6 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Grad-tts: A diffusion probabilistic model for text-to-speech,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 953b44ab-1be8-41db-a904-2aa718c0b47f · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Novel-view acoustic synthesis,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6b7b15d7-71e0-4583-aecd-21ce5fd23faa · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Como- speech: One-step speech and singing voice synthesis via consis- tency model,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 35035fff-63ed-4215-a90d-49f5ca2bf935 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Consistency models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ea639187-d110-4431-83f3-880b40a28e15 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 762799b2-c25c-48d3-be60-9f847c2a8a4d · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66992bb6-d77c-48a6-9a8a-0f57d0865c82 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge 2.5 d visual sound,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aafe6f73-5b3a-4638-86d2-35aef7133339 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Enhancing spatial audio generation with source separation and channel panning loss,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6f054485-c628-454a-8b32-7a6c2efb7461 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Multi-source spatial knowledge understanding for immersive visual text-to-speech,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation be610d70-f0ab-47ef-8361-5ed037b37824 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visually guided binaural audio generation with cross-modal consistency,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af230ae0-97c4-4739-8772-9a4b993b04f0 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Multi-modal and multi-scale spatial environment understanding for immersive visual text-to-speech,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c799f48e-d42e-4be5-9312-64db73ad783a · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Opencpop: A high-quality open source chinese popu- lar song corpus for singing voice synthesis,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6a410bcf-9fb6-404e-828c-75bea8fa0ee0 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Deep residual learning for image recognition,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c29d9c-fdd8-4f27-b17a-584b1dd191b9 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Diffusion schr¨odinger bridge matching,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dc9a9d56-171f-4a64-914a-df7ef556c927 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Likelihood training of schr ¨odinger bridge using forward-backward sdes theory,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1489dcf6-c11b-40ad-b9d3-f44e0dcf7b08 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Simplified Diffusion Schr\"odinger Bridge
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4713961d-e9ee-432c-b691-bece85ce1578 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge I 2sb: Image-to-image schr ¨odinger bridge,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6c5ec59b-df2f-4f53-b61d-e06d87539de2 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge The unreasonable effectiveness of deep features as a perceptual met- ric,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 69fd5a5a-6235-4337-99af-87d15d8e698a · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visual acoustic matching,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 82807bfe-5b9f-44c1-a1c7-768225211a68 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Ima- genet: A large-scale hierarchical image database,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1ab969e2-2260-4507-92e2-0b1980acaa83 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2e876622-dc92-42a8-bd94-09c3e13b7a74 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Self-supervised vi- sual acoustic matching,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation feda8574-5c2b-480e-bd39-342582047854 · outbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Sep-stereo: Visu- ally guided stereophonic audio generation by associating source separation,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ea639187-d110-4431-83f3-880b40a28e15 · inbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.