Pith. sign in

Paper Citation Record · LEDGER

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

As of 12 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2506.16020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16020 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:44.257683Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:53:39.652813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T23:53:44.519215Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52b214e2-5f4f-4332-a795-61f44f5b65f1 · outbound

This paper cites As the continuous development of diffusion models [6–12], the naturalness and fluency of syn- thesized singing speech have now approached those of human performances.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge As the continuous development of diffusion models [6–12], the naturalness and fluency of syn- thesized singing speech have now approached those of human performances

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:53.294893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:39.570507Z digest=sha256:fea93faeb452377096462444bc46d68ed06a5753ada65bf9222f27024ad640ce

Observation 1e4c4461-0d32-4f73-83e0-d8b00da30345 · outbound

This paper cites an unresolved cited work.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:53:53.052980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:39.747869Z digest=sha256:ccc4d54b97305e69d6f46c55f51ef5d8b45255a70a27746e7591ef2d1b823b96

Observation 506c1bba-3090-49ed-9ec5-2322f40c60c0 · outbound

This paper cites test- seen.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge test- seen

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:52.830799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:39.855427Z digest=sha256:373f5eb5df37643518b1a76c45cf9d87e16f16f0c5961a308cdfb59f8d3c2c59

Observation 277be0cc-4485-447a-a6c9-487e60bc5453 · outbound

This paper cites VS-Singer consists of a modal interaction network, a decoder based on consistency Schr ¨odinger bridge and a spatially-aware feature enhancement module.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer consists of a modal interaction network, a decoder based on consistency Schr ¨odinger bridge and a spatially-aware feature enhancement module

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:51.984203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:40.204764Z digest=sha256:302ba52059c5b2fb317469101fd4a67c8ef14acc1538f3a0ef880e263a399b33

Observation ab92c3dd-d228-4d7b-8c13-724f6c3a1018 · outbound

This paper cites an unresolved cited work.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:53:52.547797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:39.987942Z digest=sha256:014fd2a37d326d6ed3c74bd9e6b5168ab30427454c6a43c46636eb62929ad918

Observation d4bddf8b-9472-435b-b305-b47b5ffe890b · outbound

This paper cites 7) Sep- Stereo [35], a model that converts mono audio to binaural audio using scene images.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge 7) Sep- Stereo [35], a model that converts mono audio to binaural audio using scene images

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:52.268194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:40.128440Z digest=sha256:9bef93b53682b534c77f5f998b98a20476dc27f8ad3e3cb11e166e23318be9ca

Observation 52f1e380-921e-4605-b131-3406651a4f90 · outbound

This paper cites Video diffusion models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Video diffusion models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:40.946806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:40.946806Z digest=sha256:463564dc9b340a66302908e98c7421f5425db90f2b27166b8f85526751e3d495

Observation d0f9af91-738f-4989-9de4-300298b27e02 · outbound

This paper cites Diffsinger: Singing voice synthesis via shallow diffusion mechanism,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Diffsinger: Singing voice synthesis via shallow diffusion mechanism,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:51.637949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:40.290437Z digest=sha256:a7590c024e7ecb827efa4aee4b0f85d6a140ec1f06e0cdab66eff42753546c5f

Observation d98e1918-bcd3-461b-ba9e-42dd8f65eb67 · outbound

This paper cites Visinger2: High-fidelity end-to-end singing voice synthesis en- hanced by digital signal processing synthesizer,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visinger2: High-fidelity end-to-end singing voice synthesis en- hanced by digital signal processing synthesizer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:51.216181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:40.400396Z digest=sha256:adcd1f620716e01e0553d353017846d789eb0e1b08b7ae32b4d8ce9b6fe8397d

Observation 6020805f-aa77-4bcd-80b9-29d618d53d66 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Audiogpt: Understanding and generating speech, music, sound, and talking head,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.874815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:40.478105Z digest=sha256:f9d24f9e59d53e34d4ec9ac2f2fc2a98b30cfd8531837d5de93880fea0982488

Observation 09c9318b-5542-4054-b79f-2687f059b699 · outbound

This paper cites An End-to-End Approach for Chord-Conditioned Song Generation.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge An End-to-End Approach for Chord-Conditioned Song Generation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:53:44.430171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:40.561586Z digest=sha256:3abfa6ff21fdb9126c0d90b99394205393046733500eb37ec0d38eda410a5067

Observation c088f08b-dec3-49b5-b3b5-ef3c8cc19328 · outbound

This paper cites Unisyn: an end-to-end unified model for text-to-speech and singing voice synthesis,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Unisyn: an end-to-end unified model for text-to-speech and singing voice synthesis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.765455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:40.663144Z digest=sha256:6844fe3d6ba93ff8aa400cf4b9f142c12a35b6724767ad5df14567c0d1a219a1

Observation 773e026a-481e-4ded-9e08-7b799704bf1b · outbound

This paper cites Elucidating the design space of diffusion-based generative models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Elucidating the design space of diffusion-based generative models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:40.784747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:40.784747Z digest=sha256:f2286008e67d118bfe0a6c477bb8e645b4759236db31c43c4f8eb1c866aaafec

Observation f3e481f1-25ae-4764-bc6f-92bfbadd4674 · outbound

This paper cites Score-based generative modeling through stochas- tic differential equations,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Score-based generative modeling through stochas- tic differential equations,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.046653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:41.697669Z digest=sha256:317caf0923c7cf4f6b94162be9a44fb89f6140423fff85447d2c721c3ac631dc

Observation c524afbb-bf39-4a08-a685-cf7d4cd61a08 · outbound

This paper cites Learning the beauty in songs: Neural singing voice beautifier,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Learning the beauty in songs: Neural singing voice beautifier,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.499591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:41.075668Z digest=sha256:bdd586403086979fd968468e52f1e18c86c2a769d2d0b7fb0813fed3aaf49f81

Observation 1d54c7ef-8809-4841-9688-b83bfd164986 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Align your latents: High-resolution video synthesis with latent diffusion models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.317937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:41.185792Z digest=sha256:4ea61c56c9e6fb01497a5c1d10778c75456ef820d0dded8fd16f0e73c16ab4c9

Observation 99d8ab57-2439-4cf3-a06a-dcfcecd73e90 · outbound

This paper cites An image is worth one word: Personalizing text-to-image generation using textual inversion,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge An image is worth one word: Personalizing text-to-image generation using textual inversion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:50.033407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:41.315721Z digest=sha256:e67ed4963b9d74dd8a6a832dda5bc3f4b8d7b47e1429eddba689d227e0f6f34d

Observation 199a1995-ba54-4999-876c-43f6e67dbf4d · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.783454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:41.409879Z digest=sha256:23fd5296b99db9618c67f3410f66ea2b50f5b3c702aaaf886c3e6f15bf9f3bfa

Observation 1f057377-e4ef-48be-9bb6-1bff537e9148 · outbound

This paper cites Scalable diffusion models with transform- ers,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Scalable diffusion models with transform- ers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.548243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:41.541143Z digest=sha256:3b89b3c0121f74c12d70b63bf95284891c3124111468e9674c81b7db632e3e5f

Observation c9d916d6-6dcf-47b5-962b-70de28f610e6 · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Grad-tts: A diffusion probabilistic model for text-to-speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:49.292215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:41.601848Z digest=sha256:b834dde759742fdd20a96536ccca725fe68dc069da19f11face33c17a1c5893c

Observation 953b44ab-1be8-41db-a904-2aa718c0b47f · outbound

This paper cites Novel-view acoustic synthesis,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Novel-view acoustic synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.582153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:42.602814Z digest=sha256:0576bd04b97f29b0a2f7dd8d809dbd9c3c96ba4f665eb8b437eee1e77e9c4cd4

Observation 6b7b15d7-71e0-4583-aecd-21ce5fd23faa · outbound

This paper cites Como- speech: One-step speech and singing voice synthesis via consis- tency model,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Como- speech: One-step speech and singing voice synthesis via consis- tency model,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.915571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:41.831441Z digest=sha256:e89920f85d14231433f422e66e53df70f47dbb28c2c409cfd0359cfb35ef4f04

Observation 35035fff-63ed-4215-a90d-49f5ca2bf935 · outbound

This paper cites Consistency models,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Consistency models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.652424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:41.994604Z digest=sha256:cff11f0c695aadfda5efc4ef64755946f9baa96226c966656ff46c837ed0c255

Observation ea639187-d110-4431-83f3-880b40a28e15 · outbound

This paper cites VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:53:44.614311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:39.652813Z digest=sha256:4f6383802814b5445d7636b0a6cbfc8805bfdab2bb7d0d4a1bbea81a5776e034

Observation 762799b2-c25c-48d3-be60-9f847c2a8a4d · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:42.127548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:42.127548Z digest=sha256:c41f8b7c66b6faec0398f6d9410defff19c1ef2461af2756c9da4f3456967262

Observation 66992bb6-d77c-48a6-9a8a-0f57d0865c82 · outbound

This paper cites 2.5 d visual sound,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge 2.5 d visual sound,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.339568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:42.249129Z digest=sha256:595ca73c8879e2837ab5b787928668a4d2eaca551f86366516cd67a531953e0e

Observation aafe6f73-5b3a-4638-86d2-35aef7133339 · outbound

This paper cites Enhancing spatial audio generation with source separation and channel panning loss,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Enhancing spatial audio generation with source separation and channel panning loss,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:48.071677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:42.339805Z digest=sha256:0a95421314f8bf8028290da76d8622802324a351294965ade92c889e23679c8a

Observation 6f054485-c628-454a-8b32-7a6c2efb7461 · outbound

This paper cites Multi-source spatial knowledge understanding for immersive visual text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Multi-source spatial knowledge understanding for immersive visual text-to-speech,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.807211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:42.456963Z digest=sha256:50488315794191db1d36f8c6612948e8b9694a638370feb7fe7b42b217ab9002

Observation be610d70-f0ab-47ef-8361-5ed037b37824 · outbound

This paper cites Visually guided binaural audio generation with cross-modal consistency,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visually guided binaural audio generation with cross-modal consistency,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.365066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:42.753463Z digest=sha256:ee9746b5eb58ff954d1bf9528ed0f0e628d40c142761f72ebdf5146660119102

Observation af230ae0-97c4-4739-8772-9a4b993b04f0 · outbound

This paper cites Multi-modal and multi-scale spatial environment understanding for immersive visual text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Multi-modal and multi-scale spatial environment understanding for immersive visual text-to-speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:47.055369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:42.846974Z digest=sha256:bb0b7d1e482ce2bede1d4797e7030e07178b9e34b8579533e93f36912835487c

Observation c799f48e-d42e-4be5-9312-64db73ad783a · outbound

This paper cites Opencpop: A high-quality open source chinese popu- lar song corpus for singing voice synthesis,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Opencpop: A high-quality open source chinese popu- lar song corpus for singing voice synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.792475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:42.944905Z digest=sha256:0778814fdb104e99c14c2bf9ac7855a9b4a810259c4cbf12448b0671759fb9cc

Observation 6a410bcf-9fb6-404e-828c-75bea8fa0ee0 · outbound

This paper cites Deep residual learning for image recognition,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Deep residual learning for image recognition,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:43.052802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:43.052802Z digest=sha256:803818da6f4059cde13437cf74c57a777102adfde7fe7f2f8afe38946ea70c6b

Observation e5c29d9c-fdd8-4f27-b17a-584b1dd191b9 · outbound

This paper cites Diffusion schr¨odinger bridge matching,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Diffusion schr¨odinger bridge matching,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.625428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:43.170695Z digest=sha256:bf336c4743e0c63ae92520edc33ccc15b9263a1505ca0ab54018c5f6fec4bd03

Observation dc9a9d56-171f-4a64-914a-df7ef556c927 · outbound

This paper cites Likelihood training of schr ¨odinger bridge using forward-backward sdes theory,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Likelihood training of schr ¨odinger bridge using forward-backward sdes theory,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.384510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:43.271303Z digest=sha256:0a3ea607df1fcd5489b52971f2bd4e66458c76f0d27dd1ca7a9b2faac000091f

Observation 1489dcf6-c11b-40ad-b9d3-f44e0dcf7b08 · outbound

This paper cites Simplified Diffusion Schr\"odinger Bridge.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Simplified Diffusion Schr\"odinger Bridge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:43.454000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:43.454000Z digest=sha256:fa809fe59a975dabb36997c48dbc7bf95b49e9899d7f16df5d38cabc8785ad38

Observation 4713961d-e9ee-432c-b691-bece85ce1578 · outbound

This paper cites I 2sb: Image-to-image schr ¨odinger bridge,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge I 2sb: Image-to-image schr ¨odinger bridge,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:46.162553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:43.577638Z digest=sha256:cd018230fcef8da0cf4a093a4b512b687bea907bc20b1710c86466a4db9fcb7f

Observation 6c5ec59b-df2f-4f53-b61d-e06d87539de2 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual met- ric,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge The unreasonable effectiveness of deep features as a perceptual met- ric,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.994718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:43.667055Z digest=sha256:03587652fb94eebe8d1d70740d46656ca3683829d0a8b88b87158d41fe5d5186

Observation 69fd5a5a-6235-4337-99af-87d15d8e698a · outbound

This paper cites Visual acoustic matching,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Visual acoustic matching,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.754757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:43.819082Z digest=sha256:a96846da6036c4bee02646e134b70c14e0318b3c4e2eb3120e25b5c8668e0c2f

Observation 82807bfe-5b9f-44c1-a1c7-768225211a68 · outbound

This paper cites Ima- genet: A large-scale hierarchical image database,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Ima- genet: A large-scale hierarchical image database,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.514262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:43.931528Z digest=sha256:7b81dd74a09b3d0cd1f7d8ea73fc5169e1c1d8bab4b60e6b89a63a22932d8f73

Observation 1ab969e2-2260-4507-92e2-0b1980acaa83 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.249542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:44.050995Z digest=sha256:2edfb35a8dee892cd8f34294501070513bd5784c7bb0a3d4429ef2cb8aa08193

Observation 2e876622-dc92-42a8-bd94-09c3e13b7a74 · outbound

This paper cites Self-supervised vi- sual acoustic matching,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Self-supervised vi- sual acoustic matching,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:45.071008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:44.171093Z digest=sha256:236e60f01dfdaa0407c249aff9875537e8b074dfb24d99c2dbcd6312f3066ac0

Observation feda8574-5c2b-480e-bd39-342582047854 · outbound

This paper cites Sep-stereo: Visu- ally guided stereophonic audio generation by associating source separation,.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge Sep-stereo: Visu- ally guided stereophonic audio generation by associating source separation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:53:44.843365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:44.257683Z digest=sha256:c99233571487871039326571f36b58a417f772e544d87099dc9d4b5fa4484de2

Pith citing papers

Observation ea639187-d110-4431-83f3-880b40a28e15 · inbound

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge cites this paper.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:53:44.614311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T23:53:39.652813Z digest=sha256:4f6383802814b5445d7636b0a6cbfc8805bfdab2bb7d0d4a1bbea81a5776e034