Pith. sign in

Paper Citation Record · LEDGER

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 2 inbound Pith citation observations for arXiv:2506.02414.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02414 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:41.644243Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:37.819963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T11:26:01.472807Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2df8c3c7-e932-4355-a298-55ec9b527926 · outbound

This paper cites StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:37.819963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:37.819963Z digest=sha256:699540e07d1d3e7b6e64cb80d5574e16ff6046970088493c2b37e2c986ceee10

Observation d0b4d78d-0dda-44f6-a87d-c6650ede0036 · outbound

This paper cites System Architecture Our proposed StarVC framework is designed to jointly model speech conversion and text generation in an auto-regressive manner.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion System Architecture Our proposed StarVC framework is designed to jointly model speech conversion and text generation in an auto-regressive manner

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:44.222921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:37.909035Z digest=sha256:9bb43fcda302a32610bf944f2d0aea9d8a701c40c32503c5dd7b9c34db2f193b

Observation dc1e7153-c55d-4ca3-902b-70bfe306697e · outbound

This paper cites an unresolved cited work.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:28:44.016396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:38.028580Z digest=sha256:9be476dd5cc902d833c67b3ea770bdd3c9104b1dfc8b5e96d5a0740f7185861f

Observation be2e4822-4938-407a-b34d-dcbfc2997b5b · outbound

This paper cites an unresolved cited work.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:28:43.575685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:38.232977Z digest=sha256:43dd85b76595be15ea3e91903522721e1ba231567b89a82cdf16bdb6bdbbedf4

Observation 97e42094-fa92-427b-9c45-7ef4b090d700 · outbound

This paper cites an unresolved cited work.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:38.357940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:38.357940Z digest=sha256:d1479a6372abb86f3f43b24eefc65a877231fa1c2d2827142a9fd0ff774faeda

Observation 0cd0cc7e-cfb5-4cac-8002-73ea19f1c9e0 · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:38.988059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:38.988059Z digest=sha256:284fc746d5c9b94909617111edb9ae9bbdadd4d4188fc87b1ca1cba88195c259

Observation 083bfd3a-15f4-40df-877c-9ff9bc513108 · outbound

This paper cites Continuous probabilis- tic transform for voice conversion,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Continuous probabilis- tic transform for voice conversion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:43.375754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:38.470719Z digest=sha256:9907b5300db3c8226c0b6833b17f7237be2660152409046f734ad5f1c337e484

Observation cd16316b-df7e-472d-9fc1-abb2b2823e3f · outbound

This paper cites Reimagining Speech: A Scoping Review of Deep Learning-Powered Voice Conversion.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Reimagining Speech: A Scoping Review of Deep Learning-Powered Voice Conversion

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:28:42.299528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:38.614095Z digest=sha256:292655b723bc82da624e1cf8ee1a80813a663240b805290ac42d9d413f325d83

Observation fcd138b0-e68e-4169-8bd6-f62d62ee0058 · outbound

This paper cites An overview of voice conversion and its challenges: From statistical modeling to deep learning,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion An overview of voice conversion and its challenges: From statistical modeling to deep learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:38.731606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:38.731606Z digest=sha256:b9de928de1cd5ac58140d69e6f9b1239367cdea5034b08cec7df0fcc05a6a90a

Observation 195e3a43-85d4-472c-a547-3a17620df70f · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:38.812994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:38.812994Z digest=sha256:a7a7eb8c2fed510c6650d46f49b665c8bd192b0dd718889af97ab8276c5bb8c7

Observation 18398b93-4b38-4768-a86a-7eac87ab7345 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:38.897234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:38.897234Z digest=sha256:79361175532b157e34122c728177c893ecdcf6ac5bc516e024212799e10ee38d

Observation ea4a4626-d600-476b-be27-4031ff45383f · outbound

This paper cites OpenVoice: Versatile Instant Voice Cloning.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion OpenVoice: Versatile Instant Voice Cloning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.528807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.528807Z digest=sha256:631c4c5dd7fcffee9ca7c44936694069f3734ad8229a9d37d1b07cf374b83624

Observation 046b8f8f-619e-4b20-9a7f-093be92c6539 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.064394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.064394Z digest=sha256:6cf436d9be76c11cff9a4d5efc471761edf43f3756bc8b20f9960be1cfe940cd

Observation 54b065c9-800b-4167-9be1-c71ec560cc6a · outbound

This paper cites Neural codec language mod- els for disentangled and textless voice conversion,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Neural codec language mod- els for disentangled and textless voice conversion,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.137014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.137014Z digest=sha256:aa5b89422936ed95952e2e55e8311392ecfc066671b9b774520cad6537718da8

Observation 9ddda902-e377-42e3-8b8b-28ce480d2a1e · outbound

This paper cites Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:43.142109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:39.230437Z digest=sha256:d83e52b800e7f20b8500cf97c392f6f57a9714109ee27ffa0efc2bf66900ba85

Observation b8bb26c0-d798-4ad7-830a-f0d1443dacdc · outbound

This paper cites Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.352136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.352136Z digest=sha256:6d6de39050c73c733feb4a2c25ce3934cf481c786f9da6fa1fbec8a981588744

Observation 3b16057f-5c85-4021-8d28-792b3fb0e1d0 · outbound

This paper cites StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.448947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.448947Z digest=sha256:17f3d34e7847496f7a9b5e7d266b82c3a0ae2929c72174ce61a5c19686443662

Observation 08679d1f-dc81-4a25-aba9-1d489ffcb276 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Robust speech recognition via large-scale weak supervision,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:40.063613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:40.063613Z digest=sha256:83c784f2ccdbbf97c02e6463aebd7c94fb20a390749102d9300ed0b96f1e9619

Observation fd1cb69f-cfcf-4442-9706-a202b435bc43 · outbound

This paper cites Lm-vc: Zero- shot voice conversion via speech generation based on language models,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Lm-vc: Zero- shot voice conversion via speech generation based on language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:42.954005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:39.621560Z digest=sha256:ca65d7845514a17d402912f63d6ab512cd65e0f2509cabf6c2c71123a7bc70ce

Observation 36b70be6-0df8-4d65-bf04-5b5a7e66200c · outbound

This paper cites DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:28:42.047432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:39.747987Z digest=sha256:70d7a7c2b466d7a11e221ed095dd6413794d3a1ec34aae66fc60be311f1ab4ee

Observation af841cf4-fb46-4464-a0d0-427ca241a113 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Moshi: a speech-text foundation model for real-time dialogue

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.846201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.846201Z digest=sha256:2168a9d2cef069c0cd9b47ec64ae606ddd49744db4455269de41dc1859c883ea

Observation 43b79292-ecd0-4ef7-8638-a3368fa496ff · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.892026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.892026Z digest=sha256:0faf969ba804c63e31c463237f13e20308746a423a262227dc15c5edf0a9a52b

Observation a38ced2f-e6bf-428d-9ab4-bc652ad434f4 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.960930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.960930Z digest=sha256:0161cb728f340fad8281e1c5e36ed9f14e44606a94354e1cecfd80e619c24896

Observation f4cc7b23-095b-4950-b013-3e40c7124f70 · outbound

This paper cites SNAC: Multi-Scale Neural Audio Codec.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion SNAC: Multi-Scale Neural Audio Codec

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:40.523017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:40.523017Z digest=sha256:3b5ef20210c84886f5c09331e1070c064952135796c498f2e86bdb4634b92238

Observation 076302f5-652f-4c82-81b6-a0d019d2fe83 · outbound

This paper cites An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:40.158378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:40.158378Z digest=sha256:74638bdcbbfe7097378266aabeb4bcefeb11b7ba7496d3b35f7048f7b5e3fa63

Observation 8966055c-0600-48f5-842f-1b7441585887 · outbound

This paper cites Simple and controllable music gen- eration,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Simple and controllable music gen- eration,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:42.734742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:40.207492Z digest=sha256:93d74a22ecf64aca2529a5bc39b6eca9de460ae1a05370304c31e153c7582127

Observation 5a2f2a56-bbbb-481e-ae2d-1f7952068cde · outbound

This paper cites A learning algorithm for contin- ually running fully recurrent neural networks,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion A learning algorithm for contin- ually running fully recurrent neural networks,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:42.649884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:40.293600Z digest=sha256:941b70751f3e80095b67de4cdd27a8e328de3c29a6a3e1a8c1dd2fc7aa219a71

Observation 7ffb179f-d376-4940-a5e0-4b2a70c8daa0 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Roformer: Enhanced transformer with rotary position embedding,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:40.350975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:40.350975Z digest=sha256:b395a714ab6db55cbc1e4f09585b70da0001e225862a3d4acfe06ecd6211264f

Observation 3af7947a-e4ec-4113-bfab-1b49e1d85185 · outbound

This paper cites High Fidelity Neural Audio Compression.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion High Fidelity Neural Audio Compression

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:40.451256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:40.451256Z digest=sha256:bcc652f78ce6a60cba3a261cc2b0d3623cb2f8ea3f34ba287c9968ad991eadab

Observation d21c2aa6-c7f5-4489-9a50-e5509b91f13c · outbound

This paper cites GigaSpeech covers diverse domains and acoustic conditions, while LibriTTS is cleaner and more consistent.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion GigaSpeech covers diverse domains and acoustic conditions, while LibriTTS is cleaner and more consistent

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:43.750906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:38.125123Z digest=sha256:3317bf55ce5e13069bc046032644b6f906454ce0292cbf8e124bfe9d72951f64

Observation 79e7c961-e095-421a-9902-fe11c8cc0f37 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion High-fidelity audio compression with improved rvqgan,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:40.567207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:40.567207Z digest=sha256:2a0f6ed49ff6d931882de9cd61e38231ea18891e0143ce60e697315d8db94d88

Observation 1b25b703-7d5c-4ef4-82f1-5e2b0b7e8614 · outbound

This paper cites SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:40.718286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:40.718286Z digest=sha256:858b0b451a98668bdf068d7a8c4b8cb40d44b5d8e75b02ece01c0422aa7b63db

Observation ea9bd571-082f-46ac-a8d4-4cb9ca20937e · outbound

This paper cites Attention is all you need,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Attention is all you need,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:42.529219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:28:40.798667Z digest=sha256:314fff286ec55fda5828a8f0ad219b9b779d6118e5064e7c4cdcc939673cca59

Observation 7c0f7bc1-a43d-481f-b55b-e24c4e539021 · outbound

This paper cites Qwen Technical Report.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Qwen Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:40.863401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:40.863401Z digest=sha256:e13c7cd5a939da0d3c26dfcf34e2b88cf674fd66923c039e282c01e2e08fdf15

Observation bc33c64e-9538-4134-81bf-c2465e948553 · outbound

This paper cites Qwen2.5 Technical Report.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:40.987790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:40.987790Z digest=sha256:725ae2f5b10e8eee84d72cde4b319bbae21b7b09e896e4e47a2419651d06d771

Observation 3d25116b-720e-4bed-b17b-d3100c804aa7 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:41.066944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:41.066944Z digest=sha256:df5c1e43bab32ac817a9d3211e0b65da7280a6f6ae9648ed70db73e2f01d9b06

Observation bc77ec91-1540-4805-b166-5340f8844ec1 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:41.167406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:41.167406Z digest=sha256:216cb06ac1f076d5a076f64f5c1cc2db6f14d13840b5b71922bf5531d7b75473

Observation 7a88d34e-8cf9-4b50-9f13-e91273c79b3e · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Lib- rispeech: an asr corpus based on public domain audio books,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:41.253071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:41.253071Z digest=sha256:b4bf7357e9d72c38f4b278c7b113d486afc89346beaa2993d8ebb4fdfbd5e4da

Observation be796505-843a-434d-988c-72a53d0e1f13 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:41.354168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:41.354168Z digest=sha256:bbfe8a5ea3acaa0b1d985b9a5db10fab08364e7d09cb682f4d3c70cf55f65199

Observation 8f54052a-3d2c-46cd-a514-9554e24980f9 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:41.470407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:41.470407Z digest=sha256:af6da900a0ac2d63102b829c0a4654000956523b7d0e0599839fa2e327f0f37b

Observation 148c0d86-6b60-40c9-842f-178b7be97ab7 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:41.552257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:41.552257Z digest=sha256:ddd54834a15f8f35435954ffe6514786fa9144572e6cadd6781b0eebd9d85858

Observation 7c2b4fcd-d666-4124-ba7a-1ca69d74738e · outbound

This paper cites Triaan- vc: Triple adaptive attention normalization for any-to-any voice conversion,.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Triaan- vc: Triple adaptive attention normalization for any-to-any voice conversion,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:41.644243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:41.644243Z digest=sha256:213968e60e4dfef0b31d6cfa7b62bbd28fe8c20ecc28e810a2fb5974b7d42126

Pith citing papers

Observation 2df8c3c7-e932-4355-a298-55ec9b527926 · inbound

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion cites this paper.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:37.819963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:37.819963Z digest=sha256:699540e07d1d3e7b6e64cb80d5574e16ff6046970088493c2b37e2c986ceee10

Observation cbaacff1-6a1a-4a1e-872c-bb26dfe00aea · inbound

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora cites this paper.

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:01.479201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:57:07.894455Z digest=sha256:a26b396a7f140768fffb44bf72d1df96691288afb26a84413345ecd21fe628bb