Pith. sign in

Paper Citation Record · LEDGER

Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2303.03926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.03926 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:50.926162Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.393626Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c31f0988-d576-4e59-9964-a9dc3a1e4237 · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 272

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:52.719788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:63c27898a342d8ef9155f7909a62071e5f8ddd67fa6f70b57bff673a778c6130

Observation 979bebc8-0d3e-45fd-8ca7-22bb17eac15d · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:26:37.352136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:5981ee1c06a07425df44c3ca621679b46bd33952162cb93b6f0ea408c1fb033b

Observation 8685f695-97b5-4632-b1b4-a12fb11e367b · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.570236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:a1481f4045c08cf0707eecce7a4fcc1453a8e5e2e78672fb9b0ff28e9a35f443

Observation 240147d3-4c80-4585-b45b-8d478dd37758 · inbound

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits cites this paper.

Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:50.926162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:50.926162Z digest=sha256:96430f45683fdc1cdadc72146bb510db9a588c184bd64a2399e06da51da33151

Observation 64fe4a74-6347-4730-9265-6015099d6104 · inbound

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding cites this paper.

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:15.604779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:15.604779Z digest=sha256:1ac02c026981a0da387d1099b0730c06370fb32d9152118803b90ed3497229e2

Observation 7e1872e8-7f8e-468d-bf75-a5acd1a15bc1 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.462342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.462342Z digest=sha256:3b9d2cd6e68a3420c573ceb47798dfc84e9cdf76efc08f95ee97f718d4f70ad3

Observation fab4a743-7c6c-42bb-a1ab-0ef3ae3e293a · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.227783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.227783Z digest=sha256:4464e069f6e78bc44a43cc1b1f7e497d4f7fb0e7bdf62c94cfc5b591746c4ffa

Observation f148466c-20d9-4340-a9ed-b9aedfd7a3f6 · inbound

Voice Adaptation for Swiss German cites this paper.

Voice Adaptation for Swiss German Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:07.320866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:07.320866Z digest=sha256:ca63363b36679446d34ecbcc798380d8e5dafcefff3bccfd0dccfb192521e00a

Observation 7ab4c217-426d-4035-a07d-ab117c4c1ffc · inbound

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes cites this paper.

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:31.620520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:31.620520Z digest=sha256:31221b0b4ca47da84a8055e7c0c05b88e584c43f1e262de3629d20b10786bea4

Observation 9a7b921a-32ce-4f59-9765-7ed4dbf538f6 · inbound

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation cites this paper.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.091198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.091198Z digest=sha256:391b969a14908715a2ee5ddfa3e598694a22f44750db734dc2cadfb1ae13c156

Observation 09225bfb-1e78-4efb-bf6e-b9e4b684cb17 · inbound

Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages cites this paper.

Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:30.568112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:58:30.568112Z digest=sha256:93476cb40b32c47f18734d5cb507ddcd5752ca46f702cd6c631f85bfaab6e471

Observation e58d9ddf-40d0-4d10-893d-ff661ecaf7eb · inbound

OpusLM: A Family of Open Unified Speech Language Models cites this paper.

OpusLM: A Family of Open Unified Speech Language Models Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.334752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:36.334752Z digest=sha256:a138781b55105ac93f52d070077edb11198706c65cb05d41207a3c33209ac364

Observation b87f9175-72f9-4052-ba6c-4c17668cba53 · inbound

Differentiable Reward Optimization for LLM based TTS system cites this paper.

Differentiable Reward Optimization for LLM based TTS system Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:03.927154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:03.927154Z digest=sha256:b93e3304b61f809c431db918b0ced63a970eb06c3b6b369096302671719580c9

Observation 78b9d3b1-68ef-4d18-ae2e-3aabd768eb0f · inbound

SecureSpeech: Prompt-based Speaker and Content Protection cites this paper.

SecureSpeech: Prompt-based Speaker and Content Protection Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:36:07.615777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:36:07.615777Z digest=sha256:aed0dc9ffcbe90bc680d82f9b1f582d4e89d10ed8d5496d679f90a1f402d1692

Observation d4517a79-8173-4f65-8551-ffcd886f9f98 · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.950994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.950994Z digest=sha256:ee953d5afb602ebf165073b9e68cfd63f5837dcb93c1a946a6dbfdf74eaaceee

Observation 64a33c59-d966-4a47-b022-18342097d1b9 · inbound

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation cites this paper.

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:17:38.301601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:17:38.301601Z digest=sha256:3c59ab8e9e584434b81fb0f1e9abfa32f8cfba574673bff575585ad040c4e9e7

Observation d9c9843c-e0b0-4d37-9727-f0540ccd7d32 · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.883671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.883671Z digest=sha256:b572590c6a84beb69f574604c922f28a7e3bfb91a0afec01a2b8cf9775f4accd

Observation 9dc06219-a9e9-4be9-b4e1-e28e44f6aed5 · inbound

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching cites this paper.

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:21.136655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:51:21.136655Z digest=sha256:420796e71feadcfed798cd1617e9e751bd66d62229cdfc150ad93538648e2800

Observation c93a56ef-7037-49ee-95c9-33e0a5070a7b · inbound

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance cites this paper.

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.826038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:34:15.220263Z digest=sha256:e534cc54432682e04a76a1b27566515cd90eb1354dabf9d8bd09e0f027b72177

Observation 488ae25c-b8e7-4ae2-b441-d6cc262f33b1 · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 87

Resolution
malformed identifier
no resolver link, observed 2026-08-03T00:11:18.974337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:18.974337Z digest=sha256:289046c17deb84e8313636068dd3fbb3a213ef5c1154e15d7518e763fdb3fcb4

Observation 84b0e2ae-172b-42a0-a79f-bcfdf1f1c229 · inbound

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization cites this paper.

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:59:59.597146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T11:56:13.914121Z digest=sha256:3a7fa63e0e8ed769b72fc8b874af2426ff928fdf27093a3024683bdc59be4b25

Observation a5fd5a40-14db-4ffa-b6e1-e1094da1f026 · inbound

StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection cites this paper.

StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:05.181855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:43:16.998414Z digest=sha256:7c7cd481749be66463ff2e5441aabf94d5fafe041205b63a1537b1423e46a4b8

Observation 3a70b231-b3aa-409c-87f2-0ca65bf922f5 · inbound

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation cites this paper.

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.473621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:36:07.090627Z digest=sha256:88479a22a5e1acc9a34e13e3668dfe76a15dd5b1bad9beac170642b06f12cf4d

Observation 22772b80-26cd-4144-8f72-ad74a903bf89 · inbound

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations cites this paper.

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:13.516306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:16:47.522489Z digest=sha256:93406ff9017180b3e1cac25828da313f4974b4232ec143ba2009fb427bf66500

Observation dddbbbb0-4ae5-4e67-9b5d-be997bff803b · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:16.042449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T04:35:58.032597Z digest=sha256:b9922a3448d3be01745d187b73ade484e8327949f7fe368b5c64b0e033d69154

Observation 8abf31f5-c8c4-4ebf-ba4c-3459173fd0dd · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:13.915078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T01:43:48.555523Z digest=sha256:07e94930ecde09e0258aa01970af4e7f9dcb9ef8bb72cdc490559e0ae04c6072

Observation 11430ff3-1019-4358-86c2-e879e4a3d3d3 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.333971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:82a51070513086c23de75b3a6450c858a1768ae86dbf9044b581b4a11446a5f5

Observation f3ef0fae-32a5-48a0-bb78-c852c98aca91 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.842971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:6a2569766d37a84f17740dee15e32a57151945d6bf053c01e031ab4aab335e71

Observation 03a2684a-eeb9-4fbf-8e07-d08c633d1569 · inbound

Do speech foundation models perceive speaker similarity as humans do? cites this paper.

Do speech foundation models perceive speaker similarity as humans do? Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:57:04.871209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T00:06:34.841809Z digest=sha256:2a72ed949b3645c6ba72109c5d4475f1485c9f24e800bba8421bc39f4601ca74

Observation b51ec42e-5a70-42bc-bcd1-f49e35a5c523 · inbound

UniVoice: A Unified Model for Speech and Singing Voice Generation cites this paper.

UniVoice: A Unified Model for Speech and Singing Voice Generation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:08.306137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T23:56:14.198308Z digest=sha256:d249824d528939451d49928089a452554cd72ec49c91d6749a186a9d4ee65af4

Observation eff54b84-0495-4ced-a6c8-cb7fbd651d5d · inbound

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations cites this paper.

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.156796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T20:21:47.681999Z digest=sha256:91b02b70822e62e45ce881210995a244bb9795419637271cf33f15b664cf2bbd

Observation 23bea1f1-e3f1-4cb7-a49c-5a9264df1192 · inbound

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation cites this paper.

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:39:57.036118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T03:26:22.529879Z digest=sha256:3ff297e47961dcb175349f08fc260179a2daf4092314356b8c4a894f9158d156

Observation 03b696fa-7ee1-4897-96cd-1ece412df834 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 235

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.394850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:f6603ebe2bebef5963f0223e7fa7c6e27288bd5b63246baeb3734c2ba5deb25c

Observation e6fa88f4-b57b-4f87-ade2-356a59d72afa · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 235

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:a657cfe2e7c7294b664194abe34e07e82112a41ff0dc55d095197e5bd195a1eb

Observation 487ed536-8ec9-4116-a1b7-b26693617478 · inbound

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection cites this paper.

Component-Level Ensemble Fusion for Speech and Environmental Sound Deepfake Detection Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:50:08.442744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:50:08.442744Z digest=sha256:da3bdd39a4b470c7b21e34e865436257c1d8814429a9c09c32c4c781ac42319a