Pith. sign in

Paper Citation Record · LEDGER

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification

As of 8 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2506.14226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14226 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:39.522256Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:34.275071Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:23:40.027554Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b9f8a7c-1599-4d43-b0e3-545013616521 · outbound

This paper cites Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:23:40.118664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:34.275071Z digest=sha256:d0ca39f15a49062a91fcf3b64d3a41eac35d3c4406cd8fd1a0b33529a6b975cc

Observation 9973e933-9aa8-47a3-94dc-ffd25c2d0d66 · outbound

This paper cites Figure 1 illustrates the entire workflow along with implementation details.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Figure 1 illustrates the entire workflow along with implementation details

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:46.916252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:34.359464Z digest=sha256:50bae2f9eff8f15697a288f9ffcf12e964d08c9f5d087f8ee4e0e0d241ecbac3

Observation a44aaad9-7fe4-4cab-80fd-52a8424a1a83 · outbound

This paper cites an unresolved cited work.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:23:46.580564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:34.476186Z digest=sha256:2dace62dc8c9b144979b615aafc68738b7a082faaca4259977edd729b9bd14d1

Observation 7fa08d6c-1160-481a-b3c6-1ba4cfe359b1 · outbound

This paper cites In this chap- ter, we will analyze the results from four aspects: TTS models, speech durations, text content, and fusion techniques.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification In this chap- ter, we will analyze the results from four aspects: TTS models, speech durations, text content, and fusion techniques

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T00:23:46.156485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:34.652055Z digest=sha256:53d7c29d0b5c8dc7fd8d3f3e10f2d46e814bf6e72258db507ab2bd94be2e6b8e

Observation fc16b29f-4270-402a-9526-bd1d131046f7 · outbound

This paper cites an unresolved cited work.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:23:45.791905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:34.817804Z digest=sha256:1c5821ab24490f91056eb70da88d2d5bb0a65bc9af4b3f08d4caa25db3f19499

Observation cfdc2aa2-fad8-43ee-897c-2718a8beda32 · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification X-vectors: Robust dnn embeddings for speaker recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:45.438945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:34.955023Z digest=sha256:cb743de76bc3a3e5a2ad8a2a99b001af5e44defce3f59fe7dd3b031ed30c9906

Observation 297f5257-1324-411c-8c10-a9896aae4609 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:35.159496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:35.159496Z digest=sha256:3e0b54d89b15569a70bb8af0cfba0ce7795d96a67d343d0580dbf9a1e897207a

Observation f97ac33d-b4d3-4252-9882-0c1fa44a9ae7 · outbound

This paper cites BUT System Description to VoxCeleb Speaker Recognition Challenge 2019.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification BUT System Description to VoxCeleb Speaker Recognition Challenge 2019

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:35.266115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:35.266115Z digest=sha256:41d7296615aac2c4dc7c8f3d03a4f135cb053783890ff7e321bfd593a8196edb

Observation 10c0fe75-1074-40f5-8838-8a34627ac098 · outbound

This paper cites Mfa-conformer: Multi-scale feature aggregation con- former for automatic speaker verification,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Mfa-conformer: Multi-scale feature aggregation con- former for automatic speaker verification,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:45.099236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:35.483658Z digest=sha256:bfc34495cedae16ed46cb4ff5384d9121eb07e4d99bcbb82063138f330c36839

Observation 24f4ab30-36b1-4060-8fde-d6ed1896065f · outbound

This paper cites Memory storable network based feature aggregation for speaker representation learning,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Memory storable network based feature aggregation for speaker representation learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:44.808818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:35.675854Z digest=sha256:64ade77e5a6c35fa4cee719a53ecb4896c39a12f897babebd96a929b9cb05176

Observation 0ec4f438-5efa-4479-9acc-4b2a69076981 · outbound

This paper cites Whisper-pmfa: Partial multi-scale feature aggregation for speaker verification using whisper models,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Whisper-pmfa: Partial multi-scale feature aggregation for speaker verification using whisper models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:44.356554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:35.834899Z digest=sha256:67b6b7100e56b06553775ec441aedcebbf2c4075d0afb96db34484f2fea4f516

Observation 8bdc73cb-126e-45c9-b724-6b5d1ba3a97c · outbound

This paper cites Cam++: A fast and efficient network for speaker verification using context- aware masking,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Cam++: A fast and efficient network for speaker verification using context- aware masking,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:44.041649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:35.961690Z digest=sha256:7448672e6bd2cb4d35070c93f5273b604fb7353cda4c4a7fdddcbd7e981aeff5

Observation 151a4959-cb8f-428f-bf76-24f60cc67119 · outbound

This paper cites ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:36.116032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:36.116032Z digest=sha256:d4ffeca8a4bb315136a37062e7c0fe55159bd78082e0f44eeaf747a7e4e5e855

Observation d30666ed-781b-4b9e-a053-dd84df28f6e6 · outbound

This paper cites Overview of speaker modeling and its applications: From the lens of deep speaker representation learning,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Overview of speaker modeling and its applications: From the lens of deep speaker representation learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:43.653995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:36.293870Z digest=sha256:b6ba34e9212f689e35ac5443e405ecafd795ef47080007954481e554834acdce

Observation d3bc998b-e31d-40b5-aa07-a8fa0904985b · outbound

This paper cites A Deep Neural Network for Short-Segment Speaker Recognition.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification A Deep Neural Network for Short-Segment Speaker Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:36.448875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:36.448875Z digest=sha256:edecfa13acfe647892c04caa6e2a7f5b604139041dde4b821d77e998bf4872ae

Observation ee3d4389-ef59-4375-bc69-a325f536e4a1 · outbound

This paper cites Deep speaker embedding learning with multi-level pooling for text-independent speaker verification,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Deep speaker embedding learning with multi-level pooling for text-independent speaker verification,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:43.255824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:36.597158Z digest=sha256:a6349c1872bf17e9b14a40a1e7a82c4eaedf24e7f69a36c232ec7fad5d4e3959

Observation 1c5f8383-47a0-4eb4-a978-36769352e694 · outbound

This paper cites Improving multi-scale aggregation using feature pyramid module for robust speaker verification of variable-duration utterances,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Improving multi-scale aggregation using feature pyramid module for robust speaker verification of variable-duration utterances,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:42.940747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:36.708396Z digest=sha256:7864090e23c0886b062748828a29cecca19d445adfae92405e8f5cabc8e0a86c

Observation def8d12d-299e-4f01-94fb-17fcdc48ecaa · outbound

This paper cites Improving aggregation and loss function for better embedding learning in end-to-end speaker verification system.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Improving aggregation and loss function for better embedding learning in end-to-end speaker verification system

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:42.523611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:36.907874Z digest=sha256:64b65f36bcc16c566ddb23e1e29f54a5c0ecf2406f650ecf3b53ae7584e03250

Observation de3e5762-cc57-4d0a-a919-a4f3bf82ac27 · outbound

This paper cites Meta- learning for short utterance speaker recognition with imbalance length pairs,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Meta- learning for short utterance speaker recognition with imbalance length pairs,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:42.158144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:37.123479Z digest=sha256:2cd510fbfccc8b629a4d8287e2e25842836a36a5f61df30a1f6a1873a40cdc3c

Observation cb082490-694a-442f-afce-26d928525058 · outbound

This paper cites Text-independent speaker verification with adversarial learning on short utterances,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Text-independent speaker verification with adversarial learning on short utterances,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:41.927984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:37.380968Z digest=sha256:502dcda3b4a2d2e6b89958797299d1e1177002c8cce2fb098aba479857142c6c

Observation ba0ee5bc-5b31-4819-90cf-2da52a2f0e78 · outbound

This paper cites Short utterance compensation in speaker verification via cosine-based teacher- student learning of speaker embeddings,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Short utterance compensation in speaker verification via cosine-based teacher- student learning of speaker embeddings,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:37.552964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:37.552964Z digest=sha256:794fba4ae8e69aed4812178748b4ca4571aca89b17c306d4160669f403364e68

Observation e984d1ca-eb9e-47e3-8dc6-c5de88d29acc · outbound

This paper cites Open-set short utterance forensic speaker verification using teacher-student network with explicit inductive bias,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Open-set short utterance forensic speaker verification using teacher-student network with explicit inductive bias,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:41.719782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:37.669130Z digest=sha256:59dc6c9288df556b344cd95bc37046ddcfc128e1b857c44744a3a39e7b67cdf6

Observation 73aa36b4-de62-4249-9755-cc41a372d421 · outbound

This paper cites Cnn-based joint mapping of short and long utterance i-vectors for speaker verification us- ing short utterances.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Cnn-based joint mapping of short and long utterance i-vectors for speaker verification us- ing short utterances

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:41.469790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:37.792093Z digest=sha256:e68ce5415e4b3602c2251e9908c064168821faf91f0102f8f77f3015e5a81265

Observation 13daafee-65ca-4e1c-b3a5-a1dbea829d88 · outbound

This paper cites I-vector transformation us- ing conditional generative adversarial networks for short utterance speaker verification,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification I-vector transformation us- ing conditional generative adversarial networks for short utterance speaker verification,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:41.261227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:37.963583Z digest=sha256:c90a3c3cdf3fc251445492a844f2dbada3a9672d6c334c484c43532df4deed4d

Observation 2d22789f-ddf8-4565-884f-55962c6a5249 · outbound

This paper cites Data augmentation using deep generative models for embedding based speaker recog- nition,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Data augmentation using deep generative models for embedding based speaker recog- nition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:41.118213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:38.178009Z digest=sha256:58862c1ad27ff740cba6fcb13ec4378b77d27c7f11f0f325a08fc6e09cc2bbb9

Observation eb19a3e5-de23-4c3e-97b2-2d27d19b2232 · outbound

This paper cites Exploring Voice Conversion based Data Augmentation in Text-Dependent Speaker Verification.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Exploring Voice Conversion based Data Augmentation in Text-Dependent Speaker Verification

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:23:39.839245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:38.319326Z digest=sha256:9f6fb518eb9a11791f10e8fbbd2f720a2f2168ea7a27ce459d760434aaef439d

Observation 27676973-d98d-4a4b-87fb-31005c1d8553 · outbound

This paper cites Synaug: Synthesis- based data augmentation for text-dependent speaker verification,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Synaug: Synthesis- based data augmentation for text-dependent speaker verification,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:40.915478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:38.481487Z digest=sha256:a4c791d710aa128847fb2907c53c5d7b4eddf9f353d0bb6d816a0529863e4539

Observation 8ea03eae-a2b2-4f40-8d73-c1b3035e4ac6 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:38.647520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:38.647520Z digest=sha256:755d922a141b0bd55b53929e307dbfc7bce7e7967ccebf3909425c5f925ac02d

Observation 361f4ba3-eb41-495f-acab-c941748f4b61 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:38.830041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:38.830041Z digest=sha256:ae04600c302ef39048042a54b2a217b55cb86a1656e2f8da3c220eccb92ba107

Observation 51d3e78e-d0b8-4bbf-804e-24cab88dc804 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:38.994437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:38.994437Z digest=sha256:a48c2eca6b668f92ebca09b28640bd9b1d60fe0305e91ed6a66d5c757301fc81

Observation 37eb4ea3-c883-4700-810d-680a1b0903a4 · outbound

This paper cites V oxceleb: Large-scale speaker verification in the wild,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification V oxceleb: Large-scale speaker verification in the wild,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:39.162765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:39.162765Z digest=sha256:0b8692724935060a1fcf2b15030d7850fb6143b09ce02757d7e5adceb2efdbe1

Observation 9061c3ca-6c0f-4f46-9b81-48cbe11d2ac7 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:40.645490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:39.300553Z digest=sha256:85af74fcc68862cb024d520019bc48452516a60c0de9a49c7cfa828b8f60d70a

Observation 22e0c373-417c-4506-8d24-986b3f37aa1c · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:39.416451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:39.416451Z digest=sha256:24e9dea3a2a0f22396261548c6aaaf84f4ceb44e633fbd9c2f093a84f2a49199

Observation 10f541d2-dcb5-47fb-b289-9d0f53cbba39 · outbound

This paper cites Naturalspeech 3: Zero-shot speech syn- thesis with factorized codec and diffusion models,.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Naturalspeech 3: Zero-shot speech syn- thesis with factorized codec and diffusion models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:23:40.326165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:39.522256Z digest=sha256:9a45e24654187aa08d6aeffa6230dcba949da22e990e29b5e0bc50447680394f

Pith citing papers

Observation 0b9f8a7c-1599-4d43-b0e3-545013616521 · inbound

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification cites this paper.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:23:40.118664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:23:34.275071Z digest=sha256:d0ca39f15a49062a91fcf3b64d3a41eac35d3c4406cd8fd1a0b33529a6b975cc