Pith. sign in

Paper Citation Record · LEDGER

Adaptive Duration Model for Text Speech Alignment

As of 9 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2507.22612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22612 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:09.630347Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:34:07.882293Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:34:10.011767Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact2
  • verified fuzzy12
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 769360fc-7bb6-4889-b742-f172efc4899d · outbound

This paper cites A typical TTS system includes an encoder, a decoder, and an alignment mechanism linking linguistic and acoustic representations [3–6].

Adaptive Duration Model for Text Speech Alignment A typical TTS system includes an encoder, a decoder, and an alignment mechanism linking linguistic and acoustic representations [3–6]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:12.419908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:07.829227Z digest=sha256:d5e03d33c80de09d0af195c0434a87b50f60fbaba051360ebd843658a0f5a146

Observation e822238e-0ba3-457f-bee2-0dc2c323cba7 · outbound

This paper cites Adaptive Duration Model for Text Speech Alignment.

Adaptive Duration Model for Text Speech Alignment Adaptive Duration Model for Text Speech Alignment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:34:10.121918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:07.882293Z digest=sha256:af91b4de2bf44095e39ea65370b8e41595966f84a2944dc0eee20f21a9478d59

Observation 50b46a03-b979-44fc-86c3-0b452df8c76e · outbound

This paper cites We use Premium and Basic subsets of Wenet- Speech4TTS as our experiment dataset.

Adaptive Duration Model for Text Speech Alignment We use Premium and Basic subsets of Wenet- Speech4TTS as our experiment dataset

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T11:34:12.271952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:07.942094Z digest=sha256:42e1c6c413800b2e14c71ddbd22154091a8d07800c1de40b95ed639716c8801f

Observation 458ac7de-2dbe-4a37-80c5-c02a70b6ea43 · outbound

This paper cites Durformer outperforms baseline methods with respect to efficiency and accuracy.

Adaptive Duration Model for Text Speech Alignment Durformer outperforms baseline methods with respect to efficiency and accuracy

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:12.084758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:08.046867Z digest=sha256:30110e15155c395771ee8863d17ed9def89bbe02bf60b4db6bf6526897e763ca

Observation 0e2b5885-beda-4b19-8418-c051f936a75c · outbound

This paper cites One tts align- ment to rule them all,.

Adaptive Duration Model for Text Speech Alignment One tts align- ment to rule them all,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.898860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:08.114869Z digest=sha256:b808409d52d920dc775d51ab858fb1850707589fa13fdbc17ad95af6cbdcd519

Observation 0c5d9294-2c5c-4955-820b-fd37b48e7a1b · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Adaptive Duration Model for Text Speech Alignment Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.167239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.167239Z digest=sha256:03f925176fa5b526fa516dbad40b9316a9dd2199dd61985a8a83f95c21f688cc

Observation d5b85757-2026-4848-bdb3-4bc26b481cdf · outbound

This paper cites Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis.

Adaptive Duration Model for Text Speech Alignment Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.232019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.232019Z digest=sha256:3246ce126b1a5bf3d5db4f63bfa77cc9013b0f8bc2dae2df0c1e722685288c8e

Observation 1dc12977-4893-42a9-9c18-67a74ec62837 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

Adaptive Duration Model for Text Speech Alignment Fastspeech: Fast, robust and controllable text to speech,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.701373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:08.320168Z digest=sha256:a139fe36e005219d6b9e9691f43d33381b353d8f8513647d8a654ecb2d5d3c49

Observation 03f6445e-4212-4a15-8f2a-1c798a8804c9 · outbound

This paper cites Fastpitch: Parallel text-to-speech with pitch prediction,.

Adaptive Duration Model for Text Speech Alignment Fastpitch: Parallel text-to-speech with pitch prediction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.524555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:08.378954Z digest=sha256:c64a2ef27527d85951404fd683f34b21d2f20643ab6f8899ff301e15f5da00eb

Observation 10183caa-b808-4807-92b6-767070d9fb8e · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Adaptive Duration Model for Text Speech Alignment FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.523005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.523005Z digest=sha256:3e5a5c8b057f19b8652dd6b7215cee9f5101cfa38b0ffaa3f195ca90ece5034c

Observation 849ba4da-4596-46cd-8456-028e7175abd4 · outbound

This paper cites Location-relative attention mechanisms for ro- bust long-form speech synthesis,.

Adaptive Duration Model for Text Speech Alignment Location-relative attention mechanisms for ro- bust long-form speech synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.359121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:08.586589Z digest=sha256:f11140f3179e13dc26ec2b937b350af79c5157f406bcf4fad562d78d72eafc5b

Observation 368df377-e60e-4603-b996-af6b64354af5 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Adaptive Duration Model for Text Speech Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.681690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.681690Z digest=sha256:7009631813a79fa7546b290a4f6aa25a96c35ea079dab5113e8b317165f61701

Observation 71ddde11-7728-46f0-97c8-631b877d30c0 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

Adaptive Duration Model for Text Speech Alignment NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.767437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.767437Z digest=sha256:eea1ab9582650eeff0402fc80f2ca3160823c6170ba058a89fffde43360761b0

Observation 2a3f2b53-4432-4704-bf13-5ce8b7d01f08 · outbound

This paper cites Non-autoregressive neural text-to-speech,.

Adaptive Duration Model for Text Speech Alignment Non-autoregressive neural text-to-speech,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.191926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:08.866084Z digest=sha256:b9e5c27987ef8768ecd50973fd3eec72ac548a4d5687cf824e4aac365e3cbdf4

Observation 21c31295-447d-46b1-8bc5-31f1d935c5ab · outbound

This paper cites Durian-e 2: Duration informed attention net- work with adaptive variational autoencoder and adver- sarial learning for expressive text-to-speech synthesis,.

Adaptive Duration Model for Text Speech Alignment Durian-e 2: Duration informed attention net- work with adaptive variational autoencoder and adver- sarial learning for expressive text-to-speech synthesis,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:11.038847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:08.959203Z digest=sha256:79ce9e946c88fb77db68f77c2637039f3b694ac69b5095a2faa0ac13403fb6e5

Observation 6cbc73b0-781d-4ae4-99b2-0c2dc1b537e5 · outbound

This paper cites Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features.

Adaptive Duration Model for Text Speech Alignment Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:34:09.863491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:09.044690Z digest=sha256:5de884c63408511ec4608c05ecdbb28ee7b3a60dc888f1851830b92ef0468d4f

Observation 10f337af-c4fa-42b6-bcc3-fe35ea5ccb08 · outbound

This paper cites Simple-tts: End-to-end text-to-speech synthesis with latent diffusion,.

Adaptive Duration Model for Text Speech Alignment Simple-tts: End-to-end text-to-speech synthesis with latent diffusion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:10.878139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:09.097184Z digest=sha256:077c615b06fec6b98b2e251ae234e12d1d1d32b539d6acda7885513fc878805c

Observation 7f2d6d4c-ffdd-4fe9-8d08-605b21eeb050 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Adaptive Duration Model for Text Speech Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.160142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.160142Z digest=sha256:a3b2ada2498f6a8b82cc9dfc031e0984fce4049efcf02554ccdae53fbffa6a4c

Observation 06f29d60-5a31-40d7-a324-22c2f553eea3 · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,.

Adaptive Duration Model for Text Speech Alignment Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:10.727750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:09.231739Z digest=sha256:2f8c3fdf75897d91bb7e955092e18355e4ea4f24f8dc64dd53b5bc71fe8cc46c

Observation cbfaa89b-f5e3-42c5-8734-c4c308da99e2 · outbound

This paper cites Portaspeech: Portable and high-quality generative text-to-speech,.

Adaptive Duration Model for Text Speech Alignment Portaspeech: Portable and high-quality generative text-to-speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:10.564583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:09.286028Z digest=sha256:533f4506a5376603ebe32f4b17d14b1eba9b5a75a2716606965e7239a02659d0

Observation a65be6c9-43c0-465b-a52b-8a8010999c00 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

Adaptive Duration Model for Text Speech Alignment V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:34:10.353445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:09.356034Z digest=sha256:25134aa5a43422f51af3298a8c21049b76e8bbdafece74099b6b010ee287fbac

Observation 752c8ba0-a959-4247-8d1e-c3737aee4cf3 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Adaptive Duration Model for Text Speech Alignment MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.403232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.403232Z digest=sha256:9c363ea6447b31957115a2e1bb3c463e49b902ad09aa109eab340c42b2a707db

Observation 9cb08e55-b39c-4f59-ab51-4d28892893d5 · outbound

This paper cites SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models.

Adaptive Duration Model for Text Speech Alignment SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.476364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.476364Z digest=sha256:f0e790b2f4609b3beed41fa159f6d219ec0c4700616b0dbda7d853d88bee4d0e

Observation 56b91bcd-6dbd-4bf0-b070-a152a73ffcc4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Adaptive Duration Model for Text Speech Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.568289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.568289Z digest=sha256:64f27ccf05504101b47bd1b586973876c438b1a6e4d80f773110ab712ac823d2

Observation f5a86135-f78e-4daf-8f2e-11ed8cd36458 · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

Adaptive Duration Model for Text Speech Alignment WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.630347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.630347Z digest=sha256:0af087f450de2f72ee1085befca3c24f277a8415d7115d2197ac1c546e5bdac3

Pith citing papers

Observation e822238e-0ba3-457f-bee2-0dc2c323cba7 · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment Adaptive Duration Model for Text Speech Alignment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:34:10.121918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:34:07.882293Z digest=sha256:af91b4de2bf44095e39ea65370b8e41595966f84a2944dc0eee20f21a9478d59