Pith. sign in

Paper Citation Record · LEDGER

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration

As of 16 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2412.08112.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08112 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:16:26.973724Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f560fb69-7657-4e82-ab6e-b306216a25a9 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration Fastspeech: Fast, robust and controllable text to speech,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.917953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.917953Z digest=sha256:e4f843de5e09535beb2c9412f8909f8aeae9ba83563df7f7628b221d33a97168

Observation 2ec1ae0d-b63b-44fb-a6c7-f2372c92b755 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.922799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.922799Z digest=sha256:8a81b0f03c0e73ce538fdbd0fc1afb63647df2706ce3596d18061850fbeca409

Observation 7275fe30-9ac6-4f4e-8d1d-bd7f647be84f · outbound

This paper cites Stylespeech: Parameter-efficient fine tuning for pre-trained controllable text-to- speech,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration Stylespeech: Parameter-efficient fine tuning for pre-trained controllable text-to- speech,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.927864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.927864Z digest=sha256:02e3eaa721366cbad24737070e4ad500ce64442bf93c922dcb13c2783490ae41

Observation 900cc78f-fec9-4a2c-80eb-0e049fa537bc · outbound

This paper cites Montreal forced aligner: Trainable text- speech alignment using kaldi.,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration Montreal forced aligner: Trainable text- speech alignment using kaldi.,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:16:27.137777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T18:16:26.932348Z digest=sha256:e75fdf901043b7c5a54df95fefab2f1b271f2cf730f9c0b0c1eb13143c124e8c

Observation f98cb569-ce62-4360-a6be-238a161353b0 · outbound

This paper cites The kaldi speech recognition toolkit,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration The kaldi speech recognition toolkit,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:16:27.121236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T18:16:26.936995Z digest=sha256:b30763161b57efb4094f5a37bab7f28a70077e3f8b67fa9b1055feeb69c2f2de

Observation 22e600d4-296e-4cec-bfe6-afe035b083f1 · outbound

This paper cites An introduction to hidden markov models,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration An introduction to hidden markov models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:16:27.105994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T18:16:26.941294Z digest=sha256:bf2508c0c42b361c54f311c24a9532caf22c0b09c2cfca276dfbeab78f6288d6

Observation 0e3f3873-eb07-40ac-a0b9-703aaf58d2ac · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:16:27.091344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T18:16:26.946079Z digest=sha256:0aa7b64e4d8bd2e7426a3683ee032956f79271065cab8c28ae5a280be5ddc31a

Observation b0d1201b-80cf-4fc9-8afc-5498299bf0bb · outbound

This paper cites RAVE: A variational autoencoder for fast and high-quality neural audio synthesis.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration RAVE: A variational autoencoder for fast and high-quality neural audio synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.951656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.951656Z digest=sha256:5d55d2a1d9c2030042c3332d8d8607ddc46228baede204958d392aee70ef8d8a

Observation 73789886-b9b6-4ba0-bd39-246065e0b6db · outbound

This paper cites Chinese mandarin female corpus,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration Chinese mandarin female corpus,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.956497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.956497Z digest=sha256:6adeaa0dc2151b9532e648ce185f51f98c4274bdfef0bb7f3913dec1e65eb402

Observation ff00fc4c-716c-4322-a7ac-86e254bc9e3f · outbound

This paper cites Attention is all you need,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration Attention is all you need,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.960896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.960896Z digest=sha256:d0024092901599bc4f428c95d52b8f7736a4d0536240ba069838aff4a8c4dafb

Observation 31ebd09f-9c89-4825-b76d-ea16e2de42b4 · outbound

This paper cites Mel-cepstral distance measure for objective speech quality assessment,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration Mel-cepstral distance measure for objective speech quality assessment,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.965277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.965277Z digest=sha256:ad004eeaa90036d6a1497eb7c72077f6f7a9450667f3508ac053e3c4d4881445

Observation 1d3fd358-e60e-4d17-ba07-f4cd4b06bca8 · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.969170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.969170Z digest=sha256:3d634235b6d92f4e844a294dd196b2b7dd7bf34cdf34a2d7b30b19634f143d62

Observation 0386afa9-ea54-4870-bb66-e2e0c4e35193 · outbound

This paper cites Robust speech recognition via large- scale weak supervision,.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration Robust speech recognition via large- scale weak supervision,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.973724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.973724Z digest=sha256:47d8aa2c3f0b6d1da666d26b835db46f46edd1ea8a7ddb458bfcfd0a989a5abb

Pith citing papers

No inbound Pith citation observations are available.