Pith. sign in

Paper Citation Record · LEDGER

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis

As of 18 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 0 inbound Pith citation observations for arXiv:2502.01084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01084 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:43:08.911696Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

72 of 72 outbound references displayed

  • verified exact3
  • verified fuzzy26
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 996989fd-0ad4-471a-8883-8a02d55370d1 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.900635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.642508Z digest=sha256:45861184fc8e45a849d05f4c61288362e7cf6f4c8992956031a37a7ac1941ce8

Observation 2c20c57d-0727-4a1e-a2f9-1923e9fcec21 · outbound

This paper cites Better speech synthesis through scaling.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Better speech synthesis through scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.647563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.647563Z digest=sha256:29d8cc486f001321c458bd367a8244066be1acaa4b71d53b84a43f0955ae0a49

Observation da4d6059-d957-4a40-b646-ed79e10e0d91 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Audiolm: a language modeling approach to audio generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.888991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.652180Z digest=sha256:aa359363a2cc71a183cf255295328a7d66ddae09fc65ab207e668a347b9c1982

Observation 4da0a028-7968-46a2-afc0-d9ba82130442 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis SoundStorm: Efficient Parallel Audio Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.656006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.656006Z digest=sha256:523b8891b6d9b67f40febaa03e61c8574f0ed88f9649e5bf47f6d13ac0111efa

Observation 7b55b257-0531-481b-a2a2-e35416a02f23 · outbound

This paper cites Language models are few-shot learners.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.660039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.660039Z digest=sha256:ce2b27797b97c153ede3901e741fe601af7a7bbd762762d284e426e78da3837d

Observation 9a4a5a56-897e-4fa7-b9fc-06b5aa821aa5 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.663955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.663955Z digest=sha256:61ed1f0e9242a01d9b82bbd1794942e8fb50a44f3a33e0e9757617aa1b3ef70c

Observation 5919ee56-1eeb-4d1f-91bd-83f1f5086090 · outbound

This paper cites A vector quantized approach for text to speech synthesis on real-world spontaneous speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis A vector quantized approach for text to speech synthesis on real-world spontaneous speech

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.870716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.668091Z digest=sha256:2e7207339ede8adaeddddf06ca2d738b8b858bb5d3d773815e2c4262f138ec69

Observation 10b39969-f666-4c5a-b2ea-c2d769da6ada · outbound

This paper cites ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:43:09.467295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.671551Z digest=sha256:95d538e37cea672ce8886aa01bcc5812b6ef8e8743056b3eaabcaf989abd186e

Observation a98efccd-d824-4a27-a431-609124b0805e · outbound

This paper cites WavLM : Large-scale self-supervised pre-training for full stack speech processing.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis WavLM : Large-scale self-supervised pre-training for full stack speech processing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.858766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.675588Z digest=sha256:de148755f5c0097c584369a29bc4231bfb17e8da51e0edfbf01f21417a706095

Observation 64796b72-25ba-44ed-891b-acb12ae1f5d4 · outbound

This paper cites Monotonic Chunkwise Attention.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Monotonic Chunkwise Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.679399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.679399Z digest=sha256:9fd5a6e92bfa5319ac4464fe2bdcc8e5ffb7ad2e66adf5b05d23a2216b61a527

Observation 683d8685-7551-4494-b098-9d356f18ddc0 · outbound

This paper cites Self-Supervised Speech Representations are More Phonetic than Semantic.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Self-Supervised Speech Representations are More Phonetic than Semantic

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.683147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.683147Z digest=sha256:b7284840e63733a0f336264cabfed8a45c03fd6a44c2d46ac869eaf3bf2919e7

Observation 5139dc36-18d9-44d5-ba50-46c874ecb99d · outbound

This paper cites Unsupervised speech representation learning using wavenet autoencoders.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Unsupervised speech representation learning using wavenet autoencoders

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.846716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.687036Z digest=sha256:f63d2a27dd2350c600299a8936f4cab275e9ca229aaef712cc96ca9015550b85

Observation c92f5921-bb4d-4d35-94f6-ed3b4f674af5 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis VoxCeleb2: Deep Speaker Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.690593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.690593Z digest=sha256:cb1c6e51ecde278568feb38f7ddfa7965fac0bf59523df52d12ec12dd084e898

Observation ae839e87-2861-4dfb-95c4-6c5d3455ef4d · outbound

This paper cites Simple and controllable music generation.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Simple and controllable music generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.834722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.694045Z digest=sha256:1127e6f521fa152704c5bcb8f3b2978799a94a722693ed8976993b8afb0b31e1

Observation 42fdbbb7-daf1-409b-a88d-65dfee838a79 · outbound

This paper cites The Road Less Scheduled.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis The Road Less Scheduled

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.697680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.697680Z digest=sha256:24a97e095be5f83c5174722c173d4f21934af5b569bae789e7b7b096e2791b15

Observation 78e54ebb-0f1c-4cfa-abdc-d34a57012c50 · outbound

This paper cites High Fidelity Neural Audio Compression.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis High Fidelity Neural Audio Compression

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.701299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.701299Z digest=sha256:2725ca8511b120466a1c0628e438d35acd1e01970e0655bde784e86b2dafc1dc

Observation bed1d918-2f40-4e9c-bc3c-40d2c9cb9140 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.705119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.705119Z digest=sha256:6a6c0bcb030d6abb1ec5145de2849a7b7fa719b8c7e8d88fe8d32ddf3ba6afc9

Observation 145b50cf-72dd-4570-b187-3558c0682659 · outbound

This paper cites Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.709109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.709109Z digest=sha256:8eacb54b98d72ad01ece3f43d57fc827e78ac4a40f266cefb7630bfec71b43b7

Observation 80e6761d-857f-4f89-8d2c-491996350f99 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Taming transformers for high-resolution image synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.713005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.713005Z digest=sha256:5e29b802e3d8f03d278ed9c5ec4fd739ef292b474b3a36b18282129e2d3a3bfc

Observation 0676804b-3905-4acf-9af2-bc1561793e37 · outbound

This paper cites Stochastic Backpropagation through Mixture Density Distributions.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Stochastic Backpropagation through Mixture Density Distributions

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:43:09.362124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.716434Z digest=sha256:53d72899395115d8eb77ad6a8fa136c6f33b6453a863fee4d1d2abac6a3bb4d6

Observation 07aef043-1885-4302-95ac-421c5f748845 · outbound

This paper cites Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:43:09.344609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.720156Z digest=sha256:8c96fb29d0cbe13f740f1b5a446d4857b22e62b25e8778d83453de9836da9f52

Observation 51194fae-1175-4cfc-9e8f-699f97b9b500 · outbound

This paper cites Visqol: an objective speech quality model.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Visqol: an objective speech quality model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.816607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.723947Z digest=sha256:2a8bc8ad2ecce8a6e615d1dd8f700556315805dbbde8ab613c5fc88071d48bb4

Observation 64c63cf4-faab-4277-91d3-8e1179154fa6 · outbound

This paper cites Reducing the dimensionality of data with neural networks.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Reducing the dimensionality of data with neural networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.727357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.727357Z digest=sha256:509c0432f9e69e1df81d1e70fc8df4a7534fb75bbe065cb058f37ec848a8128c

Observation 044afde6-530a-473f-b1fe-04bb4432fcf4 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.799283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.730734Z digest=sha256:6d3ba72cb231e608688c5d3b27c0c0122d02bd0136c90c26f6f9badef84ea6d8

Observation 15a21aac-fb50-487a-9384-360993f62109 · outbound

This paper cites Prodiff: Progressive fast diffusion model for high-quality text-to-speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Prodiff: Progressive fast diffusion model for high-quality text-to-speech

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.787597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.734163Z digest=sha256:6e6a8e53fd7186dceb02599e440d132c21bb984a49c6cbe4b2e835f7a5758a0e

Observation 45b7547d-39b6-4a26-8dcc-0c14c4635439 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Categorical Reparameterization with Gumbel-Softmax

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.738045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.738045Z digest=sha256:d931705056668d6c8f3c6033f5d965ad9a570f03d93dc7add862b9735c8a33c4

Observation caa7e61c-5546-4f50-bd15-c1fdc9ccb94e · outbound

This paper cites Diff-TTS: A Denoising Diffusion Model for Text-to-Speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.742096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.742096Z digest=sha256:c367aee1ce5ae84f44f94cf7454048ab7b49aaf0fbc1a380c61929c245e82ace

Observation 45da75a7-23fb-4c7e-9b49-c8ce4f8aeff1 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Libri-light: A benchmark for asr with limited or no supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.745908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.745908Z digest=sha256:56960ee706a8257aa875b2bea2fbed5855d7984b78444cd9929cf84ba1974acd

Observation ce42680e-7ba0-415e-bf82-450e533ba805 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.769374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.749904Z digest=sha256:7e65c7ff614dd118dd1859b0f3055eb1f25ad98ec46c114a24882c36c3b65a4f

Observation fffee71c-5db3-4f11-bb80-845d9355e870 · outbound

This paper cites Kingma and Max Welling.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Kingma and Max Welling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.757721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.753289Z digest=sha256:08ed4499e84ab44ff1f09d2435a83531d4a92758614808767de10db41d7555c8

Observation 14cfdd7c-8655-4e06-9a7b-bdb201ac574c · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.745077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.756871Z digest=sha256:04aa0d13f4582015a748b61ab13e5a1e9a80aa4632cdc854b828ca77cc91609c

Observation 1dd0fe29-c507-42f8-b1eb-2d084abfdc8c · outbound

This paper cites High-fidelity audio compression with improved rvqgan.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis High-fidelity audio compression with improved rvqgan

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.733435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.760481Z digest=sha256:41f152ac03ed01ea3c41e927d197bf17a31710e045bd56d310a996a6a93f1f7c

Observation 982da5a7-6077-45ab-9c46-f5f50fe8ca29 · outbound

This paper cites Robust training of vector quantized bottleneck models.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Robust training of vector quantized bottleneck models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.722906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.763769Z digest=sha256:1dd186a6372478b0ea437fb026b3de586cd5087fe0a277f8eba2c6bb71abf968

Observation 996e2d5b-fdd1-48d5-88b8-1a4ee0220a00 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.768464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.768464Z digest=sha256:558f164a76eb11e1901b100356d5256d0aeb8d7d25581cd073ca64f4ddd813eb

Observation ceaace19-f2da-4f7a-97fc-c90eb85c1b64 · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.772217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.772217Z digest=sha256:1136cfa9acb8088f423a162a83e0fb184060c7da62c1ca079bf0f3ae4daa5ab7

Observation 46734ece-16fc-4464-b22c-f7d03c37f3cb · outbound

This paper cites StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.775875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.775875Z digest=sha256:302c3105acef4bfe176c38e1a3c7b5f343b353bc22bfb26c7cc8681d82bf4eb0

Observation a772db78-97fc-488d-8f3f-958bc6fe0134 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.779643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.779643Z digest=sha256:87d285bd731fc6398f3d447ff12e41d0db0b017c5fc772daea7d9137c467c684

Observation 2f6bd274-9071-430b-be9e-a69f1ff53512 · outbound

This paper cites JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.783086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.783086Z digest=sha256:f4b807218fdacd04c5d9522e6cd6883883918b93843281e2fab6597c15459a06

Observation 0c3d4016-2a18-45d4-946d-7359cefedb8c · outbound

This paper cites DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.786889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.786889Z digest=sha256:276c09ab8aecca0e18d7387ef51058e71a31a92368778e367e37b1f8c799c86e

Observation cd010cce-1125-4d87-ad1e-da6c7c2af2ce · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.790672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.790672Z digest=sha256:152928637b95dd5819ac4edb59a79b6f18f9452ddc889b629ca65d58fd6e2066

Observation 8215d38d-45e5-428f-88f1-99515e29df55 · outbound

This paper cites Adversarial Autoencoders.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Adversarial Autoencoders

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.794667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.794667Z digest=sha256:7c9ae16c811cfa99b59ecbebf6f05e26f46651b29f0382d655e5ec6f9a482f11

Observation 674f8a57-6c60-49c6-974d-a887ccba5baa · outbound

This paper cites Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.798686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.798686Z digest=sha256:980bc09bbfde7ed0f3ffa5eb76a69af9a1f334810b39c597ba04b52c82a5ee03

Observation 6975252e-218a-41d4-9a86-7e84fdf33352 · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Matcha-tts: A fast tts architecture with conditional flow matching

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.705074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.802541Z digest=sha256:a506f0ab34fd5979be564e851ef3b6f809d9e6d9280c5a966b647029e7393c64

Observation 88f18359-782a-4cf9-a878-cb14507d49a8 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Finite Scalar Quantization: VQ-VAE Made Simple

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.805847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.805847Z digest=sha256:90742e76bc51fd27cbfad0f0cab50a6868dd12bf9ca5237bfddd1e21e8227658

Observation bbac4ca4-2011-4d0d-8c2e-0f07b4bba779 · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis VoxCeleb: a large-scale speaker identification dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.809872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.809872Z digest=sha256:3643feb797155cb0cdf4501c5eb873d3217ed9bb7d121efde90316d58caed40a

Observation c8705b49-25fa-4c1c-bb45-e288cfe5198a · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Librispeech: An ASR corpus based on public domain audio books

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.693535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.813617Z digest=sha256:fcc061aa4be0f6c5ad459d0eb1e201be5ed6a8eb486b97ae43bfe47eb1d6ca9f

Observation dc3fa77f-3c20-498f-ba72-4a1c1b184875 · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.817148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.817148Z digest=sha256:8848c480db574a27be9e913711394abafd987bcfce1f18da5c0813ac1cf98037

Observation 24a6b9c0-c32b-4f2e-8ef5-16e955fb5f9a · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Grad-tts: A diffusion probabilistic model for text-to-speech

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.682228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.820769Z digest=sha256:0872fe2e742f923ed7cd7a61c236a39bab49dc687dff21c180ce729b980e60ac

Observation cfa424c9-5e9f-4027-a925-a560a4f0adb0 · outbound

This paper cites Autovc: Zero-shot voice style transfer with only autoencoder loss.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Autovc: Zero-shot voice style transfer with only autoencoder loss

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.670485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.824187Z digest=sha256:ecf8deca09da956bff8a37316f4e67167fe09477be769601ffaa0bd018e0c9a3

Observation 10f2d27b-c147-4bc3-aa7d-83456ddd948e · outbound

This paper cites Language models are unsupervised multitask learners.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Language models are unsupervised multitask learners

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.827803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.827803Z digest=sha256:b23f3b48780b9039c65e8032abf10a3d710c2ce32b3ea43bec248ed823e654f4

Observation e5c149c6-979a-4f34-b0d0-349eee89c42f · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.831448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.831448Z digest=sha256:df52bb7a55758f50816316cfd9fb13e8f6aa7aa99ce98bc29b90499afccc9232

Observation a5b6c15f-b284-4725-b56f-4229803b937d · outbound

This paper cites Online and linear-time attention by enforcing monotonic alignments.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Online and linear-time attention by enforcing monotonic alignments

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.649749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.835730Z digest=sha256:cc4ff8ee5c4e1f7ed10719c5514b8478145296cb5b35a7f8511dd4741c592ee8

Observation af00bcea-17c3-4890-a58c-a23e17cf1b72 · outbound

This paper cites Zero-shot text-to-image generation.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Zero-shot text-to-image generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.839506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.839506Z digest=sha256:ad4b317d849549c5c8311849e6a8e3ef92c432ab509c6ee3dcc8da5d3f02f3db

Observation 8554702d-e5e9-48f3-ba43-23f1330914cd · outbound

This paper cites Multi-task self-supervised learning for robust speech recognition.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Multi-task self-supervised learning for robust speech recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.629301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.843091Z digest=sha256:124624c9c64cd91d3adc4ece348a227792169de9b1551f0c2cf5c36a294624a3

Observation 95547e88-ccb3-45cb-a924-52d7e9dc8f73 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis SpeechBrain: A General-Purpose Speech Toolkit

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.846652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.846652Z digest=sha256:40ee48116faa0bfaee2ce550de2620301042914f841f14ea48e33d88aa6721fd

Observation 1efd17cb-4407-4dad-9e1f-c1b4422746be · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Fastspeech: Fast, robust and controllable text to speech

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.850283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.850283Z digest=sha256:ef6d885b40aaa8eb849812a8993bcab5d970b52640e41b8b987dc7853ce9c433

Observation 30595c40-9843-41ae-99a8-e0c97676db0f · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.853890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.853890Z digest=sha256:de40d6f17c2d90976b289118e8101e175951b45c73f8dde1af091726b5470eaa

Observation 527985ac-02ac-42d5-b1a4-e2b334c35561 · outbound

This paper cites PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.857847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.857847Z digest=sha256:27d7146f93384532d9a8d231b756922ab85e2c725d6bb04a3845a9105831603b

Observation 9e23dc9b-2bd6-467a-b18c-9d3f3979ccca · outbound

This paper cites Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.609529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.861735Z digest=sha256:4e2a84f5148c1ad8f03de39cd5654ae747d5e7b766e4f4365b64ccb5ca4aac5e

Observation 1e86ae95-b1a8-4634-8fa5-ed966586445e · outbound

This paper cites Vae with a vampprior.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Vae with a vampprior

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.597826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.865243Z digest=sha256:21b14e653885b07d6d95a6b631316beb13ede3ec3738b76776b7b1279c64d0ca

Observation 71b940c0-1f75-4b68-9f1e-4b0f1d62fbba · outbound

This paper cites Givt: Generative infinite-vocabulary transformers.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Givt: Generative infinite-vocabulary transformers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.586523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.868954Z digest=sha256:d29bbefd6ba7f08599161ef822fcb3ab95e155229a26be0455dd85927c6347cd

Observation 15dee1eb-a0c1-4600-9ac4-f2164726861f · outbound

This paper cites Neural discrete representation learning.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Neural discrete representation learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.574677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.872477Z digest=sha256:f014e58cf27ba366b866bbdac8a22fc5b44a37d1e5125b74f7c3c7fd0efdff4d

Observation 854a2750-16f6-4ce2-99a7-b7fef754ed88 · outbound

This paper cites Attention is all you need.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Attention is all you need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.875975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.875975Z digest=sha256:d24681a7902634203ce4fb7d40d89ca2c9dc57cdedb5b3e44cdb207b27656395

Observation 7b5d92ce-8e0a-4d35-98be-214931992f8a · outbound

This paper cites Extracting and composing robust features with denoising autoencoders.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Extracting and composing robust features with denoising autoencoders

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.879511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.879511Z digest=sha256:a547fd0db154442a53ce1efc850b94a9df416114a5435d765f742912a35cbb32

Observation f8a9b278-1a99-4327-897c-8e491fe085be · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.882979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.882979Z digest=sha256:7827dd0b409a2263512ee5c5d1273925955a030658324111fbfeee47604fa0e5

Observation 5f3c0e47-517f-4968-9acf-ece206c86d17 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Wespeaker: A research and production oriented speaker embedding learning toolkit

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.548600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.886878Z digest=sha256:a628329bb41cc217a1f239944812fffb228e9a6fb735f5301e59ac76407529f2

Observation 399ba5e1-aac2-4875-8197-67492879902a · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Tacotron: Towards End-to-End Speech Synthesis

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.890594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.890594Z digest=sha256:93def18df984ad1231ad18b0e2692f401bc9a2b704cb230630d58beff0835718

Observation 68e823fd-aa63-4bdf-baef-9f1b03201468 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Soundstream: An end-to-end neural audio codec

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.536680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.894452Z digest=sha256:b1d63cdde3018f67a847df4ef7a0e863d3597f277fb37ef02e1cd9d436b2bfc8

Observation 56c7d2ab-042b-413e-8564-bb6deebe436a · outbound

This paper cites write newline.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis write newline

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.897920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.897920Z digest=sha256:50bcc9d81771f1ab83b0cdf799918b001a12bd65333286f2c5775954111abd08

Observation 79083430-95b0-4b2e-9170-49f5e7593ec0 · outbound

This paper cites @esa (Ref.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis @esa (Ref

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.903102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.903102Z digest=sha256:81c7db7b9d23a8173bdd41752f74c694f74929b1f19c53b5639dd29c2abd1ee2

Observation a613cc14-712e-429d-9e52-60008a2d4564 · outbound

This paper cites an unresolved cited work.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.907509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.907509Z digest=sha256:ddce4e3bd4fc7be0c04a30d9e97e3833891f2d669acb6665d587a8db2de04c86

Observation 0cf9e3f8-fc38-42de-ac12-fc96e9c01c22 · outbound

This paper cites 1.0" encoding=.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis 1.0" encoding=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.911696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.911696Z digest=sha256:d8f3b2106a10f4283579e9fbf760d3bcb3c52d7fe0ab4cadbac578e98fcf2ade

Pith citing papers

No inbound Pith citation observations are available.