Pith. sign in

Paper Citation Record · LEDGER

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis

As of 10 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 0 inbound Pith citation observations for arXiv:2502.01084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01084 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:43:08.911696Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

72 of 72 outbound references displayed

  • verified exact3
  • verified fuzzy26
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 996989fd-0ad4-471a-8883-8a02d55370d1 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.900635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.642508Z digest=sha256:aae55d35bad769e0ac4c63b6af241d1c1dadca7c7b8fa205a3a7603592d9681a

Observation 2c20c57d-0727-4a1e-a2f9-1923e9fcec21 · outbound

This paper cites Better speech synthesis through scaling.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Better speech synthesis through scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.647563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.647563Z digest=sha256:17c9f6e6ed95874deeed943dddf206dbcf027474627075ac41f598969fd039d8

Observation da4d6059-d957-4a40-b646-ed79e10e0d91 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Audiolm: a language modeling approach to audio generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.888991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.652180Z digest=sha256:f147b2985392257638f360b03dbdf0353ea19ac1301cbf72904805c2aa5e2a88

Observation 4da0a028-7968-46a2-afc0-d9ba82130442 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis SoundStorm: Efficient Parallel Audio Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.656006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.656006Z digest=sha256:d6d0bb4ffe09d992f87fa239275b604c5ba6a5ff77d898e322f4a585d86c9dd5

Observation 7b55b257-0531-481b-a2a2-e35416a02f23 · outbound

This paper cites Language models are few-shot learners.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.660039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.660039Z digest=sha256:c09bf4e6f604a2874be79b51e5cf2ffe535782ccdaa1f5843a12faea5ef6ab8b

Observation 9a4a5a56-897e-4fa7-b9fc-06b5aa821aa5 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.663955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.663955Z digest=sha256:3f7ff1e5bec91cf96589602e00b1ada7e1b7f4deafd9801b43c5f5343c98b8a6

Observation 5919ee56-1eeb-4d1f-91bd-83f1f5086090 · outbound

This paper cites A vector quantized approach for text to speech synthesis on real-world spontaneous speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis A vector quantized approach for text to speech synthesis on real-world spontaneous speech

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.870716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.668091Z digest=sha256:8b5f4781b9aa41385ca6bf7883fd9d0f19fa2b6ec8cf7cc5fcb35ecc2ebe0438

Observation 10b39969-f666-4c5a-b2ea-c2d769da6ada · outbound

This paper cites ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:43:09.467295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.671551Z digest=sha256:440f1498de0c1bd8ad0c672954434a42d3f07eed18cddee6df8632a556ae83eb

Observation a98efccd-d824-4a27-a431-609124b0805e · outbound

This paper cites WavLM : Large-scale self-supervised pre-training for full stack speech processing.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis WavLM : Large-scale self-supervised pre-training for full stack speech processing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.858766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.675588Z digest=sha256:b1b837f72027b6c0bbc5588e02cfb6cccffd29a83ed8563de6f0286d752754f4

Observation 64796b72-25ba-44ed-891b-acb12ae1f5d4 · outbound

This paper cites Monotonic Chunkwise Attention.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Monotonic Chunkwise Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.679399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.679399Z digest=sha256:b87a469c72f017305263033def23d1abb0fcd8a89c75bf365fd91d22f293901e

Observation 683d8685-7551-4494-b098-9d356f18ddc0 · outbound

This paper cites Self-Supervised Speech Representations are More Phonetic than Semantic.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Self-Supervised Speech Representations are More Phonetic than Semantic

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.683147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.683147Z digest=sha256:bce3783f3d1fceb18dcf364f25f3e8bf17707cbea5f0896a7cab868ac20d5356

Observation 5139dc36-18d9-44d5-ba50-46c874ecb99d · outbound

This paper cites Unsupervised speech representation learning using wavenet autoencoders.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Unsupervised speech representation learning using wavenet autoencoders

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.846716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.687036Z digest=sha256:539441cb6bdf16a17406a1480affd533f1b69d22677c11557c2f8a4d9872632a

Observation c92f5921-bb4d-4d35-94f6-ed3b4f674af5 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis VoxCeleb2: Deep Speaker Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.690593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.690593Z digest=sha256:b008f9e76b58b49892f6b594e82530d3bc3acbc7b7faa0da495aea5bed394858

Observation ae839e87-2861-4dfb-95c4-6c5d3455ef4d · outbound

This paper cites Simple and controllable music generation.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Simple and controllable music generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.834722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.694045Z digest=sha256:ee6db3a751c8998502aa36c33a43f3ba2d3a8abfa5cc1c279051c66e53c5a358

Observation 42fdbbb7-daf1-409b-a88d-65dfee838a79 · outbound

This paper cites The Road Less Scheduled.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis The Road Less Scheduled

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.697680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.697680Z digest=sha256:f717ca6d9070718f1f9657eb7e6827d12cbebecf9a44d4fe8c2493e76298f754

Observation 78e54ebb-0f1c-4cfa-abdc-d34a57012c50 · outbound

This paper cites High Fidelity Neural Audio Compression.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis High Fidelity Neural Audio Compression

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.701299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.701299Z digest=sha256:150db93775513a75d5371cebd9bcfbd6c2dbc5b42668bb8388904992e518d5bf

Observation bed1d918-2f40-4e9c-bc3c-40d2c9cb9140 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.705119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.705119Z digest=sha256:2c333816ea1f0612ea55ecde791343933f82fa009656096259f26a6f55f84cb3

Observation 145b50cf-72dd-4570-b187-3558c0682659 · outbound

This paper cites Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.709109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.709109Z digest=sha256:9b9687a39ae58e5ab0809b41f554045f28b14486ee4399ec9d58a3d5a6380c01

Observation 80e6761d-857f-4f89-8d2c-491996350f99 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Taming transformers for high-resolution image synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.713005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.713005Z digest=sha256:50379a50eb936811351026ec1c698b6751ca42ac871aeafb58cc18d81a649bad

Observation 0676804b-3905-4acf-9af2-bc1561793e37 · outbound

This paper cites Stochastic Backpropagation through Mixture Density Distributions.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Stochastic Backpropagation through Mixture Density Distributions

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:43:09.362124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.716434Z digest=sha256:cc4615ec5af623d96ccffe72b3b3d269e867d2d91872d80f2128f64efef8cc4f

Observation 07aef043-1885-4302-95ac-421c5f748845 · outbound

This paper cites Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:43:09.344609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.720156Z digest=sha256:28e7be6e2fa2de027d2a3710b1efb117d8dc190daf14fda48a11941acf266a12

Observation 51194fae-1175-4cfc-9e8f-699f97b9b500 · outbound

This paper cites Visqol: an objective speech quality model.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Visqol: an objective speech quality model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.816607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.723947Z digest=sha256:ba40530d5c90bd0657d6185c32848d85e5ab9258e25c02562c53f55ead7dff0d

Observation 64c63cf4-faab-4277-91d3-8e1179154fa6 · outbound

This paper cites Reducing the dimensionality of data with neural networks.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Reducing the dimensionality of data with neural networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.727357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.727357Z digest=sha256:0799444eb51097da8bee2de6c6236c78652b96ff6ac197797b839d2b78be831b

Observation 044afde6-530a-473f-b1fe-04bb4432fcf4 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.799283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.730734Z digest=sha256:dc3c57a03c0204f35659818f2d2442da66a085268f55a1ff86c7a2e735f09171

Observation 15a21aac-fb50-487a-9384-360993f62109 · outbound

This paper cites Prodiff: Progressive fast diffusion model for high-quality text-to-speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Prodiff: Progressive fast diffusion model for high-quality text-to-speech

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.787597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.734163Z digest=sha256:ecdfa95f66e02641bb61b5347328fc344976e2e2f23bea8f215c94bf7e65a3e4

Observation 45b7547d-39b6-4a26-8dcc-0c14c4635439 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Categorical Reparameterization with Gumbel-Softmax

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.738045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.738045Z digest=sha256:6e17a70a972ea80e889eeb6e0990c6f9c291b8278a8fbc761ba436439ff6430b

Observation caa7e61c-5546-4f50-bd15-c1fdc9ccb94e · outbound

This paper cites Diff-TTS: A Denoising Diffusion Model for Text-to-Speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Diff-TTS: A Denoising Diffusion Model for Text-to-Speech

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.742096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.742096Z digest=sha256:a97306f3e870b196a7156f42cd22f6ba4cbf1f0a937f17ec0b201ea643a0f190

Observation 45da75a7-23fb-4c7e-9b49-c8ce4f8aeff1 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Libri-light: A benchmark for asr with limited or no supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.745908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.745908Z digest=sha256:b0c26b30a709bb12469029c69ba1220ac30308f61370f39e9869145f534fb562

Observation ce42680e-7ba0-415e-bf82-450e533ba805 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.769374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.749904Z digest=sha256:db2db45c61f3b34c4c8f4af4bb0881a7ff9a1ef9324e32674a8a8840eb46036a

Observation fffee71c-5db3-4f11-bb80-845d9355e870 · outbound

This paper cites Kingma and Max Welling.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Kingma and Max Welling

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.757721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.753289Z digest=sha256:9d548c3e11373b3951e0766327de4fd1ee72477030d84e63a571b809cabc3ba1

Observation 14cfdd7c-8655-4e06-9a7b-bdb201ac574c · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.745077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.756871Z digest=sha256:9a4e072d7379b9b1d5b448c9216fe7cebe1a05a347ff9e16e9dda31db24d1787

Observation 1dd0fe29-c507-42f8-b1eb-2d084abfdc8c · outbound

This paper cites High-fidelity audio compression with improved rvqgan.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis High-fidelity audio compression with improved rvqgan

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.733435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.760481Z digest=sha256:8ca41ff8e1943798416d94319278a6789ffbae73dbadb1f8acbaad06fed460b2

Observation 982da5a7-6077-45ab-9c46-f5f50fe8ca29 · outbound

This paper cites Robust training of vector quantized bottleneck models.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Robust training of vector quantized bottleneck models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.722906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.763769Z digest=sha256:0aa365619045f357f2c3d087bd9c5b363cb86c506e265b925e7d6aae429b8a98

Observation 996e2d5b-fdd1-48d5-88b8-1a4ee0220a00 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.768464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.768464Z digest=sha256:68aa965b886794af32ff971e625217aa483d55cbb4c173f621e8f8433dd7cf4c

Observation ceaace19-f2da-4f7a-97fc-c90eb85c1b64 · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.772217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.772217Z digest=sha256:a6e606c311c010d4183a1bcaa9d3597e5ac2ff1c69e2837f8b10170b8678379e

Observation 46734ece-16fc-4464-b22c-f7d03c37f3cb · outbound

This paper cites StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.775875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.775875Z digest=sha256:a53204341924eaaf40dc9ccf712d51c7c179a2b2d24ae7fb30d764916e39bf87

Observation a772db78-97fc-488d-8f3f-958bc6fe0134 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.779643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.779643Z digest=sha256:1f1e4056aec9a550f0c48583131b286c504a9b110519c8b1f39d70cee8b2826c

Observation 2f6bd274-9071-430b-be9e-a69f1ff53512 · outbound

This paper cites JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.783086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.783086Z digest=sha256:c4c0cb0cf421f0961bfe9f6eb2e479c6776a646825a161f8e301881e352f4a09

Observation 0c3d4016-2a18-45d4-946d-7359cefedb8c · outbound

This paper cites DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.786889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.786889Z digest=sha256:ce68d02518592827c779774a8d9240c241be11b3e7f5ba3a2f552097dcf90a05

Observation cd010cce-1125-4d87-ad1e-da6c7c2af2ce · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.790672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.790672Z digest=sha256:3a624cf4b149ee3df414a49c4176afdc0dff4ae0e58e45fd3e1c26a6d1e4ffaa

Observation 8215d38d-45e5-428f-88f1-99515e29df55 · outbound

This paper cites Adversarial Autoencoders.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Adversarial Autoencoders

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.794667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.794667Z digest=sha256:c8e5282e34e3e81f3dbf7e9ca7a9dec939bde229374a2ffeeba460691db75752

Observation 674f8a57-6c60-49c6-974d-a887ccba5baa · outbound

This paper cites Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.798686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.798686Z digest=sha256:5e9f65c51d46acc8c89f26340d57d1d9a13473323de6776b9c5004a5fd70c3a6

Observation 6975252e-218a-41d4-9a86-7e84fdf33352 · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Matcha-tts: A fast tts architecture with conditional flow matching

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.705074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.802541Z digest=sha256:0cdd454c7ba317a3fc81b13de057df7f3623d8eae368f448344194ef4faeff4c

Observation 88f18359-782a-4cf9-a878-cb14507d49a8 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Finite Scalar Quantization: VQ-VAE Made Simple

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.805847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.805847Z digest=sha256:c7d6fb3b75202139e66378e43d13f94cd3b5d4a2d3940116b6a624a71922f4d5

Observation bbac4ca4-2011-4d0d-8c2e-0f07b4bba779 · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis VoxCeleb: a large-scale speaker identification dataset

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.809872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.809872Z digest=sha256:52c9abca6994c8f35d1a8dfaa65309972fb0444899bfa9e84356d48030a96ec6

Observation c8705b49-25fa-4c1c-bb45-e288cfe5198a · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Librispeech: An ASR corpus based on public domain audio books

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.693535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.813617Z digest=sha256:d389f3f6fcf1d840790a49876ff917dd228dda156790c2209f29fffc823607ed

Observation dc3fa77f-3c20-498f-ba72-4a1c1b184875 · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.817148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.817148Z digest=sha256:d4b0dabacb5b1d888bd4d07bdbd1001209deef50ef2bb905933c426e64d857a0

Observation 24a6b9c0-c32b-4f2e-8ef5-16e955fb5f9a · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Grad-tts: A diffusion probabilistic model for text-to-speech

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.682228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.820769Z digest=sha256:c4ee77afeb29f496d8e3d417106cd1a76dfb061571eb90dd88c97b515ae3e549

Observation cfa424c9-5e9f-4027-a925-a560a4f0adb0 · outbound

This paper cites Autovc: Zero-shot voice style transfer with only autoencoder loss.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Autovc: Zero-shot voice style transfer with only autoencoder loss

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.670485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.824187Z digest=sha256:d5b3b3103bd565bd0220ebaf44e569f55a9bfd136a1df4456dd80857f3fa1336

Observation 10f2d27b-c147-4bc3-aa7d-83456ddd948e · outbound

This paper cites Language models are unsupervised multitask learners.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Language models are unsupervised multitask learners

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.827803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.827803Z digest=sha256:e6d719948c068278aacfadc620f566555084e269b2a0dc15ab84def4cd670267

Observation e5c149c6-979a-4f34-b0d0-349eee89c42f · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.831448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.831448Z digest=sha256:7636fe1bcb8253573d3024662597eae2146ec4c918e1d822ae7ac3dabe442616

Observation a5b6c15f-b284-4725-b56f-4229803b937d · outbound

This paper cites Online and linear-time attention by enforcing monotonic alignments.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Online and linear-time attention by enforcing monotonic alignments

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.649749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.835730Z digest=sha256:a32ff0c50e5d88d3f3b04e140aca7403d3b658b65d1ae5a18d645db31c0b8888

Observation af00bcea-17c3-4890-a58c-a23e17cf1b72 · outbound

This paper cites Zero-shot text-to-image generation.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Zero-shot text-to-image generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.839506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.839506Z digest=sha256:2ea4fcbb4d834bee4c994e1e3981fdeeae6f2627c7b756888cbf8e8911b5d90b

Observation 8554702d-e5e9-48f3-ba43-23f1330914cd · outbound

This paper cites Multi-task self-supervised learning for robust speech recognition.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Multi-task self-supervised learning for robust speech recognition

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.629301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.843091Z digest=sha256:317d4c1de81d90beb873f2c6adb81d9aee3ca92ce879b6f62c75c6c092cc4647

Observation 95547e88-ccb3-45cb-a924-52d7e9dc8f73 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis SpeechBrain: A General-Purpose Speech Toolkit

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.846652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.846652Z digest=sha256:cb92323001c6fa01ae915fdfa632aff8b7a4c6bcde07896a84df41aaa6aca704

Observation 1efd17cb-4407-4dad-9e1f-c1b4422746be · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Fastspeech: Fast, robust and controllable text to speech

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.850283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.850283Z digest=sha256:392bb11897a6ff0293e059ef9877f604a3c7a4251be6244279e98d4619cc5c7e

Observation 30595c40-9843-41ae-99a8-e0c97676db0f · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.853890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.853890Z digest=sha256:e3847eaad56d9674e73c2cbf5723db50e31229fd3ae0c116cdef9d78847fadae

Observation 527985ac-02ac-42d5-b1a4-e2b334c35561 · outbound

This paper cites PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.857847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.857847Z digest=sha256:5efd66f209241aa0427ed99c81bbdd1c84a442d972354642fc90408eabea7bbf

Observation 9e23dc9b-2bd6-467a-b18c-9d3f3979ccca · outbound

This paper cites Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.609529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.861735Z digest=sha256:8571d14938589eb07aad602d85a0d8ad4eeae11860f039ff9edef313a963ce69

Observation 1e86ae95-b1a8-4634-8fa5-ed966586445e · outbound

This paper cites Vae with a vampprior.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Vae with a vampprior

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.597826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.865243Z digest=sha256:d74dd5924892ae61ce8fd2e6bb4b1da32c13fe73763167519cd1da0baabf86ca

Observation 71b940c0-1f75-4b68-9f1e-4b0f1d62fbba · outbound

This paper cites Givt: Generative infinite-vocabulary transformers.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Givt: Generative infinite-vocabulary transformers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.586523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.868954Z digest=sha256:d1b14707f94c2b4fa5a15591f9ac1a4983bc3103285a95a0c8dd96d8d75a2359

Observation 15dee1eb-a0c1-4600-9ac4-f2164726861f · outbound

This paper cites Neural discrete representation learning.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Neural discrete representation learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.574677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.872477Z digest=sha256:7d31a779e2526126b13bffeabcbb1c778a914cbe9526a85401c320794286dc75

Observation 854a2750-16f6-4ce2-99a7-b7fef754ed88 · outbound

This paper cites Attention is all you need.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Attention is all you need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.875975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.875975Z digest=sha256:f3004cecd5c6444086741001d76cd5deb52ad33e84cc25e1283c00b48a33a9ce

Observation 7b5d92ce-8e0a-4d35-98be-214931992f8a · outbound

This paper cites Extracting and composing robust features with denoising autoencoders.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Extracting and composing robust features with denoising autoencoders

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.879511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.879511Z digest=sha256:21c6e55ceb74e3f795a690d9e4021e27a1dfacf20a2db69f748a8769d14027ab

Observation f8a9b278-1a99-4327-897c-8e491fe085be · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.882979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.882979Z digest=sha256:b1fa4660e1e1cca035cf1628fafc1ce16abff9a06cf021cb596d2853fbbc6b54

Observation 5f3c0e47-517f-4968-9acf-ece206c86d17 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Wespeaker: A research and production oriented speaker embedding learning toolkit

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.548600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.886878Z digest=sha256:36693004fbac31eb40ee1095c1d657fd8f3b66f88e438f6acb0858cfe4181b98

Observation 399ba5e1-aac2-4875-8197-67492879902a · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Tacotron: Towards End-to-End Speech Synthesis

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.890594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.890594Z digest=sha256:38b8ae0f589dd1299582e3bffdd8f7efaf15013d105f6ca0ec42c93aefacb739

Observation 68e823fd-aa63-4bdf-baef-9f1b03201468 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Soundstream: An end-to-end neural audio codec

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T16:43:09.536680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T16:43:08.894452Z digest=sha256:eb5419e6792d23cf15eb2393721ca286139cd120af9cf073517dec83593f41b8

Observation 56c7d2ab-042b-413e-8564-bb6deebe436a · outbound

This paper cites write newline.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis write newline

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.897920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.897920Z digest=sha256:cf8725a89e415e511b9c5cd762a31c78c61a7710d6e5e483a52c326304538292

Observation 79083430-95b0-4b2e-9170-49f5e7593ec0 · outbound

This paper cites @esa (Ref.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis @esa (Ref

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.903102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.903102Z digest=sha256:1891f2031dd4623e5b85feaa62bfa460303287749714ae97fb9f0b65c497468e

Observation a613cc14-712e-429d-9e52-60008a2d4564 · outbound

This paper cites an unresolved cited work.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Unresolved cited work

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.907509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.907509Z digest=sha256:6af7dd16c140d90a1cedc0176df8b6da5a74d76353a46832bd2715957a5abc73

Observation 0cf9e3f8-fc38-42de-ac12-fc96e9c01c22 · outbound

This paper cites 1.0" encoding=.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis 1.0" encoding=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.911696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.911696Z digest=sha256:0d6ccba7afd20f60f54bc1ba44ebcc89ec5a77b788cbce26be9453e882d0cd82

Pith citing papers

No inbound Pith citation observations are available.