Pith. sign in

Paper Citation Record · LEDGER

Taming Audio VAEs via Target-KL Regularization

As of 4 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2605.17085.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17085 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T14:53:15.718359Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T14:53:15.718359Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-20T14:53:23.198073Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact17
  • verified fuzzy25
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2801ff57-1e74-4225-b0f4-afd3a9f4636a · outbound

This paper cites an unresolved cited work.

Taming Audio VAEs via Target-KL Regularization Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-20T14:53:23.496571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:0d60058aae6894ad1cb9c977770092f2ed13ff604e7c9dfe09e394319de86f18

Observation 771bc962-b180-478b-a295-4c2b60751b29 · outbound

This paper cites an unresolved cited work.

Taming Audio VAEs via Target-KL Regularization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-20T14:53:23.492603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:f577e4b0821091c090d48594b23e4b0d991aba700cb090d12404dfb6f307d45c

Observation 37ab0a92-c010-47af-95f0-b251a54e9137 · outbound

This paper cites an unresolved cited work.

Taming Audio VAEs via Target-KL Regularization Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-20T14:53:23.494633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:f6d94e25248438238b096524eb29828d8e3042bd615d69190bb1893368bc0558

Observation 87d74ee6-96e8-40ac-bede-460fd3eb4ead · outbound

This paper cites an unresolved cited work.

Taming Audio VAEs via Target-KL Regularization Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-20T14:53:23.490382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:c0feb54a8a8934fdefe7603f0a98b422105380114f505eea737e4ff248ce476d

Observation 08e5bc35-139f-4b17-88df-1dfb9801d680 · outbound

This paper cites Taming Audio VAEs via Target-KL Regularization.

Taming Audio VAEs via Target-KL Regularization Taming Audio VAEs via Target-KL Regularization

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:53:23.199683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:f83095df64aec0ccd59443e18c16b4bd858909031bb74ebb7f7bc373a6ea110d

Observation 684f0f92-7476-44c4-8ded-005f1e3a7e21 · outbound

This paper cites Model architecture Our model is built on the same framework of neural audio codec models, except we replace the quantization bottleneck with a gaus- sian regularization.

Taming Audio VAEs via Target-KL Regularization Model architecture Our model is built on the same framework of neural audio codec models, except we replace the quantization bottleneck with a gaus- sian regularization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.528340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:a4f975de9e13c93f11e8151df0ec79280ca21534d34f04d77a6530ca135ae4c6

Observation 2611bae7-9537-42a6-a7ee-8479da6ff6ae · outbound

This paper cites an unresolved cited work.

Taming Audio VAEs via Target-KL Regularization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-20T14:53:23.551137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:4dbe0899a8856709f5875a58910196e720e7aacdbb0630ce73e49c342f048948

Observation 91155a2c-46dd-4c89-8208-5490d10a79d8 · outbound

This paper cites This allows for direct comparison to discrete neural audio codecs and enables systematic study of the rate-distortion trade-off for continuous audio compres- sion models.

Taming Audio VAEs via Target-KL Regularization This allows for direct comparison to discrete neural audio codecs and enables systematic study of the rate-distortion trade-off for continuous audio compres- sion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.498854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:7722f80c46f104c3d9bb6eb06fcdc5732b71943a8eb1a913b2c851f987d36e21

Observation f3f913b4-64b3-4bad-affe-0c2bd62ec039 · outbound

This paper cites Neural discrete representation learning.

Taming Audio VAEs via Target-KL Regularization Neural discrete representation learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.545442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:40821ddd2a2fbe54e49d0adc39ec481b8ee72a001845fb14613c48e310a6016c

Observation 1ed826fb-88ba-4124-8c89-3ed2c5b9db55 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Taming Audio VAEs via Target-KL Regularization High-resolution image synthesis with latent diffusion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.547408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:103dcafea410c1436d3b0ac27a35d5d9ba619c3bb54e0bd2285c31a244999a6d

Observation 92576b4d-5b78-447a-ad0d-bee40344e3ca · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

Taming Audio VAEs via Target-KL Regularization Audiolm: a language modeling approach to audio generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.509831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:3588e47f098d8d7f7da6bb46528eb0f206b942e4dd504f1c6b5467f99c086149

Observation 4770f34f-9e9f-4ada-8db9-fd99cd8978ec · outbound

This paper cites Stable audio open.

Taming Audio VAEs via Target-KL Regularization Stable audio open

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.530193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:3e3f2ca64da30f3303c8b4a83490d4d29ed65a3b88bf44c60c693111308d785a

Observation 77cf3431-ac9d-4fce-b064-c5ca517ad311 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Taming Audio VAEs via Target-KL Regularization Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.187655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:8425c4fb5adb5cbfc2fca6522e0cd03196a00a7a51eb66b52c0f87d382d45a72

Observation fd56e338-9c70-4e3b-b60e-b98286b6bad7 · outbound

This paper cites VampNet: Music Generation via Masked Acoustic Token Modeling.

Taming Audio VAEs via Target-KL Regularization VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.218495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:42cb3bae9132d537bdcd80ea48b12070a7a235020d29ac56149f4254b8e200f9

Observation b8583be2-aef0-4be5-98d7-8b637f38461b · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Taming Audio VAEs via Target-KL Regularization MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.231264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:14afc8455d764ae59be8c8d8e8354bdbabed4a5b69d5f27f71f0bd5743355834

Observation 76880868-1f65-483d-a52e-60e57fb1c33d · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Taming Audio VAEs via Target-KL Regularization SoundStorm: Efficient Parallel Audio Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.221489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:61d2aac30572df45aeda5b751de323e870f38b699723a07c1236dd92e2cc5763

Observation 7d7aca55-741a-471d-add7-dd127d086ea2 · outbound

This paper cites Auto-Encoding Variational Bayes.

Taming Audio VAEs via Target-KL Regularization Auto-Encoding Variational Bayes

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.205869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:10b017bc2df5d9436a3b097467ab0cd84d2b1debf7d360711feda6d8aeaef4cb

Observation 7a56c53d-6cf0-4c59-af0d-26f022923b05 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Taming Audio VAEs via Target-KL Regularization Denoising dif- fusion probabilistic models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.543609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:be05df8ed6fca73d2f4a56d1974b1e7b45b36ffb8ce6551305fdc7e5e2930a81

Observation 933421a4-e9a7-4c06-9768-f425aac0e273 · outbound

This paper cites Au- dioLDM: Text-to-audio generation with latent diffusion mod- els.

Taming Audio VAEs via Target-KL Regularization Au- dioLDM: Text-to-audio generation with latent diffusion mod- els

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.534344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:5dfe8220782199f277b494ba88e1e07b08eb0eabe67cc234b376c08d40151440

Observation 2d4bb04f-ccf6-4b91-85a7-7c2c06ce171b · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Taming Audio VAEs via Target-KL Regularization Scaling rectified flow transformers for high-resolution image synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.518704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:aba46699fb870b46389ca31a00a97ca52a2c7408f90d8c326644b9a7f4fa22e1

Observation 83b4282f-b643-445d-bc11-3de7a17fbed1 · outbound

This paper cites Soundstream: An end-to- end neural audio codec.

Taming Audio VAEs via Target-KL Regularization Soundstream: An end-to- end neural audio codec

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.505528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:bc428dc985179da5598a5ff66223e5f13df6a8722d76c9211a5379b5ce1247a3

Observation 9d243bcb-1239-48e8-81cd-3fc20255b94c · outbound

This paper cites High-fidelity audio compression with improved rvqgan.

Taming Audio VAEs via Target-KL Regularization High-fidelity audio compression with improved rvqgan

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.514421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:4a093427296bc6a7d537c610dc8dd3b1bd95d2d606a2d247a48e91bd4dd71493

Observation 3a2cb772-818c-41ff-aa15-7eadf2cf56a9 · outbound

This paper cites In- terpreting rate-distortion of variational autoencoder and using model uncertainty for anomaly detection.

Taming Audio VAEs via Target-KL Regularization In- terpreting rate-distortion of variational autoencoder and using model uncertainty for anomaly detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.526220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:e1924a1fd19c8ac5b0f160cd60a3a6eac9c9561264813e5843b9215a0a1689e8

Observation 88592c11-9e50-43de-ad61-c73f6bf38523 · outbound

This paper cites Practical Lossless Compression with Latent Variables using Bits Back Coding.

Taming Audio VAEs via Target-KL Regularization Practical Lossless Compression with Latent Variables using Bits Back Coding

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.237386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:9b42ab91c59c50826fa40724094a80006cff15f35d3a1255ea8ba17d98634d70

Observation e3159072-c9d8-4185-88e6-4e2bac5f83f3 · outbound

This paper cites Fixing a Broken ELBO.

Taming Audio VAEs via Target-KL Regularization Fixing a Broken ELBO

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:53:23.215261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:e855d21e433f2bcf7454d5a3d8d6b61036fd59f2799885021b1e3fe62f163612

Observation bdb16ce2-de3e-455c-9a8a-64ca6abb24a0 · outbound

This paper cites Improved variational in- ference with inverse autoregressive flow.

Taming Audio VAEs via Target-KL Regularization Improved variational in- ference with inverse autoregressive flow

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.507725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:6f0c51b113a5c0f7fe62c9e8bd57e7293ecdae0c4c2587b18edcc5cf08722150

Observation c62bb11d-3d35-44f0-bea4-db2b52abbf64 · outbound

This paper cites An introduction to variational autoencoders.

Taming Audio VAEs via Target-KL Regularization An introduction to variational autoencoders

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.539824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:f1c0358b9176d38dfc8cbb0bb2dd5406dc0271c25e5b439cdc76f220a2193581

Observation 636c0cda-1aa7-491a-803c-a712e014276f · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

Taming Audio VAEs via Target-KL Regularization BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.209087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:75f9cc18906e9ba4d4a5f681b7be9e5f6f4bde1a27d2d612b58cf3144939771a

Observation c54fba4e-1e70-448e-97f0-4089dc6858d8 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Taming Audio VAEs via Target-KL Regularization Moshi: a speech-text foundation model for real-time dialogue

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.212392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:cb4610d776d66846697adb5ca48ba66f8de070a4cb4aa8970d8955fafb792d72

Observation 9cf33edb-f31d-4617-8a83-2def3c1439e0 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

Taming Audio VAEs via Target-KL Regularization Audio set: An ontology and human-labeled dataset for audio events

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.500959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:3f78099c1dcd43897163632c64ed6b0a703b91a66f32179107c4e8497d7e00d3

Observation 20135b68-3616-4b79-bada-563d7092f5ab · outbound

This paper cites SpectroStream: A Versatile Neural Codec for General Audio.

Taming Audio VAEs via Target-KL Regularization SpectroStream: A Versatile Neural Codec for General Audio

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.193610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:8e59bed288f442f1ef332f623f6cb4392ccd30dc61070d95fa4f09cdc614dad6

Observation 6b273d94-6d2f-4367-8629-6fd27bb371f1 · outbound

This paper cites High Fidelity Neural Audio Compression.

Taming Audio VAEs via Target-KL Regularization High Fidelity Neural Audio Compression

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.184582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:a40e6b3e6c0b3a586e0092af3be96f3ceef74fd4808f07ecd99f3a4873b769b5

Observation a1ea9b96-7ea7-4f91-a008-fa96c690987a · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Taming Audio VAEs via Target-KL Regularization Progressive Distillation for Fast Sampling of Diffusion Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.234423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:61324d0aba35cc2f9768605c900fcb7aa347f11568b30b365be2073ac8ec7efb

Observation c30fea26-bd78-456d-8ad6-14f4103c5604 · outbound

This paper cites sim- ple diffusion: End-to-end diffusion for high resolution im- ages.

Taming Audio VAEs via Target-KL Regularization sim- ple diffusion: End-to-end diffusion for high resolution im- ages

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.511990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:c80a4b3c8f2348e21b56ee81898fd3876bb7cb8a14725432aa8eecfc02a0969f

Observation 7262eaa6-04e7-4c77-96c1-8d0a3fb68ae6 · outbound

This paper cites Simple-tts: End-to-end text-to-speech synthesis with latent diffusion.

Taming Audio VAEs via Target-KL Regularization Simple-tts: End-to-end text-to-speech synthesis with latent diffusion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.538092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:a9c3bede1bf54445ce50c9e0f9f1f8608d37219f538684a10ab2f9288a3171bf

Observation 9c9e3d41-425b-40ee-987f-4911387a26f6 · outbound

This paper cites Scalable diffusion models with transformers.

Taming Audio VAEs via Target-KL Regularization Scalable diffusion models with transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.532380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:fd896f641d0442b134b02c811a52202ebed3a529c41100f978f8914bd1534504

Observation 69d5631a-61d2-4e30-9e6a-6acd53f2c2cf · outbound

This paper cites Ditto-tts: Efficient and scalable zero-shot text-to-speech with diffusion transformer.

Taming Audio VAEs via Target-KL Regularization Ditto-tts: Efficient and scalable zero-shot text-to-speech with diffusion transformer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.524065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:f2297b4fcaa723b8a2f680eb86df2500966046f4327ec617dad37cb5aec8b6dc

Observation 054760b5-1092-498f-a46d-067d866ca8d2 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

Taming Audio VAEs via Target-KL Regularization Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.224858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:03cb7c751d96561c62eb4708809edc2a53ab4abf85fe25cb230e4de63d846955

Observation e0306326-c226-411b-a171-bd5f9c80c4fb · outbound

This paper cites Byt5: Towards a token-free future with pre-trained byte-to- byte models.

Taming Audio VAEs via Target-KL Regularization Byt5: Towards a token-free future with pre-trained byte-to- byte models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.521781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:c68a4b17ef46bc0407e6c00eae2797817f8af0034b9916e0e7ea0410ca68d778

Observation 9dd8e9e3-17a6-43e0-b1c3-47de3dde233d · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Taming Audio VAEs via Target-KL Regularization Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.536255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:c2b8a239412bcd39ddf4610f77098fa704df6c2b9c32c181c35b7d17a1cd17b4

Observation 6c3d5fca-e480-47fb-8de3-1a6bfa14f44e · outbound

This paper cites Phonemizer: Text to phones transcription for multiple languages in python.

Taming Audio VAEs via Target-KL Regularization Phonemizer: Text to phones transcription for multiple languages in python

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.541832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:e8b4da574516f17a4de9453574f2df606110f4a536ac822306c25bbf62ba57da

Observation 3c53ee44-f5df-41db-84f7-c92726aefef5 · outbound

This paper cites Emilia: A large-scale, extensive, multilin- gual, and diverse dataset for speech generation.

Taming Audio VAEs via Target-KL Regularization Emilia: A large-scale, extensive, multilin- gual, and diverse dataset for speech generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.227927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:bda7fe6901a1fb6c7ec00a593b43766349fffe29d3c28cc7a499fae0a49b656c

Observation 3e4bfec1-d8dd-4d70-8845-cb42683db895 · outbound

This paper cites SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation.

Taming Audio VAEs via Target-KL Regularization SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.240224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:d79584e40d99000c834a59df746a0af85bda9b9d5400bb44d02b21b65e2af171

Observation 4492068a-99c8-41aa-b65c-f503dcfe9086 · outbound

This paper cites Sketch2sound: Controllable audio generation via time-varying signals and sonic imitations.

Taming Audio VAEs via Target-KL Regularization Sketch2sound: Controllable audio generation via time-varying signals and sonic imitations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.503422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:7f791b9bdcfea960f1a2c3854c2795b4cb3f220dac1efa51509c88b2f8751ac4

Observation 2deef7fe-17c5-41e7-86e9-fd1420012cc4 · outbound

This paper cites FLAM: Frame-wise language-audio model- ing.

Taming Audio VAEs via Target-KL Regularization FLAM: Frame-wise language-audio model- ing

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.516502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:3fadda1736524533d6ffff0d0180503ecbfe39af7ddafd9f916c5abc16ec6817

Observation f9db439e-e46e-4f1f-ace9-3ae10376f717 · outbound

This paper cites Scaling instruction- finetuned language models.

Taming Audio VAEs via Target-KL Regularization Scaling instruction- finetuned language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T14:53:23.549390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:cf3291a7820bd610554114fa57ef54f0fd3b0a7c1a9e128f016e20cb9e10938d

Observation 59e97aa2-3ac4-4d1f-b1a7-bde25cad6181 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Taming Audio VAEs via Target-KL Regularization Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.190532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:155b0b0bfaf5ec2cff1a4e6efcdf52319377313e5d66491f9c2e700b46ba8e1e

Observation 76bbbf21-7af8-4adf-9c7d-07e30845c50b · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Taming Audio VAEs via Target-KL Regularization Finite Scalar Quantization: VQ-VAE Made Simple

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.196540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:ee1ebfad0bedfcf4c496fe74b2ccab2280b3d8af3d8521c7e950ba7b66cd48b6

Observation 6d3b93c9-7adf-4e9d-ae65-8fa20a0d5050 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

Taming Audio VAEs via Target-KL Regularization Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.202963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:91c1262423967d3bdef4dbad082789339017c2a32d54cb070fca9c5168291143

Pith citing papers

Observation 08e5bc35-139f-4b17-88df-1dfb9801d680 · inbound

Taming Audio VAEs via Target-KL Regularization cites this paper.

Taming Audio VAEs via Target-KL Regularization Taming Audio VAEs via Target-KL Regularization

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:53:23.199683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:f83095df64aec0ccd59443e18c16b4bd858909031bb74ebb7f7bc373a6ea110d