Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:48:46.626775Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2412.09789.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:48:46.626775Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T14:53:15.718359Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T14:53:23.238659Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 95bad160-aa81-4f78-b5ff-c140d91d4d23 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Stable Audio Open
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 541559a0-4b3b-41d3-a59b-ab3bed9e9af5 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation AudioLDM: Text-to- audio generation with latent diffusion models,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cd8b6482-9aa2-4ba4-bdb3-b51e71df9ff1 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Denoising diffusion prob- abilistic models,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f7c8f5f-b1c0-46a9-8f02-bbc5532c1cf4 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c66992e-5c2a-4837-b6c7-8a726636d3fd · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Simple and controllable music generation,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1027dca1-02c4-4226-af49-b809f9249866 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Compa: Addressing the gap in compositional reasoning in audio-language models,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f45c1f0f-cdf5-4754-bc36-44144c16a673 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation A Demand-Driven Perspective on Generative Audio AI
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eb686d45-1596-45c6-bdde-87fa5cdca707 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 241993d4-bac1-4f8f-b8e1-547179c129fd · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceaa1dff-4284-444d-a15a-1b22d0a23815 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Generative adversarial nets,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b0e16031-dc20-45c9-8d62-7e7a487310bf · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Tacotron: Towards End-to-End Speech Synthesis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3b2b77-d107-430e-95b4-d859b455c981 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Auto-Encoding Variational Bayes
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a677f8b-f682-4499-a7d2-9425e63bd858 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation AudioGen: Textually Guided Audio Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1135e13-7d02-4d6c-aed9-4297887c124d · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ff1037-5128-491d-b90c-23cad86a6c68 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Diffwave: A versatile diffusion model for audio synthesis,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c68c40c9-a15c-4e56-bbfd-8a1fbd1f1769 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Music controlnet: Multiple time-varying controls for music generation,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9effb47-a044-4fbd-aea7-e2b9298c6750 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Hierarchical Generative Modeling for Controllable Speech Synthesis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c4c451b-9dc2-4e8d-bfc3-967a60d57551 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Deep voice: Real-time neural text-to-speech,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation daa8b8fe-2100-419d-b792-eec6dbe7f438 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1e336f2c-2986-4316-a703-4f19b1ece08f · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Mistral 7B
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 995a5bcf-71d9-437f-804a-6fd8080ad16f · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ee1ac531-ddcc-4e94-8068-b0df41594ee5 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Crepe: A convolutional representation for pitch estimation,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8a25975b-8b2d-401a-9183-ac18f509fab8 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Suppression of acoustic noise in speech using spectral subtraction,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 27dcc94c-79c3-4d67-9e03-5f6f398fa878 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Fsd50k: an open dataset of human-labeled sound events,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 10696c29-daa1-4fca-9a8a-aafbfc793e38 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Make- an-audio: Text-to-audio generation with prompt-enhanced diffusion mod- els,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ffe3740a-9c03-453d-b962-963732da51d3 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Scalable diffusion models with transformers,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 672f74b3-78c4-459a-9ac3-164023fe7173 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Scaling instruction-finetuned language models,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffd59caa-dbe8-4432-b86c-877f2b71d670 · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation High-fidelity audio compression with improved rvqgan,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation beb07327-682b-48f2-9c6b-b7122c355d1a · outbound
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e4bfec1-d8dd-4d70-8845-cb42683db895 · inbound
Taming Audio VAEs via Target-KL Regularization SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.