Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:24.207276Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.23582.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:24.207276Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:21.333361Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T21:41:24.430826Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f56a6ed1-d90d-4f35-9f31-859ccba6066f · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio a dog barking behind a human speech,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e9e26b03-e55f-4363-926e-a8e6934b2c7c · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5bceae05-baab-4c89-aa9f-cbc2d88d8296 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7312557b-8972-4ae4-99b9-a421dd381213 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio 2” or higher, we added the exclusion of listeners with an average original audio rating of “6
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bca3e593-4e97-4146-829b-69f2ed23b05f · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bfb8ae9-087d-49b3-b354-e047c79026ea · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0854e656-819b-4fdc-8598-263f139e5827 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b1d044a9-a83d-4535-aae2-d532992e074b · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b5648097-787c-462b-a79d-9917e23c3e0b · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Diffsound: Discrete diffusion model for text-to-sound genera- tion,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0c15c510-645e-4f29-a594-5ed720865bd9 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Sound synthesis for impact sounds in video games,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 230d1527-c23a-4766-8c98-9676f450937b · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Challenge on sound scene synthe- sis: Evaluating text-to-audio generation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9d038414-0e20-426e-a83e-519bc03d6f4f · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio The V oiceMOS Challenge 2024: Beyond speech quality prediction,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 575f9da5-92f0-4e6c-8278-bf9af17f6c91 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Subjective-aligned dataset and metric for text-to-video quality assessment,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 105ac452-f0d2-4310-a374-341099b3a6e2 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Environmental sound synthesis from vocal imitations and sound event labels,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9ccadc8a-8fb6-4579-992f-806bef2474af · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b0ebc53f-c036-4544-a474-bcd78621ca1b · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Text-to- audio generation using instruction guided latent diffusion model,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 12b470e2-c5a9-43ce-904c-3825d3b472df · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7203e59d-5374-4215-b076-8c93d4ef1fb5 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio WavLM: Large-scale self-supervised pre-training for full stack speech processing,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c5ab75-e0cd-4bd2-af30-4519267f34f2 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Animal” category shows both of statistically significant differ- ences and interaction. Figure 2(a). shows that synthesized audio in the “Animal
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6beaa23e-2537-40ab-bcc8-1f182c8af643 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio PAM: Prompting audio-language models for audio quality assessment,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 818e93a2-30ca-4461-b3a1-625b0cfab09d · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Audio-text mod- els do not yet leverage natural language,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1f895f04-a02e-432f-bfb2-63baf86708a3 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio AudioCaps: Generat- ing captions for audios in the wild,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7dfe310c-937c-4c4f-a7fa-5adb96f9ac70 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio AudioLDM: text-to-audio generation with latent diffusion models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation df6faa69-4fec-4869-bc10-86d9f8f9d286 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio AudioLDM 2: Learn- ing holistic audio generation with self-supervised pretraining,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dfe3687a-0663-47d7-aca1-27601b5f7a6b · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 11b55f60-ec7d-4e44-80d3-aa24986f3f83 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Toward verifiable and repro- ducible human evaluation for text-to-image generation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b3b63564-ea32-4c26-af81-d021248b0511 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio On a test of whether one of two random variables is stochastically larger than the other,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a36147-225b-454a-b648-d2874a2bb133 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Use of ranks in one-criterion variance analysis,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cef4de0-9de5-49ab-bb0c-1693c021e7f6 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio A multiple comparison rank sum test: treatments versus control,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 76f4c52f-52b0-4d3e-98ec-35c0c826c1b3 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio The aligned rank transform for nonparametric factorial analyses using only anova procedures,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3b6b4a0-01fe-4463-b283-4ca45952e0ed · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio A new readability yardstick
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 69676ee5-e695-4454-967e-8d8054546456 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio BYOL for Audio: Exploring pre-trained general-purpose audio representations,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c88b6738-ae54-41de-89d5-dd104f49153c · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio LDNet: Unified listener dependent modeling in MOS prediction for syn- thetic speech,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 59f3a409-d1e9-448e-9e7d-27db25507d18 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Bidirectional LSTM networks for improved phoneme classification and recog- nition,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 32ee808a-97fb-44d5-8245-2d40a7161a49 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio MBNet: MOS prediction for synthesized speech with mean-bias network,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e5344c90-5046-4c45-baad-848b36ecb386 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Class- balanced loss based on effective number of samples,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b81dd3f9-befc-4967-a864-b4569594cff9 · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Adam: A method for stochastic op- timization,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 59903483-f256-4b56-ac5b-656a59630b3b · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2115d22c-8f3f-41b3-8b74-39f0a190b9de · outbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio CLAP: Learning audio concepts from natural language supervision,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e9e26b03-e55f-4363-926e-a8e6934b2c7c · inbound
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.