Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:57:47.433651Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2504.21815.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:57:47.433651Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:57:47.249767Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-16T04:57:47.830028Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4cf1924a-2025-401e-a7b6-1796ad82b4e8 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6d90708e-c42e-408a-b0ac-d99b91e74eb1 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a677cbcd-5180-4be5-a9bd-baa87ef75644 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cc880505-c91d-4713-ba40-49d21a8370b6 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems low quality
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f5756c59-4e1f-441c-938f-2c2b27cdce4e · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems We examine reference-based evaluation metrics on the generated dataset computed in Section 4, offering a complementary perspective to subjective assessments
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 748cb1b8-505f-46c1-8363-626873488520 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Our results reveal substantial incon- sistencies between different evaluation perspectives, highlighting the challenges of fully capturing human judgment through automated proxies
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4cad6a93-1e79-487a-961c-8255a4a1b058 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Acoustic scene generation with condi- tional SampleRNN,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1ec9efc9-90c6-45da-99da-6d7fb5062a42 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems AudioGen: Textually guided audio genera- tion,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 263790c9-df72-4cc2-b01e-e74fa0af038f · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DExter: Learning and Controlling Performance Expression with Diffusion Models,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e15067ec-328f-4ba1-a97f-963ca8c703f6 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dea05df3-4aa9-463e-90b1-868cb15c4b7f · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Qwen2.5 Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21984b2c-a8a6-4ea1-b1a0-66d8afc92516 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DSPO: Direct score preference optimization for diffusion model align- ment,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0a4454cb-d7b0-496d-835e-1e617e964cd5 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97b67597-354f-4ca0-a053-4efa3937514b · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems 1, Association for Computing Machinery, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 04fec364-3565-464e-8d24-9d25be516a88 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems BATON: Aligning text-to-audio model using human preference feedback,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 270de9c1-f454-41b2-9ba2-b54894bab1ae · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DRAGON: Distributional rewards optimize diffusion generative models,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4d53e4-f123-46d0-83fa-e685f3268f87 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems SMART: Tuning a symbolic music generation system with an audio domain aesthetic reward
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9858a7fc-3339-4148-9bae-3beae96a5fad · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Aligning Text-to-Music Evaluation with Human Preferences
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a78cb8af-1d39-436a-8cd2-25fcb146a866 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b40369f-1a92-466d-8fc3-8383255a9adf · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems WavCraft: Audio editing and generation with large language models,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6a11a3b2-b24f-4853-87a4-2747096b6ca4 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems WavJourney: Compositional audio creation with large language models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1216f0f5-e14e-4400-9816-8a87becf1229 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe47a6b5-73ee-402c-b1a6-0c9ebede57bb · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Hierarchical Symbolic Pop Music Generation with Graph Neural Networks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2943ddd5-9d71-441b-bef7-e8db6ebeeb4b · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems RenderBox: Expressive Performance Rendering with Text Control
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f301de7-9b12-4a39-abe0-7b2b434be3fc · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Leveraging pre-trained audioldm for sound generation: A benchmark study,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4d29258e-06ab-4c71-87a2-4e204c684abf · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Diffsound: Discrete diffu- sion model for text-to-sound generation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6b059db3-0c0e-465d-9a3b-635fc6f088b5 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Adapting frechet audio distance for generative music evaluation,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 55d0523a-c8ad-4a58-8426-2b6794426a8a · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Zero-shot unsupervised and text-based audio editing using DDPM inversion,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 09b81ee9-7b6f-4e12-9723-5b6787abab2b · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems AudioMorphix: Training-free audio editing with diffu- sion probabilistic models,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ad3c2866-b980-487e-ab89-467c7df63f34 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems A comparison of deep learning MOS pre- dictors for speech synthesis quality,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1558adc6-5a1b-4109-83b4-9dde11ff723e · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e846aac5-a3a4-4cd6-8382-3a1863dbbf31 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems From Audio Encoders to Piano Judges: Benchmarking Performance Under- standing for Solo Piano,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dda5b3b9-e940-4e24-916c-524783aad0f3 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Piano Skills Assessment,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 998adf1f-4074-453f-bba7-18b6777c6225 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8e010328-1240-4842-8bea-849015f58b41 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6744b11-d13b-4c25-896e-f467dc0231f9 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems How does the teacher rate? Observations from the NeuroPiano dataset,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f4fcfd6-91c2-40c6-9266-2dc96d8a1da0 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 84a339d1-68ee-441f-86ed-c914ee47a8ff · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Stable Audio Open
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19aed1bd-d670-4ec6-b619-7e7231e7a0aa · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Simple and controllable music generation,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 576d0a2c-d804-4a90-8709-69a1647f8eed · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Yue: Scaling open foundation models for long-form music generation,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 012bcc6c-7aa6-46ec-8959-cd14eaf12736 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64e66076-b27c-452f-81ff-b8209cd95143 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems LP-MusicCaps: LLM-based pseudo music captioning,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b11a7715-a5a8-4d16-a5a2-6f7b40aeebd0 · outbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6d90708e-c42e-408a-b0ac-d99b91e74eb1 · inbound
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.