Pith. sign in

Paper Citation Record · LEDGER

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

As of 22 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2504.21815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21815 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:57:47.433651Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:57:47.249767Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T04:57:47.830028Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4cf1924a-2025-401e-a7b6-1796ad82b4e8 · outbound

This paper cites an unresolved cited work.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:57:48.215884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.244102Z digest=sha256:42f5195d6c07586e2c6a50bf24a74be2c59b61ac375e837f89f58ebd8af4b0cd

Observation 6d90708e-c42e-408a-b0ac-d99b91e74eb1 · outbound

This paper cites From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:57:47.834783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.249767Z digest=sha256:db42a00ea3d4af1f8b7a56308fb8ede25ab4cbcf0d573d3db2321cac20b419b4

Observation a677cbcd-5180-4be5-a9bd-baa87ef75644 · outbound

This paper cites an unresolved cited work.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:57:48.202200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.255187Z digest=sha256:24c753f5b66356b81119fe68be3d620163e9bfcbe7aa312246d9cd152685daa9

Observation cc880505-c91d-4713-ba40-49d21a8370b6 · outbound

This paper cites low quality.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems low quality

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.187932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.259974Z digest=sha256:a5225f8a2b56dd3b925fa47b6227dd55942f0cd9c80e0803d78315be76c49176

Observation f5756c59-4e1f-441c-938f-2c2b27cdce4e · outbound

This paper cites We examine reference-based evaluation metrics on the generated dataset computed in Section 4, offering a complementary perspective to subjective assessments.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems We examine reference-based evaluation metrics on the generated dataset computed in Section 4, offering a complementary perspective to subjective assessments

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.174473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.265373Z digest=sha256:dc99d01ed55fba71daab812220fa11faaf9a289bfeb032df6afb05ef89c55c0e

Observation 748cb1b8-505f-46c1-8363-626873488520 · outbound

This paper cites Our results reveal substantial incon- sistencies between different evaluation perspectives, highlighting the challenges of fully capturing human judgment through automated proxies.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Our results reveal substantial incon- sistencies between different evaluation perspectives, highlighting the challenges of fully capturing human judgment through automated proxies

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.160227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.270241Z digest=sha256:5237bdb6d0bd11c9e25d968fdaf160ed88d6ac780f52bfbce8265de944f0ab06

Observation 4cad6a93-1e79-487a-961c-8255a4a1b058 · outbound

This paper cites Acoustic scene generation with condi- tional SampleRNN,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Acoustic scene generation with condi- tional SampleRNN,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.146152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.274640Z digest=sha256:ae8cc790007a4572854219664143c8b87005249e4e3f2ece9cff18000affbb63

Observation 1ec9efc9-90c6-45da-99da-6d7fb5062a42 · outbound

This paper cites AudioGen: Textually guided audio genera- tion,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems AudioGen: Textually guided audio genera- tion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.132578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.279168Z digest=sha256:827ba521f91d6dfea21b191050d75382ec499e0a8a67ee1ab1b4c5d55a0e0a3c

Observation 263790c9-df72-4cc2-b01e-e74fa0af038f · outbound

This paper cites DExter: Learning and Controlling Performance Expression with Diffusion Models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DExter: Learning and Controlling Performance Expression with Diffusion Models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.118563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.283916Z digest=sha256:6e8b3202fc0208114d0439a7471fdaa431fc7ca0f227387444cb9de18ff727ff

Observation e15067ec-328f-4ba1-a97f-963ca8c703f6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.288199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.288199Z digest=sha256:95b0d6448f5605d9b14fa6004dc1ef06289e29a133b56937b9847e6a088d756e

Observation dea05df3-4aa9-463e-90b1-868cb15c4b7f · outbound

This paper cites Qwen2.5 Technical Report.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Qwen2.5 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.293130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.293130Z digest=sha256:46c45eeb609c38fe23693922ce2dd953b5c307232c000d34692082defd046793

Observation 21984b2c-a8a6-4ea1-b1a0-66d8afc92516 · outbound

This paper cites DSPO: Direct score preference optimization for diffusion model align- ment,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DSPO: Direct score preference optimization for diffusion model align- ment,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.104139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.298409Z digest=sha256:b63e419084e029fc2b5d0f92c5699b3f11c68bc8a19425c9e2cf2ea6bc3181a7

Observation 0a4454cb-d7b0-496d-835e-1e617e964cd5 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.302806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.302806Z digest=sha256:1a47c30e7e07ea5bf8538b31b8aca54f78f3ac5c1b54d91afde10537a693ca29

Observation 97b67597-354f-4ca0-a053-4efa3937514b · outbound

This paper cites 1, Association for Computing Machinery, 2024.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems 1, Association for Computing Machinery, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.090417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.307649Z digest=sha256:674ac19e2305cfc962221279f0d926dbbe9b57f545a332e4ba6764013eaacf9b

Observation 04fec364-3565-464e-8d24-9d25be516a88 · outbound

This paper cites BATON: Aligning text-to-audio model using human preference feedback,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems BATON: Aligning text-to-audio model using human preference feedback,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.076969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.311697Z digest=sha256:76cfd1b0f0b2709b0fc93828a1a9dab254aed5cb11e30e12fe28313b8dd526aa

Observation 270de9c1-f454-41b2-9ba2-b54894bab1ae · outbound

This paper cites DRAGON: Distributional rewards optimize diffusion generative models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DRAGON: Distributional rewards optimize diffusion generative models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.315838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.315838Z digest=sha256:d60d680d0185bad4dd34c5dc4f48c93dbfba6dd19af1f30cb5814dfafb9cde04

Observation bf4d53e4-f123-46d0-83fa-e685f3268f87 · outbound

This paper cites SMART: Tuning a symbolic music generation system with an audio domain aesthetic reward.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems SMART: Tuning a symbolic music generation system with an audio domain aesthetic reward

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:57:47.709388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.320069Z digest=sha256:e8d0de9fdfde6cb1bb0a7198eb84a52ad9448316ab958565c7e2dd6eb7230a0f

Observation 9858a7fc-3339-4148-9bae-3beae96a5fad · outbound

This paper cites Aligning Text-to-Music Evaluation with Human Preferences.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Aligning Text-to-Music Evaluation with Human Preferences

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.324607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.324607Z digest=sha256:6d3a3484f266dd472c7ea726bf3e98dd67e8ecf05182a8f99fbd95e8bd6b30d7

Observation a78cb8af-1d39-436a-8cd2-25fcb146a866 · outbound

This paper cites KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.329053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.329053Z digest=sha256:3eac0d2ff2096f76c8a3ba8c566c15fa067fc054ea1bb126edc3b5d6600e6c45

Observation 3b40369f-1a92-466d-8fc3-8383255a9adf · outbound

This paper cites WavCraft: Audio editing and generation with large language models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems WavCraft: Audio editing and generation with large language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.063073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.333481Z digest=sha256:4ede33a2d4b7039acced9cf7702089266ac5885885d5bba188d0fecb31cdb04a

Observation 6a11a3b2-b24f-4853-87a4-2747096b6ca4 · outbound

This paper cites WavJourney: Compositional audio creation with large language models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems WavJourney: Compositional audio creation with large language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.048898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.337921Z digest=sha256:48868acb81d6c42ec48c91ff9cfc9f33924c716724e93729a865df7fe9282978

Observation 1216f0f5-e14e-4400-9816-8a87becf1229 · outbound

This paper cites Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.342303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.342303Z digest=sha256:016e1174fa447f5fde468dca5698dec9cd514757e98b012a6b5fa31a6f984104

Observation fe47a6b5-73ee-402c-b1a6-0c9ebede57bb · outbound

This paper cites Hierarchical Symbolic Pop Music Generation with Graph Neural Networks.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Hierarchical Symbolic Pop Music Generation with Graph Neural Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.346856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.346856Z digest=sha256:5b53b636f64c01991e3bb8066d21196af18256aace065bcdb665a5c6d57d246e

Observation 2943ddd5-9d71-441b-bef7-e8db6ebeeb4b · outbound

This paper cites RenderBox: Expressive Performance Rendering with Text Control.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems RenderBox: Expressive Performance Rendering with Text Control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.351237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.351237Z digest=sha256:351fcbc06cb156679dad5577da96367a4c01d7301f013b8122e7c88ef7066524

Observation 1f301de7-9b12-4a39-abe0-7b2b434be3fc · outbound

This paper cites Leveraging pre-trained audioldm for sound generation: A benchmark study,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Leveraging pre-trained audioldm for sound generation: A benchmark study,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.034883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.355679Z digest=sha256:5465ea7c98e30fdfac41d5b40b6049dabdab20b6ebdd45fff271932dba3d11a8

Observation 4d29258e-06ab-4c71-87a2-4e204c684abf · outbound

This paper cites Diffsound: Discrete diffu- sion model for text-to-sound generation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Diffsound: Discrete diffu- sion model for text-to-sound generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.021286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.360127Z digest=sha256:770e46fd4cedb7fdc15b5fff8da5822d30f823d88c9f47d1c0e05af3067ccd1f

Observation 6b059db3-0c0e-465d-9a3b-635fc6f088b5 · outbound

This paper cites Adapting frechet audio distance for generative music evaluation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Adapting frechet audio distance for generative music evaluation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.007803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.364505Z digest=sha256:fb611aaa5ed8bc9dbb4efeb52f86dd312d405631871724a4d850341f9ae77f74

Observation 55d0523a-c8ad-4a58-8426-2b6794426a8a · outbound

This paper cites Zero-shot unsupervised and text-based audio editing using DDPM inversion,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Zero-shot unsupervised and text-based audio editing using DDPM inversion,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.993797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.368632Z digest=sha256:cbf7998703377756d9f5ea7eedd6273658032e3a4446396a80af1adf5a77817d

Observation 09b81ee9-7b6f-4e12-9723-5b6787abab2b · outbound

This paper cites AudioMorphix: Training-free audio editing with diffu- sion probabilistic models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems AudioMorphix: Training-free audio editing with diffu- sion probabilistic models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.980370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.372707Z digest=sha256:94f0ec8cf830cb585f7609e892ff90e04dac7c9d22020829628b0eb130caeabf

Observation ad3c2866-b980-487e-ab89-467c7df63f34 · outbound

This paper cites A comparison of deep learning MOS pre- dictors for speech synthesis quality,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems A comparison of deep learning MOS pre- dictors for speech synthesis quality,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.965893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.377051Z digest=sha256:eb82bf697dbc6319896e02414571da3ab67fade13713a2f37ba9737eb4fc7c91

Observation 1558adc6-5a1b-4109-83b4-9dde11ff723e · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.381234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.381234Z digest=sha256:2b3f55b1f05115b4254bd026b36cc0ee1cb27a8aeb64d03d316fd284acff703e

Observation e846aac5-a3a4-4cd6-8382-3a1863dbbf31 · outbound

This paper cites From Audio Encoders to Piano Judges: Benchmarking Performance Under- standing for Solo Piano,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems From Audio Encoders to Piano Judges: Benchmarking Performance Under- standing for Solo Piano,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.952196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.385669Z digest=sha256:4105c351bbc27772f51dcf23d8ff60bc2ade31ce8e34bb73bec42dbf325b7c2d

Observation dda5b3b9-e940-4e24-916c-524783aad0f3 · outbound

This paper cites Piano Skills Assessment,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Piano Skills Assessment,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.937513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.389788Z digest=sha256:eb7f2b010c943bd6c74afd605883589c690b27a793ab60a553df6d2c105a9431

Observation 998adf1f-4074-453f-bba7-18b6777c6225 · outbound

This paper cites LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.923524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.394180Z digest=sha256:6124562bd5e09657e6647ced15726475a1624eb475fa63d8717692f00590ac12

Observation 8e010328-1240-4842-8bea-849015f58b41 · outbound

This paper cites MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.398280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.398280Z digest=sha256:478bdbe6f072e3b9c6629e5ebe751d4b618b498f23b595f23fc54e3643d1ee7c

Observation c6744b11-d13b-4c25-896e-f467dc0231f9 · outbound

This paper cites How does the teacher rate? Observations from the NeuroPiano dataset,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems How does the teacher rate? Observations from the NeuroPiano dataset,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.908177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.402747Z digest=sha256:fef9e607cdc22787f41ebce37a096435b8a6ade3f03b04da79da1588fa1c1bc6

Observation 3f4fcfd6-91c2-40c6-9266-2dc96d8a1da0 · outbound

This paper cites Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.893180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.407376Z digest=sha256:84cb5f66cdb257907662611e4f10240e62a6268398493d9b85a60258625cd4c9

Observation 84a339d1-68ee-441f-86ed-c914ee47a8ff · outbound

This paper cites Stable Audio Open.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Stable Audio Open

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.411635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.411635Z digest=sha256:e76e151dba99114e24a23a2f52af3cb6564c05371c24b860b2c188d1f962a76e

Observation 19aed1bd-d670-4ec6-b619-7e7231e7a0aa · outbound

This paper cites Simple and controllable music generation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Simple and controllable music generation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.878743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.416011Z digest=sha256:161afef3abc8ed0f5e82df1d5eed090238e8a635aa95e86d1c0842c190e2d3f8

Observation 576d0a2c-d804-4a90-8709-69a1647f8eed · outbound

This paper cites Yue: Scaling open foundation models for long-form music generation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Yue: Scaling open foundation models for long-form music generation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.420722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.420722Z digest=sha256:37118328752f1517ffc441ad9e868bbd3350e296f82b0c6cd6ba2806755008bf

Observation 012bcc6c-7aa6-46ec-8959-cd14eaf12736 · outbound

This paper cites DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.425009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.425009Z digest=sha256:34341f536e5a027fa6c69aee9a4e97f5345d38a1cc90b29ed78fe1d36032b303

Observation 64e66076-b27c-452f-81ff-b8209cd95143 · outbound

This paper cites LP-MusicCaps: LLM-based pseudo music captioning,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems LP-MusicCaps: LLM-based pseudo music captioning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.864251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.429436Z digest=sha256:e0ca69aa87a2b3370cd1263342e7c6f860159403b2a4af399c0c79e7cbd799f6

Observation b11a7715-a5a8-4d16-a5a2-6f7b40aeebd0 · outbound

This paper cites PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.849442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.433651Z digest=sha256:a4b40534c1cf6cf13d01079f7a4428edbbc1883adaee2e698163294aec5fb495

Pith citing papers

Observation 6d90708e-c42e-408a-b0ac-d99b91e74eb1 · inbound

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems cites this paper.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:57:47.834783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:57:47.249767Z digest=sha256:db42a00ea3d4af1f8b7a56308fb8ede25ab4cbcf0d573d3db2321cac20b419b4