Pith. sign in

Paper Citation Record · LEDGER

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction

As of 22 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.05899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05899 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:16:56.451078Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bc31a3e-2706-46af-8798-ee2fc057ec8a · outbound

This paper cites MusicLM: Generating Music From Text.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.146491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.146491Z digest=sha256:d5f3e305168220c2a2fbe81c42056cdaf3937113d6ae2248ec79e37a7e93054f

Observation ad315d80-4b72-4627-83d5-86978052c1e7 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.194385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.194385Z digest=sha256:f4c594d874f0078742ccd573f0a88d31dac207de35d6cfd48f2912a678ec2b32

Observation bcc3cde5-97b3-4695-8872-6c514b116744 · outbound

This paper cites Fast timing- conditioned latent audio diffusion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Fast timing- conditioned latent audio diffusion,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.613505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:54.243298Z digest=sha256:015149dc71bd02f260c031059f802b814e81050b082255ccefc41295bc377c40

Observation c2eda0e4-5e05-4345-8d6b-d655fcd3f419 · outbound

This paper cites Musiceval: A generative music dataset with expert ratings for automatic text-to-music evaluation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musiceval: A generative music dataset with expert ratings for automatic text-to-music evaluation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.450106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:54.337412Z digest=sha256:c463f68ec2221edfbc35500dec188521b1c0ec263e0a348927757908c4481971

Observation faae7500-a287-40ef-b230-0f8e7653af5a · outbound

This paper cites SSL-MOS: A Self-Supervised Learning Based Approach with A Transformer Target Model For MOS Pre- diction,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction SSL-MOS: A Self-Supervised Learning Based Approach with A Transformer Target Model For MOS Pre- diction,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.291059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:54.454075Z digest=sha256:d1a1daf570f19bee9b0922f50509b0428c2d7008491d130a532675e76198a07d

Observation c88a3958-4fb3-4d70-9463-13037c7bb64e · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.161316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:54.513674Z digest=sha256:1dba1eabe650784a0b85a01ed127e9e0c99ecb439cab8c5eb84962160200652b

Observation 009ebc16-5eae-417e-a12e-54ec420aee6e · outbound

This paper cites MOSNet: Deep learning based objective assessment for voice conver- sion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MOSNet: Deep learning based objective assessment for voice conver- sion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.028276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:54.583734Z digest=sha256:0791df1d026ed8c6ed34e5b0205e0c1cce5bf05320178ca8ffbcae520c2e8662

Observation 9d4e82e1-9659-4400-b67a-79047a6b82c5 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Robust speech recognition via large-scale weak supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.648042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.648042Z digest=sha256:1db3d1f78bb9f493f408f5f243e5d3c1156113b9f4c191f7808258acc030db4a

Observation ab3abf6f-2794-4f18-af57-8b4cd7e9445c · outbound

This paper cites Qwen Technical Report.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Qwen Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.739123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.739123Z digest=sha256:da1308cd88babeb835f8793e23043efc85d2a4c92ae058c528a4281021ee035b

Observation f12e455d-4998-41ec-b50a-b64df0ffd3f0 · outbound

This paper cites Efficient optimal transport algorithm by accelerated gradient descent,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Efficient optimal transport algorithm by accelerated gradient descent,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.868162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:54.811164Z digest=sha256:e589c1d0e2d6d0a75f780f1b1857b92feb47cda3238b969cd64e8e790feb17ad

Observation eceb604f-d319-4102-b005-320ad57c850d · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.921831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.921831Z digest=sha256:82d5b9903a43b713e5ce7d7b3faefee4c8e2723c4fc40fa68bbd6da81b15d5df

Observation 2fb0be64-6805-4edd-bb43-a0989960304e · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.010659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.010659Z digest=sha256:77eddf09acb921868d5b7c08ed225eb74055d5a0f2d68287a79c28839d14004c

Observation 6cf873c1-3055-466d-8794-ed5457ae00cb · outbound

This paper cites Simple and controllable music generation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Simple and controllable music generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.722622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:55.126990Z digest=sha256:356fb2d1a40f26a9a83d7122e135081e129812e1acb119f93d2551638158ec58

Observation b117bfbd-7da0-466e-a6db-321892dbb1ee · outbound

This paper cites Mo ˆusai: Efficient text-to-music diffusion models,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mo ˆusai: Efficient text-to-music diffusion models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.515398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:55.247126Z digest=sha256:1a9b6f4d9aec75b856c5394a8d7768556534f6d0ed866a101abda2e4dbfb0be5

Observation 110b80d9-f855-4f71-a97d-7964703822bf · outbound

This paper cites Musicmagus: Zero-shot text-to-music editing via diffu- sion models,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musicmagus: Zero-shot text-to-music editing via diffu- sion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.341182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:55.371848Z digest=sha256:088f44b25d0522ec87ce31817b4317a41861ef8b0f9882980a8aa11ae0d4cc2c

Observation 32b714ee-bd08-48cd-a7fa-e2067ef54485 · outbound

This paper cites High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.490675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.490675Z digest=sha256:914938f55ed5cf2b9e4abb3de3d61ba44010c8405136b9154c52dab4e1aa475b

Observation 948aa870-0d5d-40a3-8c25-d567d71de983 · outbound

This paper cites MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.607323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.607323Z digest=sha256:652f20963548d49a3e795c73499506141aae19a6ebc9b51c4e7f73fe1aa66a58

Observation fcda9ac2-ff87-4e34-9b23-1215077a15cb · outbound

This paper cites Mospc: Mos prediction based on pairwise comparison,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mospc: Mos prediction based on pairwise comparison,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.187547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:55.727593Z digest=sha256:bdd6f82e8cc19a6bc3a50dfb20fc6f78216eb64917373a7340cc2ae69a64ca20

Observation e78a8927-adc8-4697-9afa-7ca3b1b274bc · outbound

This paper cites Resource-efficient fine- tuning strategies for automatic mos prediction in text-to-speech for low- resource languages,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Resource-efficient fine- tuning strategies for automatic mos prediction in text-to-speech for low- resource languages,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.910254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:55.786881Z digest=sha256:6cb091a671f73db4669cf8207d5703ba8eace7b54f2196233a60e937a478e3a1

Observation 622b45c8-a352-4575-9197-7d1cc3bbd279 · outbound

This paper cites Zero- shot out-of-domain is no joke: Lessons learned in the voicemos 2023 challenge,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Zero- shot out-of-domain is no joke: Lessons learned in the voicemos 2023 challenge,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.645210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:55.905504Z digest=sha256:b6039be20d88dcd894b45aa777da7c8f03036f22d159327ca60834bf04e6dd78

Observation 83e39c97-ec97-486e-a77b-473dd56ec2ef · outbound

This paper cites APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:56.066230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:56.066230Z digest=sha256:7bf413c38c9eacc42ac2079dd2c5957a0ee258bf081732ced25bb9ea2c9e0165

Observation 0e99131c-60e1-43be-b897-8b27fd569fc6 · outbound

This paper cites LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:57.039584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:56.175634Z digest=sha256:03c2a597a06f429eb8f22719c90e71f6a017933e5ad57568dc443d004322f1e5

Observation 44d951d0-8696-4fcf-921d-c7575a48ecc7 · outbound

This paper cites U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:56.749500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:56.294362Z digest=sha256:d316338d4583cd14a6dcbdee7b09b955e9647683137cd8c34a9a2fc94e7e010f

Observation 258d12b3-f868-4a2d-b655-afd4cdcc71fe · outbound

This paper cites Cmot: Cross-modal mixup via optimal transport for speech translation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Cmot: Cross-modal mixup via optimal transport for speech translation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.395890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T10:16:56.451078Z digest=sha256:9457a113474936da23165c53443f032ae7460a55dbbc88b09d3ff5005f08f965

Pith citing papers

No inbound Pith citation observations are available.