Pith. sign in

Paper Citation Record · LEDGER

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction

As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.05899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05899 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:16:56.451078Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bc31a3e-2706-46af-8798-ee2fc057ec8a · outbound

This paper cites MusicLM: Generating Music From Text.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.146491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.146491Z digest=sha256:2f963c9047df6521ffaad6467c34e46ab63f642f892bdea58f2a1fc8992101ec

Observation ad315d80-4b72-4627-83d5-86978052c1e7 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.194385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.194385Z digest=sha256:35e9d33411c5141db2b3d305152236e4d3532a5da5262944d92ab70002bc501f

Observation bcc3cde5-97b3-4695-8872-6c514b116744 · outbound

This paper cites Fast timing- conditioned latent audio diffusion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Fast timing- conditioned latent audio diffusion,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.613505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:54.243298Z digest=sha256:d2d8aed0148476e2e84e03a33b114de26817dfc21dd9af01830bf201ebfefe0a

Observation c2eda0e4-5e05-4345-8d6b-d655fcd3f419 · outbound

This paper cites Musiceval: A generative music dataset with expert ratings for automatic text-to-music evaluation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musiceval: A generative music dataset with expert ratings for automatic text-to-music evaluation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.450106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:54.337412Z digest=sha256:7a9af82557a2a73e19ddb34300db2faed3a6b1237454a42f766f84aae9e41505

Observation faae7500-a287-40ef-b230-0f8e7653af5a · outbound

This paper cites SSL-MOS: A Self-Supervised Learning Based Approach with A Transformer Target Model For MOS Pre- diction,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction SSL-MOS: A Self-Supervised Learning Based Approach with A Transformer Target Model For MOS Pre- diction,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.291059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:54.454075Z digest=sha256:2c1ca89cd99813d126005c8d18cc87335b3407e9fe09727b5d51211423869eb4

Observation c88a3958-4fb3-4d70-9463-13037c7bb64e · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.161316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:54.513674Z digest=sha256:51d3885537476f7622eaf545eb6621218c780b52c81e35061fb8bcf7ad276699

Observation 009ebc16-5eae-417e-a12e-54ec420aee6e · outbound

This paper cites MOSNet: Deep learning based objective assessment for voice conver- sion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MOSNet: Deep learning based objective assessment for voice conver- sion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.028276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:54.583734Z digest=sha256:2a5492758f4de107ad75f2995ea46bb0e874ffae0139ef72b0519be52ae3805f

Observation 9d4e82e1-9659-4400-b67a-79047a6b82c5 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Robust speech recognition via large-scale weak supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.648042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.648042Z digest=sha256:f9764afeed99bee77871185cbaa307e77a3005c13ddcbbab5cb032b4723d3754

Observation ab3abf6f-2794-4f18-af57-8b4cd7e9445c · outbound

This paper cites Qwen Technical Report.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Qwen Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.739123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.739123Z digest=sha256:9d52f9d9c5b3ca7d59c4449696b486d6a45dead6be82ecfaf4c5c95670a06c1c

Observation f12e455d-4998-41ec-b50a-b64df0ffd3f0 · outbound

This paper cites Efficient optimal transport algorithm by accelerated gradient descent,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Efficient optimal transport algorithm by accelerated gradient descent,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.868162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:54.811164Z digest=sha256:d8900d8ed2acda2dbda4a4922d1448bd27fdf8a911786591ffbb3132bcd75111

Observation eceb604f-d319-4102-b005-320ad57c850d · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.921831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.921831Z digest=sha256:1bb3ebc12237e39befa91c3e8f7214856cbb59395ac841d62acd276081533a9f

Observation 2fb0be64-6805-4edd-bb43-a0989960304e · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.010659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.010659Z digest=sha256:19d10f396547b0f10c1427cd4b0ab5056d4f07cfa0269a86287ca0cee222dac3

Observation 6cf873c1-3055-466d-8794-ed5457ae00cb · outbound

This paper cites Simple and controllable music generation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Simple and controllable music generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.722622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:55.126990Z digest=sha256:5b9477d3596646db46c7835fa0754024796509c5f77514a880a60a60218c3db8

Observation b117bfbd-7da0-466e-a6db-321892dbb1ee · outbound

This paper cites Mo ˆusai: Efficient text-to-music diffusion models,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mo ˆusai: Efficient text-to-music diffusion models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.515398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:55.247126Z digest=sha256:9f9f88e1c75babb039776684d57b1e4e378f72d7f14548a56ee27a92342f5d2c

Observation 110b80d9-f855-4f71-a97d-7964703822bf · outbound

This paper cites Musicmagus: Zero-shot text-to-music editing via diffu- sion models,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musicmagus: Zero-shot text-to-music editing via diffu- sion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.341182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:55.371848Z digest=sha256:893ce3299bc710b0c4ef8f35cfde28ec1412cc7e218596fd0bc03c3bbb04f05e

Observation 32b714ee-bd08-48cd-a7fa-e2067ef54485 · outbound

This paper cites High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.490675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.490675Z digest=sha256:3eb7538939fbea14d8775a21e0364d0fe63adeca0b2c785472995d1bd42b2d36

Observation 948aa870-0d5d-40a3-8c25-d567d71de983 · outbound

This paper cites MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.607323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.607323Z digest=sha256:fa38c03fd5db8480db0bf83f0a0fc8dd75a80c0a3de9b10480e09093d2dbb127

Observation fcda9ac2-ff87-4e34-9b23-1215077a15cb · outbound

This paper cites Mospc: Mos prediction based on pairwise comparison,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mospc: Mos prediction based on pairwise comparison,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.187547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:55.727593Z digest=sha256:66eed4296f18082d6bb9097136a5565b7b06f55f736840f300ad8eeeabbf0287

Observation e78a8927-adc8-4697-9afa-7ca3b1b274bc · outbound

This paper cites Resource-efficient fine- tuning strategies for automatic mos prediction in text-to-speech for low- resource languages,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Resource-efficient fine- tuning strategies for automatic mos prediction in text-to-speech for low- resource languages,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.910254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:55.786881Z digest=sha256:579a56761c04d8d1f8c082e6dac777e1ba0d4641769ba627944d5f2d5015dacf

Observation 622b45c8-a352-4575-9197-7d1cc3bbd279 · outbound

This paper cites Zero- shot out-of-domain is no joke: Lessons learned in the voicemos 2023 challenge,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Zero- shot out-of-domain is no joke: Lessons learned in the voicemos 2023 challenge,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.645210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:55.905504Z digest=sha256:c08dd8f84156b7b2b7591b88a7d22208fbabb1b5626c517ede12a5db48609032

Observation 83e39c97-ec97-486e-a77b-473dd56ec2ef · outbound

This paper cites APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:56.066230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:56.066230Z digest=sha256:c6e75a5836b90f22640aa3b87f3acfca01c96b17548baefd9054aed8af93e2d5

Observation 0e99131c-60e1-43be-b897-8b27fd569fc6 · outbound

This paper cites LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:57.039584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:56.175634Z digest=sha256:a148e1554423896034d3f4bcf39abc222c1d663aab87cbe54d4d06cba167d5d1

Observation 44d951d0-8696-4fcf-921d-c7575a48ecc7 · outbound

This paper cites U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:56.749500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:56.294362Z digest=sha256:ab8c6bd5454c553be597fdfcee41bae9f19323fc9ffdf904c424bd9e17ea79aa

Observation 258d12b3-f868-4a2d-b655-afd4cdcc71fe · outbound

This paper cites Cmot: Cross-modal mixup via optimal transport for speech translation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Cmot: Cross-modal mixup via optimal transport for speech translation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.395890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:16:56.451078Z digest=sha256:a3a52898764d68d2c56c8c0e11ab31137a9df2fa5c5698b9978b1af9038578bb

Pith citing papers

No inbound Pith citation observations are available.