Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:16:56.451078Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.05899.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:16:56.451078Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7bc31a3e-2706-46af-8798-ee2fc057ec8a · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLM: Generating Music From Text
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad315d80-4b72-4627-83d5-86978052c1e7 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcc3cde5-97b3-4695-8872-6c514b116744 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Fast timing- conditioned latent audio diffusion,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c2eda0e4-5e05-4345-8d6b-d655fcd3f419 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musiceval: A generative music dataset with expert ratings for automatic text-to-music evaluation,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation faae7500-a287-40ef-b230-0f8e7653af5a · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction SSL-MOS: A Self-Supervised Learning Based Approach with A Transformer Target Model For MOS Pre- diction,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c88a3958-4fb3-4d70-9463-13037c7bb64e · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 009ebc16-5eae-417e-a12e-54ec420aee6e · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MOSNet: Deep learning based objective assessment for voice conver- sion,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d4e82e1-9659-4400-b67a-79047a6b82c5 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Robust speech recognition via large-scale weak supervi- sion,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab3abf6f-2794-4f18-af57-8b4cd7e9445c · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Qwen Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f12e455d-4998-41ec-b50a-b64df0ffd3f0 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Efficient optimal transport algorithm by accelerated gradient descent,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eceb604f-d319-4102-b005-320ad57c850d · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fb0be64-6805-4edd-bb43-a0989960304e · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cf873c1-3055-466d-8794-ed5457ae00cb · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Simple and controllable music generation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b117bfbd-7da0-466e-a6db-321892dbb1ee · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mo ˆusai: Efficient text-to-music diffusion models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 110b80d9-f855-4f71-a97d-7964703822bf · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musicmagus: Zero-shot text-to-music editing via diffu- sion models,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32b714ee-bd08-48cd-a7fa-e2067ef54485 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 948aa870-0d5d-40a3-8c25-d567d71de983 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcda9ac2-ff87-4e34-9b23-1215077a15cb · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mospc: Mos prediction based on pairwise comparison,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e78a8927-adc8-4697-9afa-7ca3b1b274bc · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Resource-efficient fine- tuning strategies for automatic mos prediction in text-to-speech for low- resource languages,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 622b45c8-a352-4575-9197-7d1cc3bbd279 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Zero- shot out-of-domain is no joke: Lessons learned in the voicemos 2023 challenge,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 83e39c97-ec97-486e-a77b-473dd56ec2ef · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e99131c-60e1-43be-b897-8b27fd569fc6 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44d951d0-8696-4fcf-921d-c7575a48ecc7 · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 258d12b3-f868-4a2d-b655-afd4cdcc71fe · outbound
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Cmot: Cross-modal mixup via optimal transport for speech translation,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.