Pith. sign in

Paper Citation Record · LEDGER

The AudioMOS Challenge 2025

As of 17 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 5 inbound Pith citation observations for arXiv:2509.01336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01336 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:41:28.171047Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:53:16.690620Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 71747265-4239-494d-9b26-0d3a57761db5 · outbound

This paper cites The V oiceMOS Challenge 2022,.

The AudioMOS Challenge 2025 The V oiceMOS Challenge 2022,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:24.747365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:24.747365Z digest=sha256:7e2fd95ba813c62b875dd82e38b0635e8e631d176b1aa9dd284b5d60900663e7

Observation 75c08112-cd56-493a-bcd9-95ed609c0cf4 · outbound

This paper cites The V oiceMOS Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple Domains,.

The AudioMOS Challenge 2025 The V oiceMOS Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple Domains,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.105692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:24.803002Z digest=sha256:b94919a1bc8298cad57c1b9db005a95799c8be6ff68b578a039b6e703827dbb8

Observation 9f15f70c-f59e-4775-97da-fec479570f35 · outbound

This paper cites The V oiceMOS Challenge 2024: Beyond Speech Quality Prediction,.

The AudioMOS Challenge 2025 The V oiceMOS Challenge 2024: Beyond Speech Quality Prediction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.090267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:24.838581Z digest=sha256:e79830a5ca2f07dec15ab9e5adaeef75e83f42e5681f417a3d0723413ea7601b

Observation bfbc1959-0a0c-4d00-8ae3-09eb3a872dc6 · outbound

This paper cites How do voices from past speech synthesis challenges compare today?.

The AudioMOS Challenge 2025 How do voices from past speech synthesis challenges compare today?

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.076034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:24.911721Z digest=sha256:060a59987076fe35bffd8ef6d989067b9dd38e606f69f45ad9a8e41d31172e99

Observation 9fee816d-8cca-43a4-89c3-636f263db927 · outbound

This paper cites Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,.

The AudioMOS Challenge 2025 Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.061380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.018687Z digest=sha256:83a9eb561712d25347e089821bb570564b52db305b643a9cb124bec00b21e074

Observation 7d788b0d-7c87-4918-9c0b-ce177cd587fc · outbound

This paper cites Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models,.

The AudioMOS Challenge 2025 Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.044903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.099194Z digest=sha256:827ec0647dc5b328d495a5ee8d67596a0a916e365038de8533764a3b1ccb9887

Observation 8cf02930-b725-4b01-84d8-8d5c8ef8c35a · outbound

This paper cites Evaluating generative audio systems and their metrics,.

The AudioMOS Challenge 2025 Evaluating generative audio systems and their metrics,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.029625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.176618Z digest=sha256:71c5acde410fdd30c920de08379013de20603424294d8c4b6051e7bd9d7906b5

Observation 444c3028-1fea-4c8b-ad67-517eb4e10b73 · outbound

This paper cites Correlation of Fr ´echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependent,.

The AudioMOS Challenge 2025 Correlation of Fr ´echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependent,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.014505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.264913Z digest=sha256:e6888d01aadf568a67a182e38054d670fd0a22bdf45649cb6e3b144bfad65f5a

Observation afbd36a3-ac50-4275-8770-a90feaf8348b · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

The AudioMOS Challenge 2025 Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:25.350124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:25.350124Z digest=sha256:025184e9d04283678028a4f622d4dc77951bc711e97915cdbf7b5747988835a6

Observation 076c667e-3f55-40e0-b9fd-ef547b50bbd3 · outbound

This paper cites MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation,.

The AudioMOS Challenge 2025 MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.995782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.426012Z digest=sha256:baa0359636159efc2d782ce079c2125bb0ab664f8234aaa491768c40d6ca603b

Observation 8a45dfc5-e190-4e62-8d36-03f44096efd7 · outbound

This paper cites LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning,.

The AudioMOS Challenge 2025 LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.980192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.499732Z digest=sha256:b79a00eb10f9e0532eff33969736ae8bdd427fb4dca6e629bc63b544daa793f8

Observation 3c74dd48-e844-46c3-a5f3-b34603c82a54 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

The AudioMOS Challenge 2025 AudioCaps: Generating captions for audios in the wild,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.964053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.614593Z digest=sha256:b652fa837c6aa2f9212ae8077b0566c0868d3933acf3b1fb999c25070e7f9942

Observation 6257b821-5018-4178-97e7-df8164efd9b9 · outbound

This paper cites MusicLM: Generating Music From Text.

The AudioMOS Challenge 2025 MusicLM: Generating Music From Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:25.791485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:25.791485Z digest=sha256:7d3c79a504754e9058d950371b1c22e3f3f4fd3d0e27f308a4c141dcd9800195

Observation 094207ec-0676-478b-b7b7-cc10dc9b3636 · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,.

The AudioMOS Challenge 2025 LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.948461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.897729Z digest=sha256:dfeecf5a9eaadc8696ca38ebebafc7eff03ebe23ef566be95e3e649f353a473b

Observation 8a3ff8ec-1567-4005-a81d-116bb2d934fc · outbound

This paper cites Hi-Fi-CAPTAIN: High-fidelity and high-capacity conversational speech synthesis corpus developed by NICT,.

The AudioMOS Challenge 2025 Hi-Fi-CAPTAIN: High-fidelity and high-capacity conversational speech synthesis corpus developed by NICT,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.932801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.973158Z digest=sha256:2addf72498390f8a2c6fd3f0c84d5af66264c0846b223b99a736c8584f8a30e6

Observation 35552078-7a96-431f-94bc-d8529290e268 · outbound

This paper cites World: a vocoder-based high-quality speech synthesis system for real-time applications,.

The AudioMOS Challenge 2025 World: a vocoder-based high-quality speech synthesis system for real-time applications,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.915871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.096740Z digest=sha256:47ef362b0ac90ff7cbe398cf3d3550d85ad5ecbe8f6ce1ac2848419a53725198

Observation c59f6e4b-80de-434f-b4ad-d173d91396c7 · outbound

This paper cites Fast Neural Speech Waveform Generative Models With Fully-Connected Layer-Based Upsampling,.

The AudioMOS Challenge 2025 Fast Neural Speech Waveform Generative Models With Fully-Connected Layer-Based Upsampling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.898082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.238332Z digest=sha256:8e54f47a1dd376ae968bf9fb447110642f44d2f30590d2bf78b0b4762be15aff

Observation 1716e873-33db-4514-9b46-f6a1c648661a · outbound

This paper cites Speech masking system based on spatially separated multiple TTS maskers with a compact circular loudspeaker array,.

The AudioMOS Challenge 2025 Speech masking system based on spatially separated multiple TTS maskers with a compact circular loudspeaker array,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.879302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.300292Z digest=sha256:902063d4b37340ab3df63b60feb65fde63eb716082131183f574335ffcfd1e48

Observation a4f16271-e27a-455a-bb07-aa84247e09b7 · outbound

This paper cites AudioSR: Versatile audio super-resolution at scale,.

The AudioMOS Challenge 2025 AudioSR: Versatile audio super-resolution at scale,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.860470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.410771Z digest=sha256:1a613b4211856f80f2c07d8504114ed6c27a1c48f47554a1dabfae2840f127d0

Observation e0184f3c-dd9c-48ff-9dc4-9d0c847682df · outbound

This paper cites pyloudnorm: A simple yet flexible loudness meter in python,.

The AudioMOS Challenge 2025 pyloudnorm: A simple yet flexible loudness meter in python,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.844989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.520713Z digest=sha256:2562a52b7a4b74108d42c36b37a224c29dd408a137a4e8f963e83e6386454436

Observation 7056e9ba-8ce3-49e2-a2c5-c4e50924387a · outbound

This paper cites Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,.

The AudioMOS Challenge 2025 Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.814314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.671368Z digest=sha256:dc3b5b1b90a63cb8f8087b5b2a94d632f7a72f7070d9547e304b8abbc4748056

Observation ed998dd8-45e3-4d88-a9e8-13eed8541ef6 · outbound

This paper cites HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection,.

The AudioMOS Challenge 2025 HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.797594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.767113Z digest=sha256:c846d352e7f36e146f64ac808df2198f20afa5745427d58502bb636d81980bc5

Observation 0a6684ce-f2dc-4ee8-ad7d-d7ddcc3aa8c1 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

The AudioMOS Challenge 2025 Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:26.859550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:26.859550Z digest=sha256:14026cb51d273851eaf46de5cd628abd2827882c7ff1d3665ce9ede47bd12838

Observation 172d95ca-1a10-4e60-866c-eebf56c26229 · outbound

This paper cites Attention is all you need,.

The AudioMOS Challenge 2025 Attention is all you need,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:26.890826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:26.890826Z digest=sha256:82b94787e97b930c11b689ef34e2fe6baae771331f01272342f254abb49f6e77

Observation 17a12fbd-41de-4400-bc07-6938649d4dfb · outbound

This paper cites Layer Normalization.

The AudioMOS Challenge 2025 Layer Normalization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:26.917094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:26.917094Z digest=sha256:506c1151167bb18aacfd2b982b1be217bbf3eb5e41fa07cb121e673e71679340

Observation 9500da5a-ebd8-4855-b507-6863ad5ae1ad · outbound

This paper cites Gaussian Error Linear Units (GELUs).

The AudioMOS Challenge 2025 Gaussian Error Linear Units (GELUs)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:27.079074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:27.079074Z digest=sha256:501b21d6144cd840fa8457151906966714879436e241d031d1ca52fcb391ac1a

Observation a227ea61-bb4b-4bb2-ae57-34891ddc50c2 · outbound

This paper cites Generalization ability of MOS prediction networks,.

The AudioMOS Challenge 2025 Generalization ability of MOS prediction networks,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:27.140555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:27.140555Z digest=sha256:61d1544300e5ee082a289c5fd19bf59b1afc2aa369ae6b39d34ba632cd52ea35

Observation 3b1348ed-832e-45ef-9242-286d6553fe4c · outbound

This paper cites PAM: Prompting Audio-Language Models for Audio Quality Assessment,.

The AudioMOS Challenge 2025 PAM: Prompting Audio-Language Models for Audio Quality Assessment,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.748591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.250841Z digest=sha256:745ccade8a59ab3a82f761f189aa123a235ba7c6d7429f36f2c62cd984fe7f61

Observation 818ff978-f69d-4b09-9641-8cef1b9b56b2 · outbound

This paper cites EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,.

The AudioMOS Challenge 2025 EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.733512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.363762Z digest=sha256:15b682e7d269b01cab63d60dd6725663e2ce0bb22622ad0ce4cf74884d3cad2b

Observation 4cb3aca0-b399-434a-b226-15d309c40397 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

The AudioMOS Challenge 2025 Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:27.466107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:27.466107Z digest=sha256:985e9f6e9a78c43ce5f6db8dd6d47dc00cd8fe36acf7b406d8f6b5b353c8e724

Observation 26bc3c32-a67e-4203-96cb-347d0b747098 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers,.

The AudioMOS Challenge 2025 BEATs: Audio Pre-Training with Acoustic Tokenizers,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.718092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.569724Z digest=sha256:a9125e0e7ca5fbb9fe08d3c9e9a42adc48968b95de250b894cb73a39d2bc0c1b

Observation 67cf87db-4185-409f-a2b8-e18bf6b0b21e · outbound

This paper cites Masked Modeling Duo: Towards a Universal Audio Pre-training Frame- work,.

The AudioMOS Challenge 2025 Masked Modeling Duo: Towards a Universal Audio Pre-training Frame- work,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.703862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.684183Z digest=sha256:7b4b02dcb70534a58cc53917ad20bc8aaa5a9929c4122a810a79b20630d820f6

Observation 7eee63f3-1fd0-41d5-aff8-9dea27ca73bf · outbound

This paper cites High Fidelity Neural Audio Compression,.

The AudioMOS Challenge 2025 High Fidelity Neural Audio Compression,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.685458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.738336Z digest=sha256:ed4b284edd6c2f1c82b8aac8361995a5c93555878c203d663007e2e795ffe4e4

Observation fbd19bd8-0b75-4f1e-8863-3654bf650e33 · outbound

This paper cites Scaling up masked audio encoder learning for general audio classification,.

The AudioMOS Challenge 2025 Scaling up masked audio encoder learning for general audio classification,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.669001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.798487Z digest=sha256:8bc444f62966014f32c8ef863a57d2ce966357ea9debd4aaa33cc9500b59723e

Observation 2c3b9206-cf74-418d-a760-716d816c152c · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,.

The AudioMOS Challenge 2025 MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.653507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.960529Z digest=sha256:fc9cd8ca80cf6f3de6d6239416fc7342d609e8768a2bc30eb18b6cac3e297164

Observation 9ea7b282-98af-4e89-a446-fe71f12d221b · outbound

This paper cites MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization.

The AudioMOS Challenge 2025 MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:28.086037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:28.086037Z digest=sha256:438c1e0b1c51ae7704e1bab923fb5d897119bc7532d8ea1216acf20ecfc171e1

Observation 1808c671-6225-400b-8cb7-30eecc0d2c3a · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding,.

The AudioMOS Challenge 2025 BERT: Pre-training of deep bidirectional transformers for language understanding,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.638751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.091853Z digest=sha256:3871a0b165c17cd93bd2e2fbfc3d945a49b1b12b86bb103afc00f8da39ede387

Observation 371f0a5f-9676-4638-98b1-7420db71a21b · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

The AudioMOS Challenge 2025 RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:28.096665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:28.096665Z digest=sha256:2dc6c45fa694e695278611c7896a8a02adc71989edd7ba72197455625f62e6f8

Observation 717265ef-56f9-4984-845e-912e8ccf3430 · outbound

This paper cites Qwen3 Technical Report.

The AudioMOS Challenge 2025 Qwen3 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:28.101543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:28.101543Z digest=sha256:db3d6fab90e74a37841d53ba835d581cd154c4b3cc100bcbe818b2bb9f2f85de

Observation decdce1d-cc57-4be4-9034-0b74cec2acf4 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Super- vision,.

The AudioMOS Challenge 2025 Robust Speech Recognition via Large-Scale Weak Super- vision,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.624216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.106411Z digest=sha256:36a6d16d279b8c670aaa19b619c4bbe7fc28de6e1ffeb19c00abeb6e40ffb447

Observation 242f1acb-fb3b-437f-ac5d-02f21b625cce · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,.

The AudioMOS Challenge 2025 HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.608435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.111341Z digest=sha256:9a6d42565fab0308df60b51c0d3f11b8ac515c781543273528a284b0fe74047e

Observation 44ff0cea-57f8-4d50-ab61-f525cc04c564 · outbound

This paper cites Scaling Speech Technology to 1,000+ Languages,.

The AudioMOS Challenge 2025 Scaling Speech Technology to 1,000+ Languages,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.590327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.115839Z digest=sha256:b4b950490f737d49f6c14fd0baa90253b81284258afa5452f2d5a23fcc9f3278

Observation 30883e1f-8453-4819-b884-c147b9987d9a · outbound

This paper cites EAT: Self- Supervised Pre-Training with Efficient Audio Transformer,.

The AudioMOS Challenge 2025 EAT: Self- Supervised Pre-Training with Efficient Audio Transformer,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.571687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.121108Z digest=sha256:2593ceff5d60cb7728619eb8851b0666f0c28b54cd20bcbf408b6a50156b01f5

Observation 66f7ec39-c620-4676-9128-c3b5ad480448 · outbound

This paper cites LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,.

The AudioMOS Challenge 2025 LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.549925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.126326Z digest=sha256:222fe39bd5fb67d332ffe79c7236cb15fc4aa46cd9cf20c3c0c077b7125bf600

Observation a147eebb-9fd7-4379-9024-91c8fc04aa03 · outbound

This paper cites KAN: Kolmogorov–Arnold networks,.

The AudioMOS Challenge 2025 KAN: Kolmogorov–Arnold networks,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.527216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.132640Z digest=sha256:01d5831e6120258e021a07d6db9cf2caee549701f92641facde9c0aa52c26626

Observation 9688a5e9-93bd-485b-a80f-a456901d0779 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality,.

The AudioMOS Challenge 2025 Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.505271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.138856Z digest=sha256:961708d3c8981aefdfd326cdec6b9b67c8239e70926f17e602f81fc8f5a09f67

Observation 5555346e-949e-49d5-b1bb-a279437c8bdf · outbound

This paper cites Sampling- Frequency-Independent Convolutional Layer and its Application to Audio Source Separation,.

The AudioMOS Challenge 2025 Sampling- Frequency-Independent Convolutional Layer and its Application to Audio Source Separation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.484447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.144515Z digest=sha256:6c0fd1dab0df7297008d3275f6e37f5b554968408f9e5e7619e13f27b0cba9f1

Observation 68362bec-02a2-44b1-b7d7-fc29fbe516ec · outbound

This paper cites Kolmogorov-Arnold Transformer.

The AudioMOS Challenge 2025 Kolmogorov-Arnold Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:28.151523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:28.151523Z digest=sha256:f30d552d03507dde6c5c9fb6a38327260a8359036cf7fb8e159dc62251088090

Observation a58e257f-295f-4f00-9adf-156dd1676430 · outbound

This paper cites VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music,.

The AudioMOS Challenge 2025 VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.457972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.157430Z digest=sha256:d03dea3ca11120ed2692d2ffa8784c6543de39eb60b4b18d2c4cb92c84583a8b

Observation fbbb3225-c21b-4a2a-afb7-d2c1f8254618 · outbound

This paper cites XGBoost: A Scalable Tree Boosting Sys- tem,.

The AudioMOS Challenge 2025 XGBoost: A Scalable Tree Boosting Sys- tem,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.436635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.165507Z digest=sha256:13214120054e12ee2b7ffef0276e6ab9df9fa74a495ffec9b92f4dc9e508dae9

Observation 37341450-6e5d-4fa2-9674-b6724494b15c · outbound

This paper cites Pseudo Label Is Better Than Human Label,.

The AudioMOS Challenge 2025 Pseudo Label Is Better Than Human Label,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.419805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.171047Z digest=sha256:e8d5f351d60180d7768bd342d945b43977daf2da086f1e9e3016b9b19631789c

Observation 635d2866-4a97-49a0-9603-f6d928cb71e9 · outbound

This paper cites an unresolved cited work.

The AudioMOS Challenge 2025 Unresolved cited work

Reference 150

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:41:28.829801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.598556Z digest=sha256:c88568b2d6c9c95556875cc86a09153b7e201c3c8fefa392035b320040c925df

Pith citing papers

Observation 08935185-338f-4cbd-a06d-1242a073ac96 · inbound

Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling cites this paper.

Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling The AudioMOS Challenge 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T09:53:16.690620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:53:16.690620Z digest=sha256:e342d966bed81364ebcf481b6bdd4b303e6a4effecc40e4f6c47dfd72c42a023

Observation e76adabc-464e-4244-bdcb-bf4f0b2cb540 · inbound

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions cites this paper.

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions The AudioMOS Challenge 2025

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:01:08.480283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T16:43:33.397158Z digest=sha256:a61b2856848170f6598c94283f5e89ab86ee1c731cc9a37d65e41547ba30e7fe

Observation a724a6a2-505c-4287-aab0-d21125a9daa1 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents The AudioMOS Challenge 2025

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.494952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:9c424a663dd9f38efb6c62073d3743d37e09b7c11c2e4dc02a7f6a294fb1d1f9

Observation 92cc6aa3-c70b-470b-8915-a73b1bfc21c0 · inbound

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation cites this paper.

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation The AudioMOS Challenge 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T13:57:36.846479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:57:36.846479Z digest=sha256:7eff5e4ac3605d8782ca973adf15a712259aeb4dc6b93b628a5ee7780bfdf5cb

Observation e086a207-662d-4702-8c06-a55293b673a3 · inbound

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves cites this paper.

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves The AudioMOS Challenge 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T20:33:36.029311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:33:36.029311Z digest=sha256:a26be5f7c7cb48fe6244292ac34a9553c1fb78db1c32753f93ccebf3633585b3