Pith. sign in

Paper Citation Record · LEDGER

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.02082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02082 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:44:17.796045Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact4
  • verified fuzzy10
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02691a0c-79b2-460b-9721-f69bf7e95692 · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.618021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.618021Z digest=sha256:71244d37212dcda36a1d6139cfceebdb73336acc353e8a27f43610faf5ca12a3

Observation c2e23147-c707-4322-a1db-88b238e4d96c · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.625066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.625066Z digest=sha256:7733a5482fa1809c37756da10ed1d07dfe7627938e35b727741790f5dedb0867

Observation 19fe8f58-2378-4a75-80be-20807df4b347 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.631566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.631566Z digest=sha256:b78d9f66b5aedd9a2509aa8eda977d0a5dd9a4c99fb5b06a696fa98050021007

Observation 0ea38d4e-bb61-43e4-974a-8cdc59922e95 · outbound

This paper cites TERA: Self-supervised learning of transformer encoder representation for speech,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction TERA: Self-supervised learning of transformer encoder representation for speech,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.390260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.638024Z digest=sha256:51d8991f26c51d41428c9d8ca4d9962680da27c4c5b94c6084dfa58ba650de8a

Observation a1983489-b854-4836-aa1b-fdb7682379f2 · outbound

This paper cites AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.644998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.644998Z digest=sha256:8e2fd60be4003abc8e61e69fd44cadcc16f682e8985269b373f13df535b6ad44

Observation 888f7599-bd55-4125-84bc-a9f43f7a0964 · outbound

This paper cites Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.651602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.651602Z digest=sha256:b24242031f0aec78406c8baa8c141c1e02f4cafe1a4353ddaae732182edd74b8

Observation d57acf9b-c798-4cf0-9b6a-35142091227b · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.365044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.659966Z digest=sha256:3aae3a1ac06893f58e2ee4ad96314d0c22c84dd01679fef5544f69e5de990952

Observation b02c79b6-39e8-437a-9d8e-d2bff1aaa8b0 · outbound

This paper cites Evaluation of speech representations for mos prediction,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Evaluation of speech representations for mos prediction,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.346302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.667199Z digest=sha256:f85a53fcf91d1f11471a81e1d7732cf73e4c998486cb01109ee4a8be1f01d5b3

Observation 9148ec3c-421a-43bd-8625-5cc092c63f54 · outbound

This paper cites Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.324451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.672517Z digest=sha256:8737633ea5dd07e62ae0ab4ec6b7959a2766d8bae39bc9ede1da0f6b8d3dab65

Observation ed080602-d705-4148-958f-1e36bc095878 · outbound

This paper cites MOSNet: Deep Learning based Objective Assessment for Voice Conversion.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction MOSNet: Deep Learning based Objective Assessment for Voice Conversion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.680497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.680497Z digest=sha256:cafa493641ba719ade916baf70628140cd754aad49481853ce65520e2e781af4

Observation dab1aff8-dbb6-417d-aaa0-f98a13c6d6fc · outbound

This paper cites MBNET: Mos prediction for synthesized speech with mean-bias network,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction MBNET: Mos prediction for synthesized speech with mean-bias network,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.302247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.688319Z digest=sha256:1cf6f3e11dad9702fb46d8c074ed47ea9cc4b2981971360d4b1a0b5b0a006312

Observation f1bc8676-32ed-4d9a-ab5e-8020b98bab0c · outbound

This paper cites LDNET: Unified listener dependent modeling in mos prediction for synthetic speech,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction LDNET: Unified listener dependent modeling in mos prediction for synthetic speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.274372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.695218Z digest=sha256:5e9253eda3519759e01880d0fbdc05b09efa5d03bed71fb29afed1d8bd750c1c

Observation 367dae72-feec-411a-8996-461f21409134 · outbound

This paper cites DDOS: A MOS Prediction Framework utilizing Domain Adaptive Pre-training and Distribution of Opinion Scores.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction DDOS: A MOS Prediction Framework utilizing Domain Adaptive Pre-training and Distribution of Opinion Scores

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:18.048488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.702839Z digest=sha256:8146e0526c8ee7d46a20244402db6562d1c41ff2568ad5a7f395b1f7f2f178b5

Observation a7e43f76-0e26-4106-a15a-d284420f2959 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.709392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.709392Z digest=sha256:907627595b34dfca9241084b303760dd6864dd4ce59cc901ced4ab85e23ca749

Observation b8b67440-aadc-440a-adc2-e0d77a008300 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.715160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.715160Z digest=sha256:9096f48be3e0ac1f2f9053d8c7dd88049da0bef0b98f8e6dc36c0bf3297f8baf

Observation b5d7a4de-583e-489f-a575-3c0c73ee414a · outbound

This paper cites Fusion of Self-supervised Learned Models for MOS Prediction.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Fusion of Self-supervised Learned Models for MOS Prediction

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:18.004321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.720703Z digest=sha256:674d58a39d09fff63a0bed2a57535341c4227433353e0a386c3b6c3f42be6c19

Observation 41d06764-3c28-4b9a-9229-41f4a80baf47 · outbound

This paper cites MOSPC: MOS Prediction Based on Pairwise Comparison.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction MOSPC: MOS Prediction Based on Pairwise Comparison

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.727766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.727766Z digest=sha256:2c1c96fcce37208b2c238cf814a1fc4cb20dcccaf13ddeb7cc27be63074338b7

Observation 220b25bd-0434-4023-9b39-cf319156dd00 · outbound

This paper cites Speech Quality Assessment through MOS using Non-Matching References.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Speech Quality Assessment through MOS using Non-Matching References

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:17.946189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.734793Z digest=sha256:e68d4f1a2ab1c6a621d5e63c5a1ff30c362aa6f361d9df78b49685b88e5662be

Observation 2dfbb2a6-1e08-49e7-a198-02301099feb2 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction U-net: Convolutional networks for biomedical image segmentation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.741382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.741382Z digest=sha256:769449cd26b9c226e9f25cf34d25b3d1dd56c1781fcc6c6ac406843abf543ab9

Observation e1f91c3c-ccd9-419b-9590-0feaf4e98c56 · outbound

This paper cites Perceptual objective listening quality assess- ment (POLQA), the third generation itu-t standard for end-to-end speech quality measurement part i—temporal alignment,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Perceptual objective listening quality assess- ment (POLQA), the third generation itu-t standard for end-to-end speech quality measurement part i—temporal alignment,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.223846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.748440Z digest=sha256:e31d0a837561c63a416f95b0f0b9437867a70f2b713ada5ba8313faa4544a9f6

Observation 52aae80b-685e-455a-9b2f-bb90c42a645f · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.199291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.755047Z digest=sha256:a2899133d5ccad6f740876d84255a1607d931de2b4325a5adff71fde69a02030

Observation b03e58a3-6fd4-4615-98b8-157532293267 · outbound

This paper cites The VoiceMOS Challenge 2022.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction The VoiceMOS Challenge 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.760799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.760799Z digest=sha256:1c49ecccb128d15555ab4ca8fc36c08b64757f7aeefe6adc365cfc89465ff4d4

Observation 69057cac-d56f-43b8-88c7-20aa16930c5c · outbound

This paper cites The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.767102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.767102Z digest=sha256:fcb12ccff32acb3e5c8b5e13d9928fe5cf1c6ea9358737705dce13dcf1f9d41d

Observation 85f38dec-661a-4107-b374-415b0001405f · outbound

This paper cites SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.773360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.773360Z digest=sha256:fdd49a19dcfb5ed6b6849a216cb67b73f0178a8380d51b3fc4a9b7ff189ef6a5

Observation a6b79536-61c2-4870-96f6-7dc5100f0201 · outbound

This paper cites InQSS: a speech intelligibility and quality assessment model using a multi-task learning network.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction InQSS: a speech intelligibility and quality assessment model using a multi-task learning network

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:17.855886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.783207Z digest=sha256:15ed1bbf5c07c819a1d58213c56aa78c7c92604ea18354fa56c62cc4848595d9

Observation b16f79fd-15d8-4d3e-8122-3f08a56ba9a7 · outbound

This paper cites LE-SSL-MOS: Self-supervised learning mos prediction with listener enhancement,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction LE-SSL-MOS: Self-supervised learning mos prediction with listener enhancement,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.180511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.790700Z digest=sha256:b61056239fcb8afc61cf557675d31e953c249fa3403513ace399032049b33196

Observation 065f0d59-b72a-410d-8802-f5139d7fc6ee · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction X-vectors: Robust dnn embeddings for speaker recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.158851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:44:17.796045Z digest=sha256:45f13be6576296f0aa27dc0679a4f8432c1d002508bd7a64cd8357b18410e3f3

Pith citing papers

No inbound Pith citation observations are available.