Pith. sign in

Paper Citation Record · LEDGER

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2411.11232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11232 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:50:08.405900Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:50:08.261278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T18:50:08.499324Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy28
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a10d66e1-013c-4a02-a300-2a91e0e2666b · outbound

This paper cites Common ob- jective evaluation metrics, such as mel-cepstral distance (MCD).

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Common ob- jective evaluation metrics, such as mel-cepstral distance (MCD)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.906628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.253206Z digest=sha256:6186138cf02365236c57cc99ccaae36efd5c67f7501b2e0a49dfbde6c6f37fb3

Observation 310222d8-cc72-43e2-8675-17cf92c5aab0 · outbound

This paper cites As a result, some ob- jective measures or models related to human perception have been proposed [3, 4, 5, 6].

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features As a result, some ob- jective measures or models related to human perception have been proposed [3, 4, 5, 6]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.893734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.257531Z digest=sha256:e74dff3d126a04e987cd2719f90ac3df0a422e9daa212ff470995df59fa242f8

Observation e87ed4fb-8977-4686-88e2-04c264b79dd2 · outbound

This paper cites Dataset In this paper, the experiments followed the same settings as the V oiceMOS Challenge 2022 [15].

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Dataset In this paper, the experiments followed the same settings as the V oiceMOS Challenge 2022 [15]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.857226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.273361Z digest=sha256:46d0507ff5860a810ec256ae477af86a11ad39bf173a4041cc9603f5c0d68eb8

Observation 780f391c-0607-4a61-aa33-3478aefec9fb · outbound

This paper cites mean-listener.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features mean-listener

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.881238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.265344Z digest=sha256:2bf6f428f5144e3972db98d07fb1baaffbe6fbe1381ae2e0910d4641f1975eee

Observation ee4f1d43-f25c-4638-bcb1-238c065bba6e · outbound

This paper cites When the rater ID is not the mean-listener, the label representing the sample is the score given by the individual rater (an integer i from 1 to 5).

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features When the rater ID is not the mean-listener, the label representing the sample is the score given by the individual rater (an integer i from 1 to 5)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.869423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.269386Z digest=sha256:f1835b754d9634ed96dc24710758f06895fc05fba73b2df477baf1a93ab525ad

Observation 1f5c6476-c622-4219-8e09-7e6423a05b1d · outbound

This paper cites NISQA: A deep CNN-self-attention model for multidimensional speech quality prediction with crowdsourced datasets,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features NISQA: A deep CNN-self-attention model for multidimensional speech quality prediction with crowdsourced datasets,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.762518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.308516Z digest=sha256:4799f181c11dd35bc097ec61a381c10adfd00ea3297517b721aae8e978a1a2b4

Observation 67bc5e51-1fd6-407d-bb72-5fb44d2d6e40 · outbound

This paper cites Comparision with baseline methods We first compare the proposed SAMOS with the baselines.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Comparision with baseline methods We first compare the proposed SAMOS with the baselines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.846329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.277306Z digest=sha256:b09ab1b3321f0ef5c37262b6c34905971ffe1a307823f4b74328e41e11c978e8

Observation 41c32c20-0186-47f1-a6b0-8cb3bec56c93 · outbound

This paper cites We can see that removing the semantic module resulted in the degradation of all the metrics on both datasets, indicating the importance of semantic repre- sentations from SSL model.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features We can see that removing the semantic module resulted in the degradation of all the metrics on both datasets, indicating the importance of semantic repre- sentations from SSL model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.834338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.281046Z digest=sha256:9a28ff8c8069db4192f1e49089d66b92e770ea68a696ff3df6ee610a5d9d75f4

Observation cdec2479-b405-4bf8-85b1-1d9413048786 · outbound

This paper cites To improve prediction accuracy, SAMOS employs parallel regression and classification heads, and finally outputs the final MOS score through an aggregation layer.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features To improve prediction accuracy, SAMOS employs parallel regression and classification heads, and finally outputs the final MOS score through an aggregation layer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.821712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.285155Z digest=sha256:1ba0172c66935779d11302a87d7e6ddc419409255e9183a4e9793e3ee197e452

Observation a4a7314c-03f6-4e6b-9fbc-0c1eacc56cf3 · outbound

This paper cites Mel-cepstral distance measure for objective speech quality assessment,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Mel-cepstral distance measure for objective speech quality assessment,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.808890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.288976Z digest=sha256:bfd6e1dcf3a8d444db91b60e2aa98ae248da47476ceed4a356c6cf22f1c597f5

Observation 9029bd3b-988a-4b5d-a2ce-96aa0b219744 · outbound

This paper cites SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:50:08.503803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.261278Z digest=sha256:ccf0c434b66743d3348f038d1dafd59924b565f6754712fb3a475a86c7d2837d

Observation 30193a09-e986-462f-9cfc-2ca600386acb · outbound

This paper cites SDR– half-baked or well done?.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features SDR– half-baked or well done?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.292704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.292704Z digest=sha256:0726b9d28ce3c61a16f47479acf31ba4ca36c7dae6582db4ca588d967704dbcf

Observation e07f6f7b-3792-4d31-b07a-1e0954f8979d · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.296480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.296480Z digest=sha256:83ca279b23297a3676042ee7700d3c92a96ec90f6882fd8ac47b2faca44511ed

Observation e44b9181-7668-4052-8766-de5ac55a6734 · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.300376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.300376Z digest=sha256:f5671d6d025c817f936466b128d3414a35971fd3e112084715f9abeb87e2ecc1

Observation f594f578-5842-4e96-9dde-2ef897188c1c · outbound

This paper cites ViSQOL: An objective speech quality model,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features ViSQOL: An objective speech quality model,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.774239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.304873Z digest=sha256:42f1a168d1fb8929df1c584b92901c8eeb305ddf368f876e85d596e9ec786308

Observation e699b29c-92a3-4aad-8d43-3aa357453f6f · outbound

This paper cites AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.312306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.312306Z digest=sha256:f93e87bbad2a93fdd571513022da056a3308970a5167204e8ba366b3d2b17409

Observation 7b82dcd7-8716-4fbe-8279-c3b28f26b027 · outbound

This paper cites Quality-Net: An end-to-end non-intrusive speech quality assessment model based on blstm,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Quality-Net: An end-to-end non-intrusive speech quality assessment model based on blstm,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.751072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.316397Z digest=sha256:cef15c407f0600077da0230fa9d6dc4955922c02ae9df76ba51040f33e5b1bb1

Observation 3e78365f-4946-4647-898c-9c39a313e118 · outbound

This paper cites MOSNet: Deep learning-based objec- tive assessment for voice conversion,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features MOSNet: Deep learning-based objec- tive assessment for voice conversion,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.320018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.320018Z digest=sha256:d0e7c77ce5702bad9f8f8a80beaff7a8053018bd17b701894360ebbbb2cd2435

Observation 3d77dae0-f23a-4757-844d-9fa4cf8e0cbe · outbound

This paper cites MBNet: MOS prediction for synthesized speech with mean-bias network,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features MBNet: MOS prediction for synthesized speech with mean-bias network,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.323526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.323526Z digest=sha256:845aee4ead3c90951c4c6b890d9bbe9b10fb289823f9d220437260addab2142b

Observation 794cebee-6b36-4bbb-b50d-89a86f18a8e7 · outbound

This paper cites LDNet: Unified listener dependent modeling in mos prediction for syn- thetic speech,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features LDNet: Unified listener dependent modeling in mos prediction for syn- thetic speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.724766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.326959Z digest=sha256:7c2a45a1f0238fbf82e12bd65a11b8cb50eb48a40a3b9d7880f3aa727dd12963

Observation 52e2e90f-d821-48e3-befb-f3f91db3124b · outbound

This paper cites Generaliza- tion ability of mos prediction networks,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Generaliza- tion ability of mos prediction networks,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.712368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.330731Z digest=sha256:e342eadd8404a8ed26c88d5a57f2ec6ed04882a71d416c2ed58cbc1dfd1eb743

Observation fb553b37-1660-4ac6-83a8-6a2253627619 · outbound

This paper cites Deep learning-based non-intrusive multi- objective speech assessment model with cross-domain features,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Deep learning-based non-intrusive multi- objective speech assessment model with cross-domain features,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.700005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.334617Z digest=sha256:899108171921523f607cb7b995824fc434e54e6c1287ef56eabb1bb0ab32f986

Observation 60af0c23-252a-4d64-97a0-0496ea529157 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.337864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.337864Z digest=sha256:46ec923753e23cf0b15ae6729f0dbef8c1d049709a6f48ea296c97fd5f5bac2d

Observation 54fe6ef4-8be8-4958-8218-b5dc94710d9f · outbound

This paper cites The V oiceMOS Challenge 2022,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features The V oiceMOS Challenge 2022,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.679439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.341380Z digest=sha256:752e3e6716ab26caa91a7a4ffbd504afaa394c5a0e5fba56a24a546678c0a71c

Observation 1c22bcfd-ac74-4282-8180-efc23733ff3e · outbound

This paper cites A transfer and multi-task learning based approach for MOS predic- tion,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features A transfer and multi-task learning based approach for MOS predic- tion,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.667934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.344950Z digest=sha256:d22fa86b59b779506a75f134a435719f638825f3c9c84c57839b73e38c5153d1

Observation 25448d39-b0e7-4b40-8abc-220ca5382bb6 · outbound

This paper cites DDOS: A MOS predic- tion framework utilizing domain adaptive pre-training and distri- bution of opinion scores,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features DDOS: A MOS predic- tion framework utilizing domain adaptive pre-training and distri- bution of opinion scores,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.656349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.348461Z digest=sha256:d6abeca63303b06dcec54c80341f918bbe3a4caff41d33403dd47af84899f42a

Observation b4ec5239-740e-4ec7-912c-06563030db64 · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.643694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.351975Z digest=sha256:069302b80dfba7607e3c22aee953ce3902408d8ea710afd02d0c6a95ec6305c6

Observation c7776e9b-87ab-4c7b-9e29-10bdc810d981 · outbound

This paper cites Fusion of self-supervised learned models for MOS pre- diction,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Fusion of self-supervised learned models for MOS pre- diction,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.631520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.355250Z digest=sha256:a58c1aecafe95904ae03fe9f4f77ba3df732aa405bb6bf0bbb59fb9040345e68

Observation dded1169-beef-48c3-a257-46a1ba2d1be0 · outbound

This paper cites Ensem- ble of deep neural network models for MOS prediction,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Ensem- ble of deep neural network models for MOS prediction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.618523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.359094Z digest=sha256:e2a30231d3a07d8f5defe46d88f698d40749579da5a570e631ee60c743fb6549

Observation 447a59fb-38f9-4672-a89c-98c72831c85d · outbound

This paper cites RAMP: Retrieval- augmented MOS prediction via confidence-based dynamic weighting,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features RAMP: Retrieval- augmented MOS prediction via confidence-based dynamic weighting,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.606030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.362859Z digest=sha256:4c024d5f0b4c5719bb1147a4d62096579872b7016bec4ff19a71c122e824a592

Observation 91352e5e-3eba-47d5-ab95-306699f19a0c · outbound

This paper cites Investigating content-aware neural text-to-speech MOS prediction using prosodic and linguistic fea- tures,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Investigating content-aware neural text-to-speech MOS prediction using prosodic and linguistic fea- tures,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.593262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.366284Z digest=sha256:860ee9197263f6f3010ea149e2b7d13a8ad1a7fd2a3fbe85c5a9057a74e49f7d

Observation 1fee44e5-7b9f-45fb-8672-4163a1ba579f · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.580675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.369808Z digest=sha256:989eb94f3bb3e652666fa53b0e7c9105d2fa5b1bdf0a1b5ff1f2968225367f71

Observation fffca9b7-87f0-4c05-9ff7-43ce8e77bfb5 · outbound

This paper cites BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:50:08.475207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.373433Z digest=sha256:de97385c6b134b88ba13f4597bff9b835ddb9156cce325978734a2c3ef521e27

Observation 1027c81d-02da-479f-ac32-dcac8a1dbf17 · outbound

This paper cites SQAT-LD: Speech quality assessment transformer utilizing listener depen- dent modeling for zero-shot out-of-domain MOS prediction,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features SQAT-LD: Speech quality assessment transformer utilizing listener depen- dent modeling for zero-shot out-of-domain MOS prediction,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.569287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.377396Z digest=sha256:f969f1bb112c4f7106869ec83b7f596946810214032a5050e074194734c42a6e

Observation 151e9e29-7922-43e7-982b-329479541b82 · outbound

This paper cites How do Voices from Past Speech Synthesis Challenges Compare Today?.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features How do Voices from Past Speech Synthesis Challenges Compare Today?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.381069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.381069Z digest=sha256:c946e9f57bdcf8af05f5b99ef61c9f524e4a8a28e5653b62c9bb55fb5a3ebf5d

Observation 02825d8f-0937-4a6c-8255-0eaafae92b5f · outbound

This paper cites The Blizzard Challenge 2019,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features The Blizzard Challenge 2019,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.385075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.385075Z digest=sha256:da1c6ebb0603a471952d79667592a4a2698bf8dc5409099202c69c710180cda5

Observation 3d4e407a-af34-485f-aa7f-04e59493193e · outbound

This paper cites Conformer: Convolution- augmented transformer for speech recognition,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Conformer: Convolution- augmented transformer for speech recognition,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.389039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.389039Z digest=sha256:a2020228671d1524c514e2f732244dbe195bebeebccc20b48bd8c0fc8d213d1c

Observation 0fd6d85c-a83c-4c24-98d7-879b107e970d · outbound

This paper cites ConvNeXt V2: Co-designing and scaling convnets with masked autoencoders,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features ConvNeXt V2: Co-designing and scaling convnets with masked autoencoders,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.543400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.393014Z digest=sha256:1d719667ec465925a72eaecd2fcddacaab5b819be8a5d89a259ffbfe3658a805

Observation 632c41a8-ad18-440c-8953-02c144402482 · outbound

This paper cites Improving Self-Supervised Learning-based MOS Prediction Networks.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Improving Self-Supervised Learning-based MOS Prediction Networks

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:50:08.445952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.397560Z digest=sha256:13711bf8eb1b5e62ddbec0c358204a6fc00cd6166e9dc153738453dff6cab48b

Observation 567d4e74-2d84-44c6-a656-ca38ed976c65 · outbound

This paper cites ESP- Net: End-to-end speech processing toolkit,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features ESP- Net: End-to-end speech processing toolkit,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.531013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.402117Z digest=sha256:fd228f90d6ee4206c732a730e39eda1be55716f94575ba02e169b6de42952250

Observation b41c0c1c-fce8-4f73-bbbc-cc8666f5f234 · outbound

This paper cites CSTR VCTK cor- pus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features CSTR VCTK cor- pus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.518311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.405900Z digest=sha256:6919ec6f78889884cc18908ff24fd66c4acee246d212732e268fd4e490fae161

Pith citing papers

Observation 9029bd3b-988a-4b5d-a2ce-96aa0b219744 · inbound

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features cites this paper.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:50:08.503803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T18:50:08.261278Z digest=sha256:ccf0c434b66743d3348f038d1dafd59924b565f6754712fb3a475a86c7d2837d