Pith. sign in

Paper Citation Record · LEDGER

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction

As of 12 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2508.21407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21407 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:22:49.287604Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84b03811-5093-4eb6-94a5-2226e005abfb · outbound

This paper cites MusicLDM: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction MusicLDM: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:53.413254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:46.856069Z digest=sha256:3372e20e2eba892f36cff98675dfda8b43f480491ae521484859102bf2060d38

Observation 28fd8169-c73d-4beb-9aae-2fbf6d893ef6 · outbound

This paper cites Masked audio generation using a single non- autoregressive transformer,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Masked audio generation using a single non- autoregressive transformer,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:53.234220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:46.931830Z digest=sha256:f4d944e394f048df5d5e9ba9db2bc06d8f82efd46298433a8084028d67c7572c

Observation 986fe835-d4af-4a6d-886d-bc6c77efe5ff · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:47.001462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:47.001462Z digest=sha256:34d498c95bf205ce8fb95a1432a9001da0976cbb0a51ff100706b2a38cbbd0e1

Observation c30e0ee2-aebe-45ec-8876-4811e67e03ed · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:47.082941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:47.082941Z digest=sha256:255b5c63b12d4d2740e79eca6a628f34b602d4e2bf2d058681df57583131998b

Observation fa93170d-47f3-4b1d-8bac-a0df666e2b12 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:53.073409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:47.171359Z digest=sha256:5e27de28656ba6f5681075c9318651648137b18e1a12710ae03d45e8c059d2a4

Observation d4b63afb-1f56-45d2-9a00-ba91da03b10e · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:47.246835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:47.246835Z digest=sha256:ae63917f20369fcc1946b1c7fb92c0b883d0c50fa8d847061b3f31d14c8d5002

Observation 6b61b1bd-8399-4eb1-9558-8f83c2327b46 · outbound

This paper cites Deep neural network embeddings for text-independent speaker verification,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Deep neural network embeddings for text-independent speaker verification,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:52.913364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:47.311643Z digest=sha256:ed4b1a146db608f1bcf7cd2a180cdc8b92dd72625a811258c35b05f735157114

Observation 80069db0-5db2-4642-8f8e-ed3b6fa9a8f6 · outbound

This paper cites Attentive statistics pooling for deep speaker embedding,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Attentive statistics pooling for deep speaker embedding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:52.770044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:47.394547Z digest=sha256:babf91f68f398f800513a4a915bd1d570dd1724d5405d352e6272eb7308380f3

Observation 303bc5b7-7073-4533-a08a-664be24aaa33 · outbound

This paper cites Attention is all you need,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Attention is all you need,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:52.649181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:47.479909Z digest=sha256:5a6ae1647dcffc294d20ced546daa0e41dfedb150e710628519b8768f3b53d45

Observation b2407de8-6d14-4c26-b8a1-0bc9c1ec0403 · outbound

This paper cites ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:52.471924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:47.586839Z digest=sha256:0e933a74da82de00cffb9a91d89272c96bd1c78b0423c0f7a9b119c54c70a237

Observation 30aa2119-5b56-43ca-b1f6-497c92a57ab1 · outbound

This paper cites CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:47.686311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:47.686311Z digest=sha256:c7ac804d3e213e4f940c38f65d1e447b55074b7f64e8e2f4ca7faecdbe1abb5d

Observation 9e2ce2ba-c6cd-4428-b5ee-e207f9ec895e · outbound

This paper cites V oiceID loss: Speech enhancement for speaker verification,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction V oiceID loss: Speech enhancement for speaker verification,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:52.357835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:47.750265Z digest=sha256:3dd18b90bd4797ef95df3ae46c3e71ff7c4a30da856ad5e005d17c52688027f7

Observation 9cefc8d0-d6c0-4304-a551-947839580113 · outbound

This paper cites Improved RawNet with Feature Map Scaling for Text-independent Speaker Verification using Raw Waveforms.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Improved RawNet with Feature Map Scaling for Text-independent Speaker Verification using Raw Waveforms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:47.829849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:47.829849Z digest=sha256:c173d3ba823756c540aa581e7a039742aefe0234562cae9c459920959e9f36fa

Observation c4da1a5b-0efb-4cc3-9ec1-50344dbe91f4 · outbound

This paper cites ResSKNet-SSDP: Effective and light end-to-end architecture for speaker recognition,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction ResSKNet-SSDP: Effective and light end-to-end architecture for speaker recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:52.213291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:47.905766Z digest=sha256:e4b209757b00ac091fcd6edf4d62b779b9913269eb87269c0d8e619a0941ece3

Observation 9e539c80-9abc-4f77-8381-0f6126dfc92c · outbound

This paper cites Deep segment attentive embedding for duration robust speaker verification,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Deep segment attentive embedding for duration robust speaker verification,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:52.075819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:47.970305Z digest=sha256:e5cc42a8b5e98a2ca82e7ba592aaff1159f87ba99e29337cc46fb5ab5e5508ed

Observation 1010ba86-8bc0-4751-b77d-2d4056acebb4 · outbound

This paper cites MusicEval: A generative music dataset with expert ratings for automatic text-to-music evaluation,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction MusicEval: A generative music dataset with expert ratings for automatic text-to-music evaluation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:51.919927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.049496Z digest=sha256:53e9c1837df897bc3fa250434cd456840b211166a36814dab2c0db111aee8927

Observation e799f815-b3ba-4efe-8770-aab7ca3e8a79 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:48.102873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:48.102873Z digest=sha256:790409c471b512a151f598d58b1ede84fd5a46489249d262afa958ae29164dc6

Observation f340da02-f107-4fcf-9261-29ea9e2f197d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:51.775500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.111551Z digest=sha256:81420f51295fb80335f45baf01289ccc0e49c7faa4ea18e6d272950e12c0e5a3

Observation 11f20ea0-7f46-4042-b0ba-f56291dc6e67 · outbound

This paper cites Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:51.644718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.201404Z digest=sha256:862267b98320ee17b35bd364a44485d2b7be664381f9e018f3be014c47080e9f

Observation c3a3fc5f-e16b-4464-9373-4d58ee843195 · outbound

This paper cites Self-attentive speaker embeddings for text-independent speaker verification,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Self-attentive speaker embeddings for text-independent speaker verification,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:51.488658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.310128Z digest=sha256:f4679f406c51dd97583fbd8132b93f25cc402d771ed959814ca7750fc430e494

Observation 76aa0763-e766-4222-910f-2041090107cd · outbound

This paper cites Self multi-head attention for speaker recognition,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Self multi-head attention for speaker recognition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:51.347918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.355394Z digest=sha256:c4aa39f59204b3ae34424b24e7acd85c03cbb6bf094572fc4334ef552966c9e0

Observation 69d5331b-0e83-447e-af62-82ef3fb87abc · outbound

This paper cites Vector-based attentive pooling for text-independent speaker verification,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Vector-based attentive pooling for text-independent speaker verification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:51.109082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.460106Z digest=sha256:3bd3cb9897fa0385a9b909476753c41c19dfab72c3a5c64518d827d6f1faa804

Observation 34c9633a-d80e-47df-b619-11db0f25d735 · outbound

This paper cites RSKNet-MTSP: Effective and portable deep architecture for speaker verification,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction RSKNet-MTSP: Effective and portable deep architecture for speaker verification,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:50.923273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.617793Z digest=sha256:b90967f86c62b58b706a860786c655cad3d21b28102b338ed2417bc21a374716

Observation d08f0789-eb84-4b38-b35c-882de7af8221 · outbound

This paper cites Multi-resolution multi-head attention in deep speaker embedding,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Multi-resolution multi-head attention in deep speaker embedding,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:50.807057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.699514Z digest=sha256:6ae22f3e49c07d05920e80d29ac4bbfa1987082871d576b42558e0b5a7bd110b

Observation 8383e420-07fa-492e-bb4d-084a7e2678db · outbound

This paper cites NetVLAD: CNN architecture for weakly supervised place recognition,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction NetVLAD: CNN architecture for weakly supervised place recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:50.587804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.820549Z digest=sha256:44e53c2870b0173fb58dca49c9f245be35b51f0ff76cdac608eddb070df49868

Observation 9e220270-3237-4e30-9fd1-53948ad89b52 · outbound

This paper cites GhostVLAD for set- based face recognition,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction GhostVLAD for set- based face recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:50.391939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:48.967002Z digest=sha256:1e49f923d121986387748eef40949ae9135f82917acc5830fefee634c36f0c27

Observation 82e405d1-2ce7-47ba-ba52-1461df4b3428 · outbound

This paper cites Neural aggregation network for video face recognition,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Neural aggregation network for video face recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:50.124959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:49.070316Z digest=sha256:0bf7f43b458809e8ad76370f958e41c76e8e667bff0d749405718428cdef75bf

Observation 17410fc8-c5de-4c95-adcd-a3307f759b23 · outbound

This paper cites Quality aware network for set to set recognition,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Quality aware network for set to set recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:49.885954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:49.168400Z digest=sha256:e3b7da0baade2c1659b4bdbfeac01f3931f38fa7bdd8dfb07c8bb8bba00064ac

Observation ad12a7fa-8fc7-4893-b1ca-71140829cde9 · outbound

This paper cites Generalization ability of MOS prediction networks,.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction Generalization ability of MOS prediction networks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:22:49.597372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-05T14:22:49.287604Z digest=sha256:528fb91bfc0594bd7126fe475a2eeaa19a30439ec2de89fc65dc8770f02d9430

Pith citing papers

No inbound Pith citation observations are available.