Pith. sign in

Paper Citation Record · LEDGER

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio

As of 23 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.23582.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23582 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:24.207276Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:21.333361Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T21:41:24.430826Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f56a6ed1-d90d-4f35-9f31-859ccba6066f · outbound

This paper cites a dog barking behind a human speech,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio a dog barking behind a human speech,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:30.509503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.279386Z digest=sha256:dd60686f9998f9ac4b75bdbeb1452324e8f3454615e66348ed31b1a6b34bd854

Observation e9e26b03-e55f-4363-926e-a8e6934b2c7c · outbound

This paper cites RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:24.575515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.333361Z digest=sha256:dcde2216317203cae4890a604f9282c94d4fce1a4ede7a66d659a17354113a37

Observation 5bceae05-baab-4c89-aa9f-cbc2d88d8296 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:21.405172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:21.405172Z digest=sha256:1f222e791472b9df7f542200b37cee43246dfc2a778b93b22dc13fbbabb30783

Observation 7312557b-8972-4ae4-99b9-a421dd381213 · outbound

This paper cites 2” or higher, we added the exclusion of listeners with an average original audio rating of “6.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio 2” or higher, we added the exclusion of listeners with an average original audio rating of “6

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:30.309407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.469080Z digest=sha256:d5a4056319ad685a04eefca4b9253a58a892f417185f1d77f801ab26dd0102db

Observation bca3e593-4e97-4146-829b-69f2ed23b05f · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:21.575339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:21.575339Z digest=sha256:f3fbc4a2681f8b93bce12cd68dcff0e9aa334224c4561a69c903e291f3f2ac5e

Observation 9bfb8ae9-087d-49b3-b354-e047c79026ea · outbound

This paper cites an unresolved cited work.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:29.964702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.623018Z digest=sha256:58dfdb16b8ef046b0836a7d3f0d3b0d5dc1b036863ce8c4a451bde824c942199

Observation 0854e656-819b-4fdc-8598-263f139e5827 · outbound

This paper cites an unresolved cited work.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:29.801710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.684278Z digest=sha256:0fd991f0be6979f0f35406c84d9877c91bea448d04f09e000e154636e42870f4

Observation b1d044a9-a83d-4535-aae2-d532992e074b · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.443852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.222125Z digest=sha256:a1f9e94212f6da20805b4d363a3d8307c6ad3643be32a864b5c27343420dded1

Observation b5648097-787c-462b-a79d-9917e23c3e0b · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound genera- tion,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Diffsound: Discrete diffusion model for text-to-sound genera- tion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:29.572302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.755408Z digest=sha256:af1422260528bf594ec307b4a398546d3aac67e25e4f5e7e957fa9b8b83868af

Observation 0c15c510-645e-4f29-a594-5ed720865bd9 · outbound

This paper cites Sound synthesis for impact sounds in video games,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Sound synthesis for impact sounds in video games,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:29.343842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.820308Z digest=sha256:4f0ac6d8234d8ad648532c56e732381494cfaa9c485a56ad37c6513fb8e8db31

Observation 230d1527-c23a-4766-8c98-9676f450937b · outbound

This paper cites Challenge on sound scene synthe- sis: Evaluating text-to-audio generation,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Challenge on sound scene synthe- sis: Evaluating text-to-audio generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:29.149522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.868292Z digest=sha256:9672050db4d80b4a53b43cdde9a49a9ba0dee8ad3a859c9c4518ed47b9fbe3f8

Observation 9d038414-0e20-426e-a83e-519bc03d6f4f · outbound

This paper cites The V oiceMOS Challenge 2024: Beyond speech quality prediction,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio The V oiceMOS Challenge 2024: Beyond speech quality prediction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.991430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.935696Z digest=sha256:3ce0a6e5d1531e61f1cf0eb44e86e4180b821851032a39f7cee282b61b0c2e91

Observation 575f9da5-92f0-4e6c-8278-bf9af17f6c91 · outbound

This paper cites Subjective-aligned dataset and metric for text-to-video quality assessment,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Subjective-aligned dataset and metric for text-to-video quality assessment,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.816240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.012364Z digest=sha256:d7b0d15ac68cf25fbc53d8e2569e9a219e95c530689c865b6975a30c7d088431

Observation 105ac452-f0d2-4310-a374-341099b3a6e2 · outbound

This paper cites Environmental sound synthesis from vocal imitations and sound event labels,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Environmental sound synthesis from vocal imitations and sound event labels,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.698899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.082880Z digest=sha256:f7136b73a8fa268d698124985fae748c2cae08806d87b344dc65c7a40cdd1042

Observation 9ccadc8a-8fb6-4579-992f-806bef2474af · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.578905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.143289Z digest=sha256:5905122c13d991efe81cdf50dd71e9df02499f1903298536a3e6f381c77939b3

Observation b0ebc53f-c036-4544-a474-bcd78621ca1b · outbound

This paper cites Text-to- audio generation using instruction guided latent diffusion model,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Text-to- audio generation using instruction guided latent diffusion model,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.266852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.788988Z digest=sha256:0143caaa678800d99d74ed4af0e186f561929a50a0736ddfc0eff0a214419d65

Observation 12b470e2-c5a9-43ce-904c-3825d3b472df · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:22.283957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:22.283957Z digest=sha256:becda9e8d520bfb5f4e093d4eab50cbc928e54ef19b619042f23fa23281cf126

Observation 7203e59d-5374-4215-b076-8c93d4ef1fb5 · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio WavLM: Large-scale self-supervised pre-training for full stack speech processing,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:22.352486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:22.352486Z digest=sha256:a7071d7d46f30d73d8964aea90ca07140e417d8b051e8c5b8c12de71066a7925

Observation 66c5ab75-e0cd-4bd2-af30-4519267f34f2 · outbound

This paper cites Animal” category shows both of statistically significant differ- ences and interaction. Figure 2(a). shows that synthesized audio in the “Animal.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Animal” category shows both of statistically significant differ- ences and interaction. Figure 2(a). shows that synthesized audio in the “Animal

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:30.166092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.515288Z digest=sha256:399e3d853c210982ba8eb2c0cca80f4e86474b9f23dd960606fb0b8beb454466

Observation 6beaa23e-2537-40ab-bcc8-1f182c8af643 · outbound

This paper cites PAM: Prompting audio-language models for audio quality assessment,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio PAM: Prompting audio-language models for audio quality assessment,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.252093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.421792Z digest=sha256:e587d62dfb5f2d78f115f375b3c2a612109a5e8aecf465a20f216b393a67d2a9

Observation 818e93a2-30ca-4461-b3a1-625b0cfab09d · outbound

This paper cites Audio-text mod- els do not yet leverage natural language,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Audio-text mod- els do not yet leverage natural language,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.017606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.493924Z digest=sha256:65186ccb11982ce30120d9a6b206d6225f4ce9a57b9798cdca0908645bd6cca0

Observation 1f895f04-a02e-432f-bfb2-63baf86708a3 · outbound

This paper cites AudioCaps: Generat- ing captions for audios in the wild,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio AudioCaps: Generat- ing captions for audios in the wild,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.791916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.559383Z digest=sha256:ebcefe42b69083222e888f8a75973dfb25f29a238a0c48d408306f342d6229f0

Observation 7dfe310c-937c-4c4f-a7fa-5adb96f9ac70 · outbound

This paper cites AudioLDM: text-to-audio generation with latent diffusion models,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio AudioLDM: text-to-audio generation with latent diffusion models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.637563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.634145Z digest=sha256:b906bb3df7173c2e8b968a49d4a52b0f77ef4d4621a94e153e863de3038a5761

Observation df6faa69-4fec-4869-bc10-86d9f8f9d286 · outbound

This paper cites AudioLDM 2: Learn- ing holistic audio generation with self-supervised pretraining,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio AudioLDM 2: Learn- ing holistic audio generation with self-supervised pretraining,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.472397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.715542Z digest=sha256:9eb2f9e6c466c3c35ee21b6386947e61e7550c9c2631cdd0e529a81660cc8f96

Observation dfe3687a-0663-47d7-aca1-27601b5f7a6b · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.012471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.873119Z digest=sha256:e393e7eccb516e477d992864e031f03ae84b78efd18ffa6d86f9528234644f8f

Observation 11b55f60-ec7d-4e44-80d3-aa24986f3f83 · outbound

This paper cites Toward verifiable and repro- ducible human evaluation for text-to-image generation,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Toward verifiable and repro- ducible human evaluation for text-to-image generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:26.723656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:22.951084Z digest=sha256:cfcbaaafa6edda1438124b14b5173361d352b9f6baa71b6cfee742c2cb2aa88c

Observation b3b63564-ea32-4c26-af81-d021248b0511 · outbound

This paper cites On a test of whether one of two random variables is stochastically larger than the other,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio On a test of whether one of two random variables is stochastically larger than the other,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:23.027753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:23.027753Z digest=sha256:26e727733bc4a9653c401bf9b40170e5a73acd5674e9939f9fe8ade199c78828

Observation d5a36147-225b-454a-b648-d2874a2bb133 · outbound

This paper cites Use of ranks in one-criterion variance analysis,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Use of ranks in one-criterion variance analysis,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:23.092781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:23.092781Z digest=sha256:ebf084627977ce5fead0026c88e61d127613408a8011279e1531413aa31efc0a

Observation 7cef4de0-9de5-49ab-bb0c-1693c021e7f6 · outbound

This paper cites A multiple comparison rank sum test: treatments versus control,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio A multiple comparison rank sum test: treatments versus control,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:26.509594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:23.171413Z digest=sha256:589eb9edda066543c747373747e0654e15a9f9e540210315479945e4a428de61

Observation 76f4c52f-52b0-4d3e-98ec-35c0c826c1b3 · outbound

This paper cites The aligned rank transform for nonparametric factorial analyses using only anova procedures,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio The aligned rank transform for nonparametric factorial analyses using only anova procedures,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:23.280968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:23.280968Z digest=sha256:38a2e7eddef35ffd7be619436b3ebbea602843827f315ca473660cd1296295d3

Observation a3b6b4a0-01fe-4463-b283-4ca45952e0ed · outbound

This paper cites A new readability yardstick.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio A new readability yardstick

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:26.272858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:23.367652Z digest=sha256:7d034db79dc99a01385b3fe33a109bae66802bfe968b8066be8a3f8ebd7b5e6e

Observation 69676ee5-e695-4454-967e-8d8054546456 · outbound

This paper cites BYOL for Audio: Exploring pre-trained general-purpose audio representations,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio BYOL for Audio: Exploring pre-trained general-purpose audio representations,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:26.061815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:23.462651Z digest=sha256:445806fb92534ea072d526db17f7a43f09b0e731225e16abc4c15f16192b051c

Observation c88b6738-ae54-41de-89d5-dd104f49153c · outbound

This paper cites LDNet: Unified listener dependent modeling in MOS prediction for syn- thetic speech,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio LDNet: Unified listener dependent modeling in MOS prediction for syn- thetic speech,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.867292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:23.569795Z digest=sha256:ba5c7218ae9f0a835b495e10bf82a846be9724a09c3ee3925a4eff0d7a3f7310

Observation 59f3a409-d1e9-448e-9e7d-27db25507d18 · outbound

This paper cites Bidirectional LSTM networks for improved phoneme classification and recog- nition,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Bidirectional LSTM networks for improved phoneme classification and recog- nition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.655575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:23.667767Z digest=sha256:db2f50854ededbb4f9f4ed5bf3554622bf773c578348d71011e36755596e4a14

Observation 32ee808a-97fb-44d5-8245-2d40a7161a49 · outbound

This paper cites MBNet: MOS prediction for synthesized speech with mean-bias network,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio MBNet: MOS prediction for synthesized speech with mean-bias network,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.520959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:23.793918Z digest=sha256:5235a8d411fcad072d537c3de7b6a1600416b39edc3f274ea10d6d0456528db5

Observation e5344c90-5046-4c45-baad-848b36ecb386 · outbound

This paper cites Class- balanced loss based on effective number of samples,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Class- balanced loss based on effective number of samples,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.345388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:23.906898Z digest=sha256:4cf61ff138f4ca5bf3833407f49057d9e32c1efa476d8c4b81d12a5848e898d3

Observation b81dd3f9-befc-4967-a864-b4569594cff9 · outbound

This paper cites Adam: A method for stochastic op- timization,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Adam: A method for stochastic op- timization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.181627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:24.002568Z digest=sha256:61c4c58127e4cac3e9619a6295aa91dee37bd1b3e8ada43ac5d23411e35c17bf

Observation 59903483-f256-4b56-ac5b-656a59630b3b · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:24.966345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:24.119652Z digest=sha256:ec0f88d0d1702ba27e7ddd28970e22a3bca2ef260795582ef7dd5263bd067007

Observation 2115d22c-8f3f-41b3-8b74-39f0a190b9de · outbound

This paper cites CLAP: Learning audio concepts from natural language supervision,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio CLAP: Learning audio concepts from natural language supervision,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:24.796055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:24.207276Z digest=sha256:2b1d5e0dfc2ff127115d780090471f783f39086bab7ec401ac32a25d4b03d8db

Pith citing papers

Observation e9e26b03-e55f-4363-926e-a8e6934b2c7c · inbound

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio cites this paper.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:24.575515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T21:41:21.333361Z digest=sha256:dcde2216317203cae4890a604f9282c94d4fce1a4ede7a66d659a17354113a37