Pith. sign in

Paper Citation Record · LEDGER

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio

As of 22 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.23582.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23582 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:24.207276Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:21.333361Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T21:41:24.430826Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f56a6ed1-d90d-4f35-9f31-859ccba6066f · outbound

This paper cites a dog barking behind a human speech,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio a dog barking behind a human speech,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:30.509503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.279386Z digest=sha256:6ded1f7cdfd10c4a7db17e19ce5359298a504053b7ec7acf394dd3da59096437

Observation e9e26b03-e55f-4363-926e-a8e6934b2c7c · outbound

This paper cites RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:24.575515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.333361Z digest=sha256:93cbd1651c84ba08229f0d0f03684cdd0a5ff649977c28629d2b99c7f1f730dc

Observation 5bceae05-baab-4c89-aa9f-cbc2d88d8296 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:21.405172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:21.405172Z digest=sha256:1f222e791472b9df7f542200b37cee43246dfc2a778b93b22dc13fbbabb30783

Observation 7312557b-8972-4ae4-99b9-a421dd381213 · outbound

This paper cites 2” or higher, we added the exclusion of listeners with an average original audio rating of “6.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio 2” or higher, we added the exclusion of listeners with an average original audio rating of “6

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:30.309407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.469080Z digest=sha256:c1c677eed534e5a5eded44ac3142a25ead77668d6a4985a0a8fa927d1beb1785

Observation bca3e593-4e97-4146-829b-69f2ed23b05f · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:21.575339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:21.575339Z digest=sha256:f3fbc4a2681f8b93bce12cd68dcff0e9aa334224c4561a69c903e291f3f2ac5e

Observation 9bfb8ae9-087d-49b3-b354-e047c79026ea · outbound

This paper cites an unresolved cited work.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:29.964702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.623018Z digest=sha256:3d7eb3dca22736ad26e50ccdc9787e5cac225d7481a872a516014d19856cd75f

Observation 0854e656-819b-4fdc-8598-263f139e5827 · outbound

This paper cites an unresolved cited work.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:29.801710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.684278Z digest=sha256:c136c9880e667da8e5c06dd46c5c740e4976e90696d28ae3ea65a1161d273cef

Observation b1d044a9-a83d-4535-aae2-d532992e074b · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.443852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.222125Z digest=sha256:3169a2a71b362dcecde89059d2a6d68769a10bb1609de16f1797d09b0d608451

Observation b5648097-787c-462b-a79d-9917e23c3e0b · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound genera- tion,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Diffsound: Discrete diffusion model for text-to-sound genera- tion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:29.572302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.755408Z digest=sha256:7411c50a647330a86b91595832985d6b26ef5f7b12c75803f8ae0dd3c1939ca4

Observation 0c15c510-645e-4f29-a594-5ed720865bd9 · outbound

This paper cites Sound synthesis for impact sounds in video games,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Sound synthesis for impact sounds in video games,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:29.343842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.820308Z digest=sha256:1560d8e213f4af747fbb226284979ef7e293f7ecccd73fd70f72ecd8af01d9c8

Observation 230d1527-c23a-4766-8c98-9676f450937b · outbound

This paper cites Challenge on sound scene synthe- sis: Evaluating text-to-audio generation,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Challenge on sound scene synthe- sis: Evaluating text-to-audio generation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:29.149522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.868292Z digest=sha256:c155369bd6172048918beb76be4caa5680a281a8115d858fe453fd10cd36e2ed

Observation 9d038414-0e20-426e-a83e-519bc03d6f4f · outbound

This paper cites The V oiceMOS Challenge 2024: Beyond speech quality prediction,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio The V oiceMOS Challenge 2024: Beyond speech quality prediction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.991430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.935696Z digest=sha256:5b11f2c596fcb2a50d19451f64e3a9a81d6bd5a11b3792385a3d3e0a57f36eb5

Observation 575f9da5-92f0-4e6c-8278-bf9af17f6c91 · outbound

This paper cites Subjective-aligned dataset and metric for text-to-video quality assessment,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Subjective-aligned dataset and metric for text-to-video quality assessment,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.816240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.012364Z digest=sha256:50a6d2e2eafc2b265d4cb02e03ff66b0dfaafc5f4fa240234fea52f3bda42763

Observation 105ac452-f0d2-4310-a374-341099b3a6e2 · outbound

This paper cites Environmental sound synthesis from vocal imitations and sound event labels,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Environmental sound synthesis from vocal imitations and sound event labels,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.698899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.082880Z digest=sha256:e6e83244ade284eaa4e3d9f07f35faf626ef6edd5229f9241148b2f90bb1e579

Observation 9ccadc8a-8fb6-4579-992f-806bef2474af · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.578905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.143289Z digest=sha256:a4ee9303299d7a591e7cdb67ceac149b4cc29bf3b0fdd15e4de31164ee9fe6f2

Observation b0ebc53f-c036-4544-a474-bcd78621ca1b · outbound

This paper cites Text-to- audio generation using instruction guided latent diffusion model,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Text-to- audio generation using instruction guided latent diffusion model,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.266852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.788988Z digest=sha256:b2ded7da525ef9ba3605e42be4f2533e6a962bcd4f8c00583107abfab7d95630

Observation 12b470e2-c5a9-43ce-904c-3825d3b472df · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:22.283957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:22.283957Z digest=sha256:becda9e8d520bfb5f4e093d4eab50cbc928e54ef19b619042f23fa23281cf126

Observation 7203e59d-5374-4215-b076-8c93d4ef1fb5 · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio WavLM: Large-scale self-supervised pre-training for full stack speech processing,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:22.352486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:22.352486Z digest=sha256:a7071d7d46f30d73d8964aea90ca07140e417d8b051e8c5b8c12de71066a7925

Observation 66c5ab75-e0cd-4bd2-af30-4519267f34f2 · outbound

This paper cites Animal” category shows both of statistically significant differ- ences and interaction. Figure 2(a). shows that synthesized audio in the “Animal.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Animal” category shows both of statistically significant differ- ences and interaction. Figure 2(a). shows that synthesized audio in the “Animal

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:30.166092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.515288Z digest=sha256:cd17b3d381b65995b6cec4f810b7fca378f6ea2a0c740a33a7292655e1629af4

Observation 6beaa23e-2537-40ab-bcc8-1f182c8af643 · outbound

This paper cites PAM: Prompting audio-language models for audio quality assessment,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio PAM: Prompting audio-language models for audio quality assessment,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.252093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.421792Z digest=sha256:98ef82d2c8fcc3d1a5dd69723c8d8800291f5d1bd27029bb1dee7811a1e4e1fd

Observation 818e93a2-30ca-4461-b3a1-625b0cfab09d · outbound

This paper cites Audio-text mod- els do not yet leverage natural language,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Audio-text mod- els do not yet leverage natural language,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:28.017606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.493924Z digest=sha256:8599395996e233068676841cf88142e06df977ec24325c8906a148bfc773b133

Observation 1f895f04-a02e-432f-bfb2-63baf86708a3 · outbound

This paper cites AudioCaps: Generat- ing captions for audios in the wild,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio AudioCaps: Generat- ing captions for audios in the wild,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.791916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.559383Z digest=sha256:c5d065907fd65f4cf0ae626bca9d37cec91ec2d994b37558c23d7cf083945514

Observation 7dfe310c-937c-4c4f-a7fa-5adb96f9ac70 · outbound

This paper cites AudioLDM: text-to-audio generation with latent diffusion models,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio AudioLDM: text-to-audio generation with latent diffusion models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.637563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.634145Z digest=sha256:e08e959f348911882bc7f2333b13496613fd961239e00d86364f461724580b0c

Observation df6faa69-4fec-4869-bc10-86d9f8f9d286 · outbound

This paper cites AudioLDM 2: Learn- ing holistic audio generation with self-supervised pretraining,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio AudioLDM 2: Learn- ing holistic audio generation with self-supervised pretraining,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.472397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.715542Z digest=sha256:893d81195a434f0daea0d5982a2aaf4a5ac61a39e27c704f29b251299e0fcded

Observation dfe3687a-0663-47d7-aca1-27601b5f7a6b · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:27.012471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.873119Z digest=sha256:2755c06a30f206aa3ddc943dd68dfeca0c4dd377c0c49042bee5564b6e254450

Observation 11b55f60-ec7d-4e44-80d3-aa24986f3f83 · outbound

This paper cites Toward verifiable and repro- ducible human evaluation for text-to-image generation,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Toward verifiable and repro- ducible human evaluation for text-to-image generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:26.723656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:22.951084Z digest=sha256:16e8722ba1e2379786a62e757519a936403448a4be4fe7baee3f9892dbee7d9b

Observation b3b63564-ea32-4c26-af81-d021248b0511 · outbound

This paper cites On a test of whether one of two random variables is stochastically larger than the other,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio On a test of whether one of two random variables is stochastically larger than the other,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:23.027753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:23.027753Z digest=sha256:26e727733bc4a9653c401bf9b40170e5a73acd5674e9939f9fe8ade199c78828

Observation d5a36147-225b-454a-b648-d2874a2bb133 · outbound

This paper cites Use of ranks in one-criterion variance analysis,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Use of ranks in one-criterion variance analysis,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:23.092781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:23.092781Z digest=sha256:ebf084627977ce5fead0026c88e61d127613408a8011279e1531413aa31efc0a

Observation 7cef4de0-9de5-49ab-bb0c-1693c021e7f6 · outbound

This paper cites A multiple comparison rank sum test: treatments versus control,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio A multiple comparison rank sum test: treatments versus control,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:26.509594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:23.171413Z digest=sha256:54a0076d6aa6df6d9ad527b24849ae2010b641fe28bb6fb0b43f92aa5dc3741a

Observation 76f4c52f-52b0-4d3e-98ec-35c0c826c1b3 · outbound

This paper cites The aligned rank transform for nonparametric factorial analyses using only anova procedures,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio The aligned rank transform for nonparametric factorial analyses using only anova procedures,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:23.280968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:23.280968Z digest=sha256:38a2e7eddef35ffd7be619436b3ebbea602843827f315ca473660cd1296295d3

Observation a3b6b4a0-01fe-4463-b283-4ca45952e0ed · outbound

This paper cites A new readability yardstick.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio A new readability yardstick

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:26.272858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:23.367652Z digest=sha256:a5829e8924dbb350f1c8b56bc35ea9b20d2858abcb4ed3f84f54af789ce61ffb

Observation 69676ee5-e695-4454-967e-8d8054546456 · outbound

This paper cites BYOL for Audio: Exploring pre-trained general-purpose audio representations,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio BYOL for Audio: Exploring pre-trained general-purpose audio representations,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:26.061815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:23.462651Z digest=sha256:e04920c36661fc3384c0f3e79b7c85fa892512ba4b370426e51f8d65ffc99118

Observation c88b6738-ae54-41de-89d5-dd104f49153c · outbound

This paper cites LDNet: Unified listener dependent modeling in MOS prediction for syn- thetic speech,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio LDNet: Unified listener dependent modeling in MOS prediction for syn- thetic speech,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.867292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:23.569795Z digest=sha256:596ad8f9a371876b2426027f528a464d7c6686b6a0655b5a3843579c4edc50ab

Observation 59f3a409-d1e9-448e-9e7d-27db25507d18 · outbound

This paper cites Bidirectional LSTM networks for improved phoneme classification and recog- nition,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Bidirectional LSTM networks for improved phoneme classification and recog- nition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.655575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:23.667767Z digest=sha256:cb4800e37e6c9d9aea9eeec50bcfa45fedfae8e99b54efac163a7a4d7c58e4cd

Observation 32ee808a-97fb-44d5-8245-2d40a7161a49 · outbound

This paper cites MBNet: MOS prediction for synthesized speech with mean-bias network,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio MBNet: MOS prediction for synthesized speech with mean-bias network,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.520959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:23.793918Z digest=sha256:fa6d56c8ac4dc6c330f30c10edbd45e36f9fcb41755240f00a09aedb6526219e

Observation e5344c90-5046-4c45-baad-848b36ecb386 · outbound

This paper cites Class- balanced loss based on effective number of samples,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Class- balanced loss based on effective number of samples,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.345388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:23.906898Z digest=sha256:359e5f11570ba69234d3453809020558ffa65bbbf282324bd001aa2892c4f546

Observation b81dd3f9-befc-4967-a864-b4569594cff9 · outbound

This paper cites Adam: A method for stochastic op- timization,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Adam: A method for stochastic op- timization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:25.181627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:24.002568Z digest=sha256:b3bf3d30de1da9c772c9422708a50e1f9747c95ad394437fb2275264f64ed5a6

Observation 59903483-f256-4b56-ac5b-656a59630b3b · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:24.966345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:24.119652Z digest=sha256:3886a858bdce3faeca703f434936c2f141d82a6e21a07a4014e1178b25b67514

Observation 2115d22c-8f3f-41b3-8b74-39f0a190b9de · outbound

This paper cites CLAP: Learning audio concepts from natural language supervision,.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio CLAP: Learning audio concepts from natural language supervision,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:24.796055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:24.207276Z digest=sha256:069b9cd1c3e20db019962192858538b17e5a2ee0b42d40e34e33117dfa5e0c56

Pith citing papers

Observation e9e26b03-e55f-4363-926e-a8e6934b2c7c · inbound

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio cites this paper.

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:41:24.575515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:21.333361Z digest=sha256:93cbd1651c84ba08229f0d0f03684cdd0a5ff649977c28629d2b99c7f1f730dc