Pith. sign in

Paper Citation Record · LEDGER

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.00475.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00475 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:53.784375Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bff84d31-ef32-4d78-84c4-69ce575a5ab6 · outbound

This paper cites Leveraging AI to Generate Audio for User-generated Content in Video Games.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Leveraging AI to Generate Audio for User-generated Content in Video Games

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.404688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.218751Z digest=sha256:37c2b0658fc43015e6fc2ddc994584bc039983a2da2986cc41aac9d79c3ff3b5

Observation ced172d0-bccc-4d23-9345-70dfbb000bfd · outbound

This paper cites How should we evaluate synthesized environmental sounds,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences How should we evaluate synthesized environmental sounds,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.881042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.286779Z digest=sha256:fea49d43cf11c0822c8b2f00d4db00b6d7692689c0f92a79fd441dc6e79c1583

Observation 82ffe15f-da9e-4bac-b782-0f8ef3f81e72 · outbound

This paper cites An effective quality evaluation protocol for speech enhancement algorithms,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences An effective quality evaluation protocol for speech enhancement algorithms,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.622417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.327726Z digest=sha256:274cbc33b35f58f8e62c9d394fc60766f302ed5430de7119c92e0a1b1b55b8c0

Observation 0201f1fa-03c6-481f-ae34-d95e75947a30 · outbound

This paper cites Mel-cepstral distance measure for objective speech quality assessment,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Mel-cepstral distance measure for objective speech quality assessment,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.378676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.405548Z digest=sha256:858bf877c2fe2ab6908c740727c1096a7dafce524653f63c08e9eeb3fc44ca3d

Observation b14f7aa4-b61b-4b6c-a8ba-379dac7f31fb · outbound

This paper cites Human-CLAP: Human-perception-based contrastive language-audio pretraining,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Human-CLAP: Human-perception-based contrastive language-audio pretraining,

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:19:54.276450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.469206Z digest=sha256:6e02051c70e6ecef205ea3581caee663ba1edcc2077b97dfdf81fb0852f39177

Observation 265ce1c3-27ec-45c7-9cf2-6d144038162f · outbound

This paper cites Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.006373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.531504Z digest=sha256:c7fe90d643025f278f437a41442e1c8c703ddd99f358ba9a3cda300124f74a37

Observation 133a3232-6ffe-4d4a-94d8-936a231fffbd · outbound

This paper cites Foley Sound Synthesis at the DCASE 2023 Challenge.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Foley Sound Synthesis at the DCASE 2023 Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.610462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:52.610462Z digest=sha256:366232c019751b93538f37195a555e1897183676cc9fc66af6df04093fd870ad

Observation 6d174fd1-5325-4c00-9d6a-46fcbfaa7a26 · outbound

This paper cites Using dynamic time warping to find patterns in time series,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Using dynamic time warping to find patterns in time series,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.177899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.646054Z digest=sha256:93fb464c04f3c5e2e858addf65fde6059bab718cc94672eae166eeef9c6c7ad7

Observation a6d0491c-3747-4e14-bcc1-f031f6ba3cdb · outbound

This paper cites Dynamic time warping under subsequence,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Dynamic time warping under subsequence,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.920610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.719921Z digest=sha256:1588a040ec2f0b5a0cd992bafed1e3dfd90487df3319abf71b72b1c5759fe874

Observation 80ddb99f-504a-460e-9362-ec6e1f1bcf00 · outbound

This paper cites Relate: Subjective evaluation dataset for automatic evaluation of relevance between text and audio,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Relate: Subjective evaluation dataset for automatic evaluation of relevance between text and audio,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.721799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.782560Z digest=sha256:610d51d54d129635028555ec0acb1a745f3e511abfe453c114774e4669ec82f5

Observation d1a964c5-3e2a-4db2-b122-421b597c6a0d · outbound

This paper cites PAM: Prompting Audio-Language Models for Audio Quality Assessment.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences PAM: Prompting Audio-Language Models for Audio Quality Assessment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.851246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:52.851246Z digest=sha256:87ad559524fd9ac92e7a2f036ca9a2541b4333870b6ab1f55e636103dafa69c4

Observation 7feaf2e8-91c9-45f7-9929-e699670a212b · outbound

This paper cites Speech- BERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Speech- BERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.481697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:52.932662Z digest=sha256:581da006f307e159b5f833bdb36d7fa92cb859257fd28c92722aa3d5048ffe1b

Observation 4eba9c28-1718-4af8-b975-33ad5893b314 · outbound

This paper cites BERTScore: Evaluating text generation with bert,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences BERTScore: Evaluating text generation with bert,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.295310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:53.023947Z digest=sha256:d2906118775bd84f718ef0537b53d62aa66b10a9e41de1843117cb3e4f145904

Observation 47102028-a538-4314-89c2-0e946891576c · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.067501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.067501Z digest=sha256:5ca2bd60bb26d022e25e533b3933f4ae159f97692031293eac0ead5f0a8eba1a

Observation b742a959-b4ce-4b31-9cbe-a9fde70627e1 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.105669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:53.127818Z digest=sha256:43e7f7366a8b54e6af2477e14e14b7f55edd037c8ab6c10baa02432d73d83c74

Observation eee17444-cae4-44fc-8873-4ab47fe9dca4 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioCaps: Generating captions for audios in the wild,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.889041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:53.208905Z digest=sha256:fe790222479fccb4506bd844d58503905a3fb3f406943d0bb29ec655b8154f77

Observation 17d59307-b38d-4257-8c92-a740fd511246 · outbound

This paper cites AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.251131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.251131Z digest=sha256:e4eeba4008401e819efb9c694651759e7dff3d8cae1120e0c83db2636fc70d32

Observation 0a6dbeca-be39-4df6-811f-c9d77ee23099 · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.652067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:53.316231Z digest=sha256:4449db622b83b6fd604ce8acf7a5788188ab29ce6028ec26991510b3739c8cb1

Observation 22a2f17d-6983-401a-9a7e-2026ed796308 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioGen: Textually Guided Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.373430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.373430Z digest=sha256:668f9715b7267f34b49fa47bd841adedadc9676cce68eee37b6010ce6584ac0b

Observation 57db31d8-1b16-40b2-b399-f4ef562e7b20 · outbound

This paper cites BYOL for Audio: Exploring pre-trained general-purpose audio representations,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences BYOL for Audio: Exploring pre-trained general-purpose audio representations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.362877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:53.409957Z digest=sha256:649e55e4dd28826220f884489a4620e541260c341176de279ac0fa86caa33fbe

Observation 7be1632e-89a8-472e-ab05-4b491b7f8565 · outbound

This paper cites Backpropagation applied to handwritten zip code recognition,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Backpropagation applied to handwritten zip code recognition,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.471791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.471791Z digest=sha256:f5d8882a15d417ded64cb2d28794689880e0a475f693fa07ad03dbf23fdd061d

Observation cadd9621-dfc4-458e-aa49-1478b4d5916c · outbound

This paper cites Attention Is All You Need.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Attention Is All You Need

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.511609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.511609Z digest=sha256:72594c33abf9d4d44cc6779ee4beca67e9fcf49dcd7c24f53666363c23430325

Observation 1ff852fa-0a3a-448d-a0ba-b0c7ac9a08ce · outbound

This paper cites AST: Audio spectrogram trans- former,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AST: Audio spectrogram trans- former,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.153487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:53.573314Z digest=sha256:58f57a2765ce3f9a57da319c504248b2b3024b97dfe6d78974364700380333d5

Observation c69a0d0a-f007-4243-84ca-d10ffcdb4952 · outbound

This paper cites W ARP-Q: Quality prediction for generative neural speech codecs,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences W ARP-Q: Quality prediction for generative neural speech codecs,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.978164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:53.630560Z digest=sha256:726f187a39f2d226e11d2a490da53b26d564aa01ab513a12ec87e2102deed78f

Observation 86d194db-b93a-432f-becf-fed237d3f5be · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.679676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.679676Z digest=sha256:51689cfcbb5c08db02c8f2f7550799b7f56cc4790b34c05e35aa7e8faba56aa7

Observation 69fde75b-1511-4bf7-9cb4-46accb36684e · outbound

This paper cites A reference-free metric for language-queried audio source separation using contrastive language-audio pretraining,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences A reference-free metric for language-queried audio source separation using contrastive language-audio pretraining,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.738962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:53.743187Z digest=sha256:6e2eb714fb406285852ea1ccc3cb7f6a087522d55b57d8e95cc678a0d807f19a

Observation c1a49af6-baf8-462e-97bd-73da735584c0 · outbound

This paper cites Conformer: Local features coupling global representations for recognition and detection,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Conformer: Local features coupling global representations for recognition and detection,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.507500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:19:53.784375Z digest=sha256:b1235eefb79b2526ec0211b620bd459948588a5d8f885aca7118a724d55bb86f

Pith citing papers

No inbound Pith citation observations are available.