Pith. sign in

Paper Citation Record · LEDGER

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences

As of 21 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.00475.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00475 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:19:53.784375Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bff84d31-ef32-4d78-84c4-69ce575a5ab6 · outbound

This paper cites Leveraging AI to Generate Audio for User-generated Content in Video Games.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Leveraging AI to Generate Audio for User-generated Content in Video Games

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.404688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.218751Z digest=sha256:da9d3b698bd152f1b26d0bfe1064ee8de2e6830101eb5ed1c8aa588d8d331de6

Observation ced172d0-bccc-4d23-9345-70dfbb000bfd · outbound

This paper cites How should we evaluate synthesized environmental sounds,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences How should we evaluate synthesized environmental sounds,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.881042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.286779Z digest=sha256:506feb6814d002e6418cc4c0601b7926de27f3d07b59777b669afc27c5a3cf92

Observation 82ffe15f-da9e-4bac-b782-0f8ef3f81e72 · outbound

This paper cites An effective quality evaluation protocol for speech enhancement algorithms,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences An effective quality evaluation protocol for speech enhancement algorithms,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.622417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.327726Z digest=sha256:965381eaad3579bc7e469fc73a5abf9edb32b483fc0b2430d18360917f547cef

Observation 0201f1fa-03c6-481f-ae34-d95e75947a30 · outbound

This paper cites Mel-cepstral distance measure for objective speech quality assessment,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Mel-cepstral distance measure for objective speech quality assessment,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.378676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.405548Z digest=sha256:ae67a009a4b5cedb74f5f0113e68c3cacff9c134f6d3c2ffc505ecd91f9f9cc2

Observation b14f7aa4-b61b-4b6c-a8ba-379dac7f31fb · outbound

This paper cites Human-CLAP: Human-perception-based contrastive language-audio pretraining,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Human-CLAP: Human-perception-based contrastive language-audio pretraining,

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:19:54.276450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.469206Z digest=sha256:b6a287d4851a6a928e18e1e4f68717961340e6419fb4e844532796a52378e41f

Observation 265ce1c3-27ec-45c7-9cf2-6d144038162f · outbound

This paper cites Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:19:54.006373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.531504Z digest=sha256:b69c714080bbea9542ac2e92098d8cf80fca5c6e13b99bab7b6940113744cf66

Observation 133a3232-6ffe-4d4a-94d8-936a231fffbd · outbound

This paper cites Foley Sound Synthesis at the DCASE 2023 Challenge.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Foley Sound Synthesis at the DCASE 2023 Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.610462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:52.610462Z digest=sha256:f5b51f42a7ae0245fc62cd5826ba8e622c7295077e08d79892fea301e0aeec58

Observation 6d174fd1-5325-4c00-9d6a-46fcbfaa7a26 · outbound

This paper cites Using dynamic time warping to find patterns in time series,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Using dynamic time warping to find patterns in time series,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:57.177899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.646054Z digest=sha256:98b94c1ee857c5f68231aa2644b645d53626c4f0fc8767d471e7db7176f30356

Observation a6d0491c-3747-4e14-bcc1-f031f6ba3cdb · outbound

This paper cites Dynamic time warping under subsequence,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Dynamic time warping under subsequence,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.920610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.719921Z digest=sha256:d688567031123c1f86e312f545b8a8abb3e763c530c054bc9c7afb71443c6ce6

Observation 80ddb99f-504a-460e-9362-ec6e1f1bcf00 · outbound

This paper cites Relate: Subjective evaluation dataset for automatic evaluation of relevance between text and audio,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Relate: Subjective evaluation dataset for automatic evaluation of relevance between text and audio,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.721799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.782560Z digest=sha256:3318565d0ba9b38fad9909673c74ccc0336ade0f744b9e5d6ec38dff1d01b099

Observation d1a964c5-3e2a-4db2-b122-421b597c6a0d · outbound

This paper cites PAM: Prompting Audio-Language Models for Audio Quality Assessment.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences PAM: Prompting Audio-Language Models for Audio Quality Assessment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:52.851246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:52.851246Z digest=sha256:79e7d9ef80b4c3ca226cc081ee6dfcdc38a4a7605e21a1874b51d09f592c64bd

Observation 7feaf2e8-91c9-45f7-9929-e699670a212b · outbound

This paper cites Speech- BERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Speech- BERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.481697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:52.932662Z digest=sha256:b7bfe58da6e96b8e0b2e5806393e51b1cf4f028b80814ae6724e144d4c812b79

Observation 4eba9c28-1718-4af8-b975-33ad5893b314 · outbound

This paper cites BERTScore: Evaluating text generation with bert,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences BERTScore: Evaluating text generation with bert,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.295310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:53.023947Z digest=sha256:fe04b511082fcfad1e1f455bbd4215e428df10179adf42ebf9b70431e93175d6

Observation 47102028-a538-4314-89c2-0e946891576c · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.067501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.067501Z digest=sha256:2873d05a4b1562730b9c5b32f93eeb10802cd9c2b6493d4b084a67a3a4c13035

Observation b742a959-b4ce-4b31-9cbe-a9fde70627e1 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:56.105669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:53.127818Z digest=sha256:832f80a5599ce5bd528a35cb00f5bf94651fc88238f025b328904eacb2b2b6d1

Observation eee17444-cae4-44fc-8873-4ab47fe9dca4 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioCaps: Generating captions for audios in the wild,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.889041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:53.208905Z digest=sha256:a7d6bd32a104efb7220127b0ac623196dcd073bd3de3050d975002c374df7ad8

Observation 17d59307-b38d-4257-8c92-a740fd511246 · outbound

This paper cites AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.251131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.251131Z digest=sha256:a499b25fe54a913a6a278b87279310c4395e585c9c709baffb2b7de55a8dbe20

Observation 0a6dbeca-be39-4df6-811f-c9d77ee23099 · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.652067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:53.316231Z digest=sha256:6386f377240a44ab9f73dd20c608202fb93d9e1a464555efed9883c48e6d3450

Observation 22a2f17d-6983-401a-9a7e-2026ed796308 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AudioGen: Textually Guided Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.373430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.373430Z digest=sha256:258206484bb7878fe088d8dcb8a20a93a60b2e9a7db722f12dcfcd09ea7a4916

Observation 57db31d8-1b16-40b2-b399-f4ef562e7b20 · outbound

This paper cites BYOL for Audio: Exploring pre-trained general-purpose audio representations,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences BYOL for Audio: Exploring pre-trained general-purpose audio representations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.362877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:53.409957Z digest=sha256:86acb69c12856ee24ff471f9b08ce5a384d73f84175faf0c86fcf773f26b9ed3

Observation 7be1632e-89a8-472e-ab05-4b491b7f8565 · outbound

This paper cites Backpropagation applied to handwritten zip code recognition,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Backpropagation applied to handwritten zip code recognition,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.471791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.471791Z digest=sha256:56fa57f20b20796b4a7e586bb8f098a6588d2e7346c0d27568dbe2fcc5c6369e

Observation cadd9621-dfc4-458e-aa49-1478b4d5916c · outbound

This paper cites Attention Is All You Need.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Attention Is All You Need

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.511609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.511609Z digest=sha256:e5439640b0b45f2370de5a29a1213fc697d04fa2a1ad41dc3a7a14e006ca977d

Observation 1ff852fa-0a3a-448d-a0ba-b0c7ac9a08ce · outbound

This paper cites AST: Audio spectrogram trans- former,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences AST: Audio spectrogram trans- former,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:55.153487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:53.573314Z digest=sha256:e3da80b3d6ec93df69901e1ba212a4bc8eaa472ca3d213260c1bb3318d69fe01

Observation c69a0d0a-f007-4243-84ca-d10ffcdb4952 · outbound

This paper cites W ARP-Q: Quality prediction for generative neural speech codecs,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences W ARP-Q: Quality prediction for generative neural speech codecs,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.978164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:53.630560Z digest=sha256:3983ecb840485babba5d20e2cf1af6017378088c0b46b335ee3669619c68c91e

Observation 86d194db-b93a-432f-becf-fed237d3f5be · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:53.679676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:53.679676Z digest=sha256:468f61a520bd4309fda05df4714c9b0dd5b18382d717d073b720477fb6a93370

Observation 69fde75b-1511-4bf7-9cb4-46accb36684e · outbound

This paper cites A reference-free metric for language-queried audio source separation using contrastive language-audio pretraining,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences A reference-free metric for language-queried audio source separation using contrastive language-audio pretraining,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.738962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:53.743187Z digest=sha256:25a79c5ac77fb1bc9b2b65c883ef76e205eaf30969f010fa62d1519b0257b7ec

Observation c1a49af6-baf8-462e-97bd-73da735584c0 · outbound

This paper cites Conformer: Local features coupling global representations for recognition and detection,.

AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences Conformer: Local features coupling global representations for recognition and detection,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:19:54.507500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:19:53.784375Z digest=sha256:285ba86b306279e4292d807124e35d85ad0db48d44a22d2facfe0f8f32eea84d

Pith citing papers

No inbound Pith citation observations are available.