Pith. sign in

Paper Citation Record · LEDGER

Discrete Audio Representations for Automated Audio Captioning

As of 8 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2505.14989.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14989 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:41.089739Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:37.637530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:29:41.315494Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99f40907-dc3f-4326-98e4-a199b836bdaa · outbound

This paper cites Discrete Audio Representations for Automated Audio Captioning.

Discrete Audio Representations for Automated Audio Captioning Discrete Audio Representations for Automated Audio Captioning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:29:41.407940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:37.637530Z digest=sha256:56019b3c513898f6c45dcb1b0bc9354d0bc2df18590d308280009e0df826cd4d

Observation 33ad1828-4fd5-4d8e-bb50-0923888c62ad · outbound

This paper cites We construct BART-based and GPT-2 XL-based AAC systems, both utilizing semantic and acoustic tokens.

Discrete Audio Representations for Automated Audio Captioning We construct BART-based and GPT-2 XL-based AAC systems, both utilizing semantic and acoustic tokens

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:44.012926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:37.685143Z digest=sha256:9d1f091438f0b89a91a6e6e7ac87e5f30873ef90238d78cc88b96ed95c906f32

Observation cefd1b20-81c4-4ea7-a9c9-c6e2de374566 · outbound

This paper cites an unresolved cited work.

Discrete Audio Representations for Automated Audio Captioning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:29:43.877697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:37.771872Z digest=sha256:f5509f7acd7ddad590fdb7f3f735ce20900c6fafe4c151a21fc89547ee1d9b6b

Observation 72d3ae5e-6dc2-4d14-8e49-086a70a44d25 · outbound

This paper cites Our findings indicate that semantic tokens significantly outperform acoustic tokens in this context.

Discrete Audio Representations for Automated Audio Captioning Our findings indicate that semantic tokens significantly outperform acoustic tokens in this context

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.729200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:37.837897Z digest=sha256:cdb59a585a60d6e5171a775c6bdcacd0776beda9fab711c7b95341d0be697822

Observation 603fc8e4-6430-4c5f-b5f4-a5dffa6819bf · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges,.

Discrete Audio Representations for Automated Audio Captioning Automated audio captioning: An overview of recent progress and new challenges,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.606953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:37.933826Z digest=sha256:c5e3b985ba1bebac3620e033cd899cedcfb43853c4d682182dfdbf2073f3b04f

Observation c8825f6b-0b1c-4c98-b2d7-44a980073a55 · outbound

This paper cites Audio captioning based on transformer and pre-trained cnn,.

Discrete Audio Representations for Automated Audio Captioning Audio captioning based on transformer and pre-trained cnn,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.460031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:38.002613Z digest=sha256:3ebc60135c57543e5d4cf29c562e5e5ebd0c7fcb716dba3bd39f9fd6cf784c5c

Observation 59b5a379-4928-4a17-ad1d-929c748ca906 · outbound

This paper cites Automated audio caption- ing by fine-tuning bart with audioset tags,.

Discrete Audio Representations for Automated Audio Captioning Automated audio caption- ing by fine-tuning bart with audioset tags,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.305054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:38.078918Z digest=sha256:025a8a143c78dc42c3cd9979b57e39ddd75b80c7263e0cff1716759ba9b140a9

Observation f8b710d6-2282-4147-9d68-64b5259cbd75 · outbound

This paper cites Leveraging pre-trained bert for audio captioning,.

Discrete Audio Representations for Automated Audio Captioning Leveraging pre-trained bert for audio captioning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.170929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:38.173354Z digest=sha256:b10f445c23ef3f1e9e65ce3cd875eadf8becada187cd8b238cbab76e90620637

Observation 83fc5ec3-ab4e-4553-bf30-3d2aac66599b · outbound

This paper cites Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,.

Discrete Audio Representations for Automated Audio Captioning Conette: An efficient au- dio captioning system leveraging multiple datasets with task em- bedding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:43.032026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:38.274473Z digest=sha256:be6ad43ce44da7b50eef891e7e13b4d46cbb10911c9c48b060ff60e06c254874

Observation e6243da2-8ba2-4ee1-ac41-5b5c5bb07e3f · outbound

This paper cites Improving audio captioning mod- els with fine-grained audio features, text embedding supervision, and llm mix-up augmentation,.

Discrete Audio Representations for Automated Audio Captioning Improving audio captioning mod- els with fine-grained audio features, text embedding supervision, and llm mix-up augmentation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.889413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:38.375285Z digest=sha256:9a6ff7236cec94815b31c6a72f02fa833f105f4b54d345d69c55aa009898920f

Observation 35d13838-760a-4507-a964-a3d39b0c42e4 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Discrete Audio Representations for Automated Audio Captioning Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.478425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.478425Z digest=sha256:9953a74e490db4a037699629815bb81cd85061461ca60fdfad9e1239fdd98b5e

Observation bb3ef9d7-fdaa-49fb-a8eb-c94b7d2dcecf · outbound

This paper cites Adapting a convnext model to audio classification on au- dioset,.

Discrete Audio Representations for Automated Audio Captioning Adapting a convnext model to audio classification on au- dioset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.746482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:38.560986Z digest=sha256:750d4df1fa438e70e9afe8eda19d5b073b2336d04ef2e847b5a5180812539295

Observation 46b92cf6-81f9-48cc-933a-4c058438af93 · outbound

This paper cites Beats: audio pre-training with acoustic tok- enizers,.

Discrete Audio Representations for Automated Audio Captioning Beats: audio pre-training with acoustic tok- enizers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.551062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:38.654264Z digest=sha256:5273271ab26028c6b716d6b4cac7e413fa1d0ab8a69282038295ae16f263f498

Observation 6b9661e5-d984-425b-8ea9-a1c8e8aff9c9 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Discrete Audio Representations for Automated Audio Captioning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.730373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.730373Z digest=sha256:6a432f9031eb3836b21a802c79694d2f4accefe549cfd312289f158e3c8f71a4

Observation f97881d5-1890-40a5-8d5a-44f7c8b214a5 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

Discrete Audio Representations for Automated Audio Captioning BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.816322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.816322Z digest=sha256:8292df8155c7c94645534307f046de65111147569fb8e82b204d44f600fe106e

Observation faeb6c81-bccb-4221-a163-554f51667a31 · outbound

This paper cites Language models are unsupervised multitask learners,.

Discrete Audio Representations for Automated Audio Captioning Language models are unsupervised multitask learners,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:38.907144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:38.907144Z digest=sha256:b65f6e35699d7f48fbc0127c351dc4f97aa3be195470bab69c7ef93c991c31e4

Observation 6c960552-04af-4b3a-ba53-1ebb82abe3a5 · outbound

This paper cites STAB: Speech Tokenizer Assessment Benchmark.

Discrete Audio Representations for Automated Audio Captioning STAB: Speech Tokenizer Assessment Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.010252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.010252Z digest=sha256:0d18920cc7203cb90394693c6c0981fa1135cf0147e5513ba28b39d8c6166c10

Observation 1bf8414e-bbe7-412c-926d-3c551e889850 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Discrete Audio Representations for Automated Audio Captioning Soundstream: An end-to-end neural audio codec,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.090136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.090136Z digest=sha256:95d8957879f36b752c4b509ead9a81bea7f11b65f4d9c0eec06e861c5a8ed8d4

Observation 604426d6-34a9-45e0-8b98-d9a7518c4e3a · outbound

This paper cites High Fidelity Neural Audio Compression.

Discrete Audio Representations for Automated Audio Captioning High Fidelity Neural Audio Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.236427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.236427Z digest=sha256:c724e228290133d1294eb5e64eb1bd085c90cc4ded9b8d0a5274c9e15a5094ce

Observation a635e2e1-ee8c-4774-a573-a3dd8fd5abe8 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

Discrete Audio Representations for Automated Audio Captioning High-fidelity audio compression with improved rvqgan,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.371761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.371761Z digest=sha256:76f66d2e74ff957406089f261574cf79ab7a1fbdbb3f706180afb5383f263344

Observation db345d99-1cc7-44c3-b2e9-cff7d83fdb29 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Discrete Audio Representations for Automated Audio Captioning Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.476168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.476168Z digest=sha256:f00528a576e55ec93fb75afd4752e592c6b04cd79244a4dcbb37c430c302038a

Observation d440acd7-51cc-45cf-99e7-d54ad8af7e6c · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Discrete Audio Representations for Automated Audio Captioning Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.544115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.544115Z digest=sha256:125889072d3c07988dbb3d889c375b04b2ddbae1435120d734b497da39e5374c

Observation 820f0797-cf0f-4bd5-aad6-9be4af9109d9 · outbound

This paper cites RepCodec: A speech represen- tation codec for speech tokenization,.

Discrete Audio Representations for Automated Audio Captioning RepCodec: A speech represen- tation codec for speech tokenization,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.358530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:39.614586Z digest=sha256:b7b6f261f3b7a55f0b648f752dbe2a161b42f954d4bf0ca3978ca017c5755e46

Observation 2365d3a4-9a06-480b-a9e9-738cec03d94e · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Discrete Audio Representations for Automated Audio Captioning Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.691319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.691319Z digest=sha256:17d76adf1895e67b53cbbe0f7ef04c1710783f19861ee460977cf58e27aa3fd4

Observation 7e1079c0-39a2-4f82-bd04-ba21a7d587b4 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Discrete Audio Representations for Automated Audio Captioning BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.917468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.917468Z digest=sha256:3e28f95a396f05b81fc7c9ef152d947b5fc1d136a12f73dcbb22c47a61263047

Observation d6434d90-b764-4b0a-b87d-a5606774bcdf · outbound

This paper cites Exploration of efficient end-to-end asr using discretized input from self-supervised learning,.

Discrete Audio Representations for Automated Audio Captioning Exploration of efficient end-to-end asr using discretized input from self-supervised learning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.206416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:40.021098Z digest=sha256:ec7dfb7e909ede3e5228f25e27b62573144dd48a7ea9b768aef6ba0ffe93fed1

Observation 9031b10f-1abd-4b2c-9429-6d768514a162 · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?.

Discrete Audio Representations for Automated Audio Captioning How should we extract discrete audio tokens from self-supervised models?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.136058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.136058Z digest=sha256:f24199c0d8a9d45f59e741848370d7e9d8084d9b8e293cf31a925b3aa13450fc

Observation f071cc86-68db-4404-ba0a-947998171708 · outbound

This paper cites Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,.

Discrete Audio Representations for Automated Audio Captioning Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:42.081468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:40.274350Z digest=sha256:2e0e5ff1771a33f19de57bd343433504f0e23afee9d2ec0b74fc49bc84860ac4

Observation e716f219-6b97-4506-b963-09661381826f · outbound

This paper cites Enclap: Combining neu- ral audio codec and audio-text joint embedding for automated au- dio captioning,.

Discrete Audio Representations for Automated Audio Captioning Enclap: Combining neu- ral audio codec and audio-text joint embedding for automated au- dio captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.951056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:40.400959Z digest=sha256:8662b32dec7a3aa37b95bd9910cb95242e9163dc5f8575d36b7c3b20049268d9

Observation 559d8f9d-1118-412f-8dc3-f2ee8e0538a6 · outbound

This paper cites Neural discrete represen- tation learning,.

Discrete Audio Representations for Automated Audio Captioning Neural discrete represen- tation learning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.513005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.513005Z digest=sha256:86a94649eb7698d26793e486ebc9eda5557e5eb506e1ec2cf866d235635cc9c4

Observation 5ca75e92-a442-4ac8-9b6b-6902dbdd3bb9 · outbound

This paper cites Investigating lo- cal and global information for automated audio captioning with transfer learning,.

Discrete Audio Representations for Automated Audio Captioning Investigating lo- cal and global information for automated audio captioning with transfer learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.823902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:40.581171Z digest=sha256:46e64f483d0973f403f2c3f128e4f671db787a3545af0ef8f0233b334eb4c8c4

Observation a11eb655-7160-47d5-9313-fe082c58d128 · outbound

This paper cites Prefix tuning for auto- mated audio captioning,.

Discrete Audio Representations for Automated Audio Captioning Prefix tuning for auto- mated audio captioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.713383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:40.658035Z digest=sha256:ac4ba060b7829967d897ec28f363cf57d05c8abd6b90339f8f11d91bb314c3ea

Observation fa09ffca-3051-4b66-b7ac-c9a49cbceec3 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

Discrete Audio Representations for Automated Audio Captioning Audio set: An ontology and human-labeled dataset for audio events,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.769821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.769821Z digest=sha256:92964160f478db263c2a5ac63aba2909e807456f4e09fa864ea8814ec914aa72

Observation 6c7a4798-49e0-4a63-8e6a-5c18735ef4c8 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

Discrete Audio Representations for Automated Audio Captioning Clotho: An audio cap- tioning dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.890406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.890406Z digest=sha256:4270eb828cb53f7995effb6e0fca2fc48e4d57d1c8197f651877870bef1a275d

Observation ca949e47-509f-4b35-9b34-a44501a28463 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

Discrete Audio Representations for Automated Audio Captioning Improved image captioning via policy gradient optimization of spider,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:40.981070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:40.981070Z digest=sha256:d9a5e544e21f1cf309eee968786fc380f4d7681e391546df7ca2fe309599acde

Observation 5e43b8e5-c476-4d03-a339-29dcf1715581 · outbound

This paper cites Can audio captions be evaluated with image caption metrics?.

Discrete Audio Representations for Automated Audio Captioning Can audio captions be evaluated with image caption metrics?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:29:41.528936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:41.089739Z digest=sha256:00d7e4310c44a5b8ec808affc85710342f10108f15fe5a516e113c1593da9539

Pith citing papers

Observation 99f40907-dc3f-4326-98e4-a199b836bdaa · inbound

Discrete Audio Representations for Automated Audio Captioning cites this paper.

Discrete Audio Representations for Automated Audio Captioning Discrete Audio Representations for Automated Audio Captioning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:29:41.407940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:29:37.637530Z digest=sha256:56019b3c513898f6c45dcb1b0bc9354d0bc2df18590d308280009e0df826cd4d