Pith. sign in

Paper Citation Record · LEDGER

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.00800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00800 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:35.993878Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:33.023609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:01:36.150100Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · outbound

This paper cites CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:01:36.208170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.023609Z digest=sha256:51faf0545f24fb8633f837c6fb968e7d8b84901a42f97086d141bd2c97ec42df

Observation c26a34a6-401c-444b-b56b-9d4426b7316c · outbound

This paper cites Automated Audio Captioning Methods using Pre- trained Language Model Pre-trained language models have been utilized in some stud- ies to improve AAC performance [4–11].

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated Audio Captioning Methods using Pre- trained Language Model Pre-trained language models have been utilized in some stud- ies to improve AAC performance [4–11]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.671760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.069426Z digest=sha256:1e17b219a338e4a831e6a02979231fd8642843507cd6c5262f8575e66a9d34de

Observation 22ad47f7-076f-4c97-94c5-2ddeb94a5f13 · outbound

This paper cites First, En- CLAP transforms an input audio waveform into a CLAP audio embedding and EnCodec discrete tokens.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer First, En- CLAP transforms an input audio waveform into a CLAP audio embedding and EnCodec discrete tokens

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.662089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.141557Z digest=sha256:cc0399d45531910c742379f6c76bbb508841c483e73eaf9194c4db7195684471

Observation 2b899848-a6e9-43d5-a3ab-5cb4e692de14 · outbound

This paper cites an unresolved cited work.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:40.652837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.268808Z digest=sha256:7abb237e3415a62b26c2c488c22a4c9dcf2c6d53d22b2272ca5cd5eba92a8d21

Observation d9178af3-636b-4d95-b1ab-9fbb316c1c21 · outbound

This paper cites Experimental Setup We conducted experiments on two AAC datasets: Audio- Caps [23] and Clotho [24].

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Experimental Setup We conducted experiments on two AAC datasets: Audio- Caps [23] and Clotho [24]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.643800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.325366Z digest=sha256:65d1a94d92f7cc075741a987e6c9be0e9c33fd8877acfbcdbfddfa0a8fd8e870

Observation c9cbd217-355f-43af-8165-8c4a49217fbf · outbound

This paper cites semantic-rich and discrete.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer semantic-rich and discrete

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.633679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.360355Z digest=sha256:ef3a317e26933ae399205928065d153c59ca239a053c962bea0d9576658423a7

Observation e5b382c7-cdf9-45c2-ae30-53407d95080f · outbound

This paper cites an unresolved cited work.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:40.624033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.427951Z digest=sha256:3b2bc4bb8a58136a2e598372bd2ac1000708d4193062a66681763c652db24fe4

Observation 2b8053d3-a6a5-4e22-bdb8-2f33fd944383 · outbound

This paper cites Automated audio captioning with recurrent neural networks,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning with recurrent neural networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.615126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.508524Z digest=sha256:737f4b9b8f28d0471797fe8c474cb68e7c9b9e0b546af734c7210d1d1027a3a1

Observation 21aaba24-c546-4ab3-a3fd-5ecc41bf560d · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning: An overview of recent progress and new challenges,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.606743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.600862Z digest=sha256:d1852ee89e0b273c20c04503110ce061c0f3c97ace896ce6b99e22e5d6dc137a

Observation 327d26fd-73b7-45c0-9581-f4d75e449ba1 · outbound

This paper cites Beyond the status quo: A contemporary survey of advances and challenges in audio cap- tioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Beyond the status quo: A contemporary survey of advances and challenges in audio cap- tioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.597069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.656600Z digest=sha256:27fdcee6d4c8ec553db0ead445496041a87b103a794864bd1bc0c829e82264e4

Observation c32a6fbe-9650-4e07-b6a3-7baa2371cee5 · outbound

This paper cites Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:33.727705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:33.727705Z digest=sha256:b09bb13d596bb4dcc3e7eae5f9b34b7ee895ccbe7f9c16889fd67958da81b23f

Observation b6605451-fc0a-494c-abe4-d485ddffa251 · outbound

This paper cites Automated audio cap- tioning by fine-tuning bart with audioset tags,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio cap- tioning by fine-tuning bart with audioset tags,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.586909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.801814Z digest=sha256:050ae8ff1085fd49729c7a4cec0f3354072274deb9e1224f356dee3b118ea410

Observation 405f177c-82a0-4e4f-b9b0-fbffb039535e · outbound

This paper cites Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.572547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.866374Z digest=sha256:4ad85bca1ec1f9003bc9067b2e526f190346c518a07067c686bc3cc198d99297

Observation b04eacd2-2fe4-4413-a3e5-caf3b85c35fb · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.432143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.905971Z digest=sha256:b9d004497e838eb98318de04cf743091ea0998d3b31aebf3d55d0bfa2a85fbf7

Observation fd3a5ab3-c46d-4b97-8b81-c063f3e49c89 · outbound

This paper cites Improving audio captioning models with fine- grained audio features, text embedding supervision, and llm mix- up augmentation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improving audio captioning models with fine- grained audio features, text embedding supervision, and llm mix- up augmentation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.220996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.966127Z digest=sha256:e594d13ca59050c3d0ac0fa10b224bf8f204e87f7e6ac4ec81466f259277aa80

Observation 0199daa7-7f7b-48a3-8f11-3e7babeef12a · outbound

This paper cites Recap: Retrieval-augmented audio captioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Recap: Retrieval-augmented audio captioning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.143102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.051316Z digest=sha256:97b5b073506705e5f8389dd83868cd88ae87a4ceb429a02d85a30a508ee6eaf6

Observation ab9921c6-7f27-4813-8fdb-9715031c8ae5 · outbound

This paper cites Taming Data and Transformers for Audio Generation.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Taming Data and Transformers for Audio Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:34.136572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:34.136572Z digest=sha256:2481944edbc0ff1e77ab6a2fc8b76e20545682f85c5a99d25e35875fa957aeef

Observation 95fd2992-8ce8-463c-ad70-4910f955bce7 · outbound

This paper cites Enhancing automated audio captioning via large language models with optimized audio encoding,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enhancing automated audio captioning via large language models with optimized audio encoding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.973182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.232063Z digest=sha256:7e83461abc81a1d4defae32c909dbff4c8a440d2d538b47ce10d19f68b5ed254

Observation e5fb0d28-9ed9-4735-b5bf-fa0b242bd93f · outbound

This paper cites BEATs: Audio pre-training with acoustic tokenizers,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BEATs: Audio pre-training with acoustic tokenizers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.792153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.312586Z digest=sha256:bf7f95effec99bb8d8465d1ebd59a86ef5047ecbbe3e370485c6baa06012adaf

Observation 1c71dbdf-29d1-4713-b39f-5cc7e576fecd · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Set: An ontology and human-labeled dataset for audio events,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.645364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.374746Z digest=sha256:7539efea01d9d081abcf60d27fecd8bf4dadbe6df7cc72f4f805c19a3919cba4

Observation 0e361f16-f008-4518-b797-c9ae4fd2cf9d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.413088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.445594Z digest=sha256:d277c835d261cea8c2caa76a2728fae71bad5ca0c5254f52c88cf8de5c0e3fc8

Observation 849d4c47-fa88-4064-a327-d07ebe4e913d · outbound

This paper cites High fidelity neural audio compression,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer High fidelity neural audio compression,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.203206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.506940Z digest=sha256:3601038c3839529832ca9d12b6d1028cec343b1ac718cbe00725b3df59d9404d

Observation 39e2a501-5016-464b-ac77-b9882674171b · outbound

This paper cites BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.968506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.589090Z digest=sha256:26685aa24e6232e169cda03dc440208bcd83d59d2f507eeb891f5e7fed73bf7e

Observation 7244c56a-c1e5-4a44-bd02-48dcc4a1123e · outbound

This paper cites SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:34.694185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:34.694185Z digest=sha256:92e3470b8dcfcdb844dd130a5c4886a9de833731ec33ba25f54aa461141d7fd0

Observation 1daff668-6f25-4c57-859f-d70c503c2e27 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech language models,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Speechtok- enizer: Unified speech tokenizer for speech language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.710089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.783763Z digest=sha256:fd48438e1380dd2d08d124fe6f4d8d5d6b0d4fa71a7ebbe8c78f3e918ffb2254

Observation 825707c1-2a9f-4374-806e-0fcdaa537847 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.416120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.886016Z digest=sha256:d77f5a54f4f2ed4dcad133dd56b3d1d7cc569110d51313e39c3a18cb554e1bcf

Observation dde80387-82ca-484d-9e41-15118b8e9377 · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SoundStream: An end-to-end neural audio codec,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.234586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.967305Z digest=sha256:6744d0483c9c19f90fe61807ea7b8cf3e8333b144d72647078e858403074366f

Observation cad56588-70cf-4805-afe2-cb9b317a6268 · outbound

This paper cites vq- wav2vec: Self-supervised learning of discrete speech representa- tions,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer vq- wav2vec: Self-supervised learning of discrete speech representa- tions,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.051717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.045177Z digest=sha256:702a6175ae49a0272a4755d1f8a176e74295bcc59845cce21011e08f66c52e66

Observation dacc8b28-076a-42f1-999d-7607c3e29081 · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer How should we extract discrete audio tokens from self-supervised models?,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.876465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.127376Z digest=sha256:2f9ae1632bf43a4781a0e69e53c6c821dc8ddc726fa732b9146455f67cb40e8a

Observation a310985e-a820-492c-96f7-f08ca33af77f · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audiocaps: Generating captions for audios in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.725909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.207642Z digest=sha256:7d151983908bda4affa74cb22b192fe93c06d620bdb61a8f79a09dd09403838e

Observation 6b35c6c1-310e-4ee6-8a4e-c1c384c400d2 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Clotho: An audio cap- tioning dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.529375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.295830Z digest=sha256:92707cc9308013f861bd51f9f49ae0c2f527c748896f6ff2df85cad1d622656a

Observation 78e4e4d8-f205-4695-a62a-f2b7e767aa50 · outbound

This paper cites Decoupled Weight Decay Regularization.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:35.359333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:35.359333Z digest=sha256:07ed6ae3ad4d4714ad2eb348495f30289eed6c8842ef27d6949b0e005d41fb11

Observation e32cd370-503b-4519-ad3a-18545a2f37e6 · outbound

This paper cites Meteor universal: Language spe- cific translation evaluation for any target language,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Meteor universal: Language spe- cific translation evaluation for any target language,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.362079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.442458Z digest=sha256:7e3fe17b93b7b161893258db3ccab9106cce87f3a77b718bde4ac9ceb7ccc2d6

Observation 61ddb5b5-50de-4892-bba4-d68103ef3633 · outbound

This paper cites CIDEr: Consensus-based image description evaluation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CIDEr: Consensus-based image description evaluation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.197754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.563960Z digest=sha256:1be1dc67a2c74422d57d809d5e145817245f9756fa1e1fafc9ba6c6592b8e0a1

Observation 89e7d13c-cac2-43b4-b452-0bba65e4a275 · outbound

This paper cites SPICE: Semantic propositional image caption evaluation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SPICE: Semantic propositional image caption evaluation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.030840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.643613Z digest=sha256:9bc805da240913bc2340247f9271a9c09e538c1e6f7544e07d6b88a87c8cc4ec

Observation a78f4136-0a21-428d-a152-04d6e1cace00 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improved image captioning via policy gradient optimization of spider,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.844380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.725864Z digest=sha256:05ecf026229c9986dd0a73ebd47f827143c8269ec9febaa8a298e0d20324b87a

Observation ffd1cc54-6e51-4411-8c84-438a46ef7573 · outbound

This paper cites Can audio captions be evaluated with image caption metrics?,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Can audio captions be evaluated with image caption metrics?,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.701552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.799846Z digest=sha256:3feffed10c29bbe32de67212e705435dacdab55ca0e56476955947d30e36ae87

Observation 0e90ee26-cb6f-4015-84ee-ea6ba9f10b91 · outbound

This paper cites Enclap++: Ana- lyzing the enclap framework for optimizing automated audio cap- tioning performance,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap++: Ana- lyzing the enclap framework for optimizing automated audio cap- tioning performance,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.552520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.906298Z digest=sha256:f67a7796883f7ee16d6217ede6cb09b8910b063d69c39c2b0f19fc4bb7357331

Observation 0a1131d4-b1e1-4a29-91fa-1082570c3239 · outbound

This paper cites Slam-aac: Enhancing audio captioning with paraphras- ing augmentation and clap-refine through llms,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Slam-aac: Enhancing audio captioning with paraphras- ing augmentation and clap-refine through llms,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.366008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.993878Z digest=sha256:1620edd9893ff3564d25162c2d83d151874edc39617df3a5a9be5f41ed5d1bde

Pith citing papers

Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · inbound

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer cites this paper.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:01:36.208170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.023609Z digest=sha256:51faf0545f24fb8633f837c6fb968e7d8b84901a42f97086d141bd2c97ec42df