Pith. sign in

Paper Citation Record · LEDGER

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.00800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00800 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:35.993878Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:01:33.023609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:01:36.150100Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · outbound

This paper cites CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:01:36.208170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.023609Z digest=sha256:a44d11a9ef4c8d863354dc1107c1f6a2d791ee77135ec3d8b2a86224c0c92309

Observation c26a34a6-401c-444b-b56b-9d4426b7316c · outbound

This paper cites Automated Audio Captioning Methods using Pre- trained Language Model Pre-trained language models have been utilized in some stud- ies to improve AAC performance [4–11].

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated Audio Captioning Methods using Pre- trained Language Model Pre-trained language models have been utilized in some stud- ies to improve AAC performance [4–11]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.671760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.069426Z digest=sha256:193edf2a5480d27e11b2a5fd9fb19309e91fd63c96ea93faaaaf9f13504f929b

Observation 22ad47f7-076f-4c97-94c5-2ddeb94a5f13 · outbound

This paper cites First, En- CLAP transforms an input audio waveform into a CLAP audio embedding and EnCodec discrete tokens.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer First, En- CLAP transforms an input audio waveform into a CLAP audio embedding and EnCodec discrete tokens

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.662089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.141557Z digest=sha256:420a5ce0f9b75d6f264bebd4d5224541c582919cf12560ea3dfb04b0094858f0

Observation 2b899848-a6e9-43d5-a3ab-5cb4e692de14 · outbound

This paper cites an unresolved cited work.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:40.652837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.268808Z digest=sha256:7bd2be27d4e2c667b8226b232cda8c61e339fe56e958046b28f002a7a4ffdc72

Observation d9178af3-636b-4d95-b1ab-9fbb316c1c21 · outbound

This paper cites Experimental Setup We conducted experiments on two AAC datasets: Audio- Caps [23] and Clotho [24].

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Experimental Setup We conducted experiments on two AAC datasets: Audio- Caps [23] and Clotho [24]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.643800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.325366Z digest=sha256:9c21a5e930825c9ba766f02f95f564846cd3e69a645817a3c52cc95f9c4bd9b5

Observation c9cbd217-355f-43af-8165-8c4a49217fbf · outbound

This paper cites semantic-rich and discrete.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer semantic-rich and discrete

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.633679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.360355Z digest=sha256:7dd97ac2949686ef8db4e1d11bdb1637ec0ee0315b496d567166f443721517bb

Observation e5b382c7-cdf9-45c2-ae30-53407d95080f · outbound

This paper cites an unresolved cited work.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:01:40.624033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.427951Z digest=sha256:ea680744014eac7244a21fc3cd3de845a62252dac1ae54233bdf90a422e0cdb8

Observation 2b8053d3-a6a5-4e22-bdb8-2f33fd944383 · outbound

This paper cites Automated audio captioning with recurrent neural networks,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning with recurrent neural networks,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.615126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.508524Z digest=sha256:01f85fe4212dacb41bb862a03303f28dd6d2c9c68696c1985b41907658ace751

Observation 21aaba24-c546-4ab3-a3fd-5ecc41bf560d · outbound

This paper cites Automated audio captioning: An overview of recent progress and new challenges,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio captioning: An overview of recent progress and new challenges,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.606743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.600862Z digest=sha256:428b12a66f5e4e73e132ffacdd4f649efabfd488a9411fff7b923796cba679c1

Observation 327d26fd-73b7-45c0-9581-f4d75e449ba1 · outbound

This paper cites Beyond the status quo: A contemporary survey of advances and challenges in audio cap- tioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Beyond the status quo: A contemporary survey of advances and challenges in audio cap- tioning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.597069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.656600Z digest=sha256:b2c734675a4cfcf9aa55db4bbcf596cd43c9b2ad52ee3516eed9042a59336989

Observation c32a6fbe-9650-4e07-b6a3-7baa2371cee5 · outbound

This paper cites Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Captioning using Pre-Trained Large-Scale Language Model Guided by Audio-based Similar Caption Retrieval

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:33.727705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:33.727705Z digest=sha256:8d18d1bca06a3f69d0c334a3b7b789f25dc6540d2bd5bc23256ce4f9394844d6

Observation b6605451-fc0a-494c-abe4-d485ddffa251 · outbound

This paper cites Automated audio cap- tioning by fine-tuning bart with audioset tags,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Automated audio cap- tioning by fine-tuning bart with audioset tags,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.586909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.801814Z digest=sha256:df39e5d53743e8a9998b5e1bc9264c6046fd5675283d5145c3529fd536180d5c

Observation 405f177c-82a0-4e4f-b9b0-fbffb039535e · outbound

This paper cites Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.572547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.866374Z digest=sha256:08fa86c6a13ec60e6e217653097153e335bb8a943d4a90a11d4b73efee3aa5eb

Observation b04eacd2-2fe4-4413-a3e5-caf3b85c35fb · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Wavcaps: A chatgpt-assisted weakly- labelled audio captioning dataset for audio-language multimodal research,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.432143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.905971Z digest=sha256:488a625a6f0caf2b7fdc58db82f50d3e95f82a6068a8070f878452790c07927c

Observation fd3a5ab3-c46d-4b97-8b81-c063f3e49c89 · outbound

This paper cites Improving audio captioning models with fine- grained audio features, text embedding supervision, and llm mix- up augmentation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improving audio captioning models with fine- grained audio features, text embedding supervision, and llm mix- up augmentation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.220996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.966127Z digest=sha256:8d5d00ddc0214a31072d7d7bdf2b7847ac4fc1550452f9657e856c43806b9a62

Observation 0199daa7-7f7b-48a3-8f11-3e7babeef12a · outbound

This paper cites Recap: Retrieval-augmented audio captioning,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Recap: Retrieval-augmented audio captioning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:40.143102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.051316Z digest=sha256:2b8f4d2e6e688ef3cf30196366c4c684f8b11cbb8901438413ebedbcdce560bf

Observation ab9921c6-7f27-4813-8fdb-9715031c8ae5 · outbound

This paper cites Taming Data and Transformers for Audio Generation.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Taming Data and Transformers for Audio Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:34.136572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:34.136572Z digest=sha256:3c38d3609d3618d7bee68658234befb970f1afd8a910569d1ac7ecd8bdd5131f

Observation 95fd2992-8ce8-463c-ad70-4910f955bce7 · outbound

This paper cites Enhancing automated audio captioning via large language models with optimized audio encoding,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enhancing automated audio captioning via large language models with optimized audio encoding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.973182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.232063Z digest=sha256:0176742d160b117ec4d4d43230784a57a51c90b72584040da40ce026f295fa1a

Observation e5fb0d28-9ed9-4735-b5bf-fa0b242bd93f · outbound

This paper cites BEATs: Audio pre-training with acoustic tokenizers,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BEATs: Audio pre-training with acoustic tokenizers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.792153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.312586Z digest=sha256:f733c1f4f25449390a7f97e3196ac7dc7acd769540373d88821966298549169c

Observation 1c71dbdf-29d1-4713-b39f-5cc7e576fecd · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audio Set: An ontology and human-labeled dataset for audio events,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.645364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.374746Z digest=sha256:6fbfb87605d7198297ed63ada4fdfae4c131082da68a763debce16b5dd3f7a16

Observation 0e361f16-f008-4518-b797-c9ae4fd2cf9d · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.413088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.445594Z digest=sha256:5814e67fda3b48e03fa829c235530771172778288d0ad97d0d75e2bafabadc68

Observation 849d4c47-fa88-4064-a327-d07ebe4e913d · outbound

This paper cites High fidelity neural audio compression,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer High fidelity neural audio compression,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:39.203206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.506940Z digest=sha256:d6c082cfa9f83ea096cf42965f07b89e8f30b65e7f1e49e18e666755d7d725ae

Observation 39e2a501-5016-464b-ac77-b9882674171b · outbound

This paper cites BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer BART: Denoising sequence-to-sequence pre-training for natural language genera- tion, translation, and comprehension,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.968506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.589090Z digest=sha256:4c42f7e4a1159d399e9f97fb42d7052fd02c32878149833c97f94ec117096d39

Observation 7244c56a-c1e5-4a44-bd02-48dcc4a1123e · outbound

This paper cites SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:34.694185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:34.694185Z digest=sha256:2be74fcf47cfd7e502170eccdeb671dac880c5124252b41e4f85880da0de86e3

Observation 1daff668-6f25-4c57-859f-d70c503c2e27 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech language models,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Speechtok- enizer: Unified speech tokenizer for speech language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.710089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.783763Z digest=sha256:cdc06434813b41d31e53a2dc8515c4e42493d7be76a300187810f8c35fb5960b

Observation 825707c1-2a9f-4374-806e-0fcdaa537847 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.416120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.886016Z digest=sha256:382c2545dffb424b2d5ea2a521648299c24864502917938bd66cc9b214f13868

Observation dde80387-82ca-484d-9e41-15118b8e9377 · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SoundStream: An end-to-end neural audio codec,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.234586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:34.967305Z digest=sha256:6080fc18ccf5e9d0a97350664ca03c15d5a7af079d3855a47b45570ecb862625

Observation cad56588-70cf-4805-afe2-cb9b317a6268 · outbound

This paper cites vq- wav2vec: Self-supervised learning of discrete speech representa- tions,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer vq- wav2vec: Self-supervised learning of discrete speech representa- tions,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:38.051717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.045177Z digest=sha256:bf5fd6a97dfb503e099ec16da972b70e5290fa1249398993b7b0403c8c0e58f6

Observation dacc8b28-076a-42f1-999d-7607c3e29081 · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer How should we extract discrete audio tokens from self-supervised models?,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.876465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.127376Z digest=sha256:03d0ae3be00e2e9b0e943a42a9dbcbf9e07d64524c60899ce9fe2bf7e37897ec

Observation a310985e-a820-492c-96f7-f08ca33af77f · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Audiocaps: Generating captions for audios in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.725909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.207642Z digest=sha256:d06cfadfda0c8680fb07b8a0d07c8659948dc34be9eddd90129dc3195faa024b

Observation 6b35c6c1-310e-4ee6-8a4e-c1c384c400d2 · outbound

This paper cites Clotho: An audio cap- tioning dataset,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Clotho: An audio cap- tioning dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.529375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.295830Z digest=sha256:484022c4cd71fcf1e04faf71da22e521b4ab26561a885adea8c825d5ba1867f8

Observation 78e4e4d8-f205-4695-a62a-f2b7e767aa50 · outbound

This paper cites Decoupled Weight Decay Regularization.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:35.359333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:35.359333Z digest=sha256:60efc869b888b42feadcf4a3f74d3ea473c840e4e3f5226d5d5d13c69a95b070

Observation e32cd370-503b-4519-ad3a-18545a2f37e6 · outbound

This paper cites Meteor universal: Language spe- cific translation evaluation for any target language,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Meteor universal: Language spe- cific translation evaluation for any target language,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.362079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.442458Z digest=sha256:33d4753baa49629fb9818b55161678fa7950e58e0e50855c3ac7337110c5d3b8

Observation 61ddb5b5-50de-4892-bba4-d68103ef3633 · outbound

This paper cites CIDEr: Consensus-based image description evaluation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CIDEr: Consensus-based image description evaluation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.197754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.563960Z digest=sha256:e8df7a4a8f11af2a126b19eaa02d02ef3d5cfbc3734e28b3a152147d3cd6b701

Observation 89e7d13c-cac2-43b4-b452-0bba65e4a275 · outbound

This paper cites SPICE: Semantic propositional image caption evaluation,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer SPICE: Semantic propositional image caption evaluation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:37.030840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.643613Z digest=sha256:b442815a3e15de8ae68239ecdbb902ac673dac196788ad26283d88f3c5cb05a2

Observation a78f4136-0a21-428d-a152-04d6e1cace00 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Improved image captioning via policy gradient optimization of spider,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.844380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.725864Z digest=sha256:2170108ea3c11f85b2c47128f7887b8d065c9c93672fc0400a4f94c182f8086e

Observation ffd1cc54-6e51-4411-8c84-438a46ef7573 · outbound

This paper cites Can audio captions be evaluated with image caption metrics?,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Can audio captions be evaluated with image caption metrics?,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.701552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.799846Z digest=sha256:8ea01a88d4a5e1d9593b8ca9d7ec4cfaec8689ec15ccdec2cff0d32a06ba64dd

Observation 0e90ee26-cb6f-4015-84ee-ea6ba9f10b91 · outbound

This paper cites Enclap++: Ana- lyzing the enclap framework for optimizing automated audio cap- tioning performance,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Enclap++: Ana- lyzing the enclap framework for optimizing automated audio cap- tioning performance,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.552520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.906298Z digest=sha256:c8123e7e3eac054bbf32e4f0b7e821b0e04d52df51c1ced88b71ed5cd7c9e52f

Observation 0a1131d4-b1e1-4a29-91fa-1082570c3239 · outbound

This paper cites Slam-aac: Enhancing audio captioning with paraphras- ing augmentation and clap-refine through llms,.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Slam-aac: Enhancing audio captioning with paraphras- ing augmentation and clap-refine through llms,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:01:36.366008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:35.993878Z digest=sha256:19b008af7d23a139ea8f941979769b8d579fc79da883965cefb72f20eb9f1579

Pith citing papers

Observation d9fec535-2a2f-45d1-8589-33a4efdc3814 · inbound

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer cites this paper.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:01:36.208170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:01:33.023609Z digest=sha256:a44d11a9ef4c8d863354dc1107c1f6a2d791ee77135ec3d8b2a86224c0c92309