Pith. sign in

Paper Citation Record · LEDGER

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning

As of 17 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2607.04619.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04619 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T16:20:13.303809Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2cb80e9-f109-4d06-a863-cdf215c098c8 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning AudioCaps: Generating captions for audios in the wild,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:0ebdeb6ea0d1fb66659663dac92e014ef37e67cd200f144ff718cbc6d45635c7

Observation b95b315f-901c-4a3f-92cb-5d6d23bcaf38 · outbound

This paper cites Clotho: an audio captioning dataset,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Clotho: an audio captioning dataset,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:fcbaeb569d4f9413e437cb5c19e978b748bebab311616a3f7e5821b705aab78f

Observation 329824d9-23c4-4478-b7f2-e52693b62a7c · outbound

This paper cites CLAP learning audio concepts from natural language supervision,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning CLAP learning audio concepts from natural language supervision,

Reference 3

Resolution
verified exact
doi, observed 2026-07-11T16:28:07.838371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:fd8e8a9c0068abc6488a193288e0bc1b7a5b32cdaddd387d6a2fe91f2481479a

Observation 497fdbf0-b212-4e35-bad6-76e291da7892 · outbound

This paper cites Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:b60bd3a8ae9bbd038c648364c1b34fa801b2310e9c58970f122b8f0fa84d73d4

Observation e60ee710-fd40-4b46-b957-5a48c1b8e6fe · outbound

This paper cites SLAM-AAC: enhancing audio captioning with paraphrasing augmentation and clap-refine through llms,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning SLAM-AAC: enhancing audio captioning with paraphrasing augmentation and clap-refine through llms,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:3a2eeb65c25eb2bb32722717212c95a1805d54c9a258b7a1303e327fe42ebf92

Observation aafbbb77-e291-4a10-8d5a-46e020e7f20d · outbound

This paper cites Drcap: Decoding CLAP latents with retrieval-augmented generation for zero-shot audio captioning,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Drcap: Decoding CLAP latents with retrieval-augmented generation for zero-shot audio captioning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:8ae4169398e5e332b9885ffd2e0da486a623ed18ffb6ec8ed7837ecc66578474

Observation cda596e6-7081-46cb-ba46-9774f1433403 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Unveiling Encoder-Free Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:081353c7cc8d7404086f415ab4f4b327d49a1a8bca2bf2ca7912f83603510c3c

Observation 73b72544-169a-4e8e-8433-1332e12f5897 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 270559398.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Available: https://api.semanticscholar.org/CorpusID: 270559398

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:fdb4e772bd08616cf0d0c4acc2e637080c9398ae047a8efbfa0385160e854bc6

Observation 499168d0-b18a-4fe7-b6a6-0fa015b37759 · outbound

This paper cites Vision as LoRA.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Vision as LoRA

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:2fdd07e7d117af7afd6158f4615b4d91522d9d91e7bb99c84cbf363b8c3e18a9

Observation 4c220655-7736-4013-8ad6-ac3f89bf8121 · outbound

This paper cites Aura: Internalizing audio understanding into llms as lora,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Aura: Internalizing audio understanding into llms as lora,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:b63858e252376ef89b8e9e8be3aefc62e190d9f9ae6ac456c53e3984f1947912

Observation 2e0f931e-e155-478c-b1d0-61010534587b · outbound

This paper cites Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:0009e1c05856d9182bee42a00ecf8927443665b4898610c00c8954a869a2775f

Observation a2113221-944c-495b-b382-6029eca27b7b · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 270621008.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Available: https://api.semanticscholar.org/CorpusID: 270621008

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:5bf84dbbcff68a0fc2e5eb99a3953fceeb04d36066d8649645b8b0d4ac237d75

Observation 751c1cd3-301b-4b3d-a1f6-ff316d9dcd59 · outbound

This paper cites CED: consistent ensemble distillation for audio tagging,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning CED: consistent ensemble distillation for audio tagging,

Reference 13

Resolution
verified exact
doi, observed 2026-07-11T16:28:07.850454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:f1fe6db4798d2e29829b093e76af86acd2d21c5e86846f38b651103ecc3b1d9d

Observation 86073d4e-5cd6-4472-b4bf-9b10a357464b · outbound

This paper cites Llm can read spectrogram: Encoder-free speech-language modeling,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Llm can read spectrogram: Encoder-free speech-language modeling,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:1e5e4b71ae305e344515ef5c2366a0fe07e37100aa359489cde77e99d1052e74

Observation f0b5ddf1-c83f-4874-b9d5-cfae8cf34cde · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 289132537.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Available: https://api.semanticscholar.org/CorpusID: 289132537

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:a035eeeabe19b7dbbd40b862f40c758ba3e0850726531e5195131056b2cda1c4

Observation 5bb30c1b-0193-43e1-90c7-d1fc821cd507 · outbound

This paper cites Fuyu-8B: A multimodal architecture for AI agents,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Fuyu-8B: A multimodal architecture for AI agents,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:f9fa12c12b610553e2dc6933e8b869ed72ba3c85542f177b6030ab78937aecff

Observation 478ee192-c146-47cc-90c4-25f65decfe5d · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre- training,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre- training,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:b5a01765a4546b8b5db37a256c3d2fdac8eb3f042c5af7a84777f1935173accb

Observation f2c2f2e9-3202-4c57-bd4b-ce561a047d67 · outbound

This paper cites Gemma 4 model card,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Gemma 4 model card,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:900f5bbe3fe08575a670fcef60485c20efb6caee255339a3881c8e4a484038f3

Observation 724a040a-aa46-4b84-b4b0-af86109996ae · outbound

This paper cites Distilling the Knowledge in a Neural Network.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Distilling the Knowledge in a Neural Network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:f8c513629c6b8c66a75a0b52d976339ca252de8ec0494533b35111924be4d3f4

Observation 89725ebb-405d-47e0-a269-d9ffd01abb86 · outbound

This paper cites Minillm: On-policy distillation of large language models,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Minillm: On-policy distillation of large language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:7af6a61122462565e392782f2f735059021b4697744d5dba9739b11f085efa33

Observation 567b2f1a-6bce-4c59-b669-576331c06f00 · outbound

This paper cites Qwen3 Technical Report.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:1673fea62d033bc663488dce039e32d697ffa79158ae5aac80bccd24902abbb1

Observation 6ba4318d-f899-4257-8af5-3b8ad590d284 · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Lora: Low-rank adaptation of large language models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:42840fc73385b6bb1a47fa3629bc6b1ac31b149a7a909cf3b7dfd1dac737cc7f

Observation cb08dd0a-306f-469d-b534-e96f607325cf · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:5347a6407851f1de670e7b93e7b34e1eb3ff4f264d1bcec9ebc286fa0e5e6eb4

Observation 70ec322f-199f-44c5-a93d-efafc958acee · outbound

This paper cites HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:6e3ad10c7322d7f00e64fdc1dc78f3ef8b7110c408066c392e45c03c351dcf6c

Observation c42d4d08-501d-49d1-989c-da3bc21bd33c · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:e6358c64954172013c64c7b8188149a18e9370bc9c7e47d656f96117be102466

Observation cbf35b7c-daea-48e4-b024-75f4d6e949ca · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:e6052cfcce94465135ab05b46c0e16e16b34dd0abfd06c1d209e6f7aac4f2fc2

Observation d3d0bf0d-195b-4907-a672-0c3d7ba10e90 · outbound

This paper cites Available: https://doi.org/10.1109/TASLP.2024.3419446.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Available: https://doi.org/10.1109/TASLP.2024.3419446

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:34a96a9226371bf305f991775f39918f8aacb6a299926371a41bd7f7c2bd6fa2

Observation 90467b0c-3fd6-40bb-8a27-5d448510226c · outbound

This paper cites Auto-acd: A large-scale dataset for audio-language representation learning,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Auto-acd: A large-scale dataset for audio-language representation learning,

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-11T16:28:07.857288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:c7549bfbb29e747a5e2d9c0e7ba7a30e85c8687f5bf3ddbc80308f3a74062402

Observation 0304028a-0b4c-4b99-8830-58044652787a · outbound

This paper cites Macs - multi-annotator captioned soundscapes,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Macs - multi-annotator captioned soundscapes,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:322f5cdb705d906815c5c438583be3d2e6f0dc47233779ba66756545fd1ed1cb

Observation a166d7e1-8ea9-4a93-b035-0d25475d890b · outbound

This paper cites Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:9258d2dde0d39ef136aec7bf50d3c84cc209c8e0d064a7f30099dd2e672d4da9

Observation df0554b9-7867-45d2-a805-fc19e7265e23 · outbound

This paper cites Cider: Consensus-based image description evaluation,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Cider: Consensus-based image description evaluation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:4ad36a615e221065e3898bf3c2972636d552abf94745a5e9a6dd88c677a2a634

Observation 9fe9598e-098c-4ea9-a1db-e947fd3bbcc5 · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Improved image captioning via policy gradient optimization of spider,

Reference 32

Resolution
verified exact
doi, observed 2026-07-11T16:28:07.838639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:a532ee791f6c9e9d9fd29b828a632f323de8ede9ab6dc93ed9dafd5c19c7c020

Observation e7b5af90-8d46-410d-8992-333c54062544 · outbound

This paper cites SPICE: semantic propositional image caption evaluation,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning SPICE: semantic propositional image caption evaluation,

Reference 33

Resolution
verified exact
doi, observed 2026-07-11T16:28:07.821580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:01c0fe32df6d6ff889092de931835e1ec2a2f52a34af4dd3e829f7b82875b98d

Observation 4fd7b5bb-e1e5-40c8-b2ec-7d21efb79f85 · outbound

This paper cites METEOR: an automatic metric for MT evaluation with improved correlation with human judgments,.

CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning METEOR: an automatic metric for MT evaluation with improved correlation with human judgments,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T16:20:13.303809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:20:13.303809Z digest=sha256:37fb9806178ad9b556f393d9e726db94d61526d1857002c33602a6d9e3b078bc

Pith citing papers

No inbound Pith citation observations are available.