Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T16:20:13.303809Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2607.04619.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T16:20:13.303809Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b2cb80e9-f109-4d06-a863-cdf215c098c8 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning AudioCaps: Generating captions for audios in the wild,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b95b315f-901c-4a3f-92cb-5d6d23bcaf38 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Clotho: an audio captioning dataset,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 329824d9-23c4-4478-b7f2-e52693b62a7c · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning CLAP learning audio concepts from natural language supervision,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 497fdbf0-b212-4e35-bad6-76e291da7892 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e60ee710-fd40-4b46-b957-5a48c1b8e6fe · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning SLAM-AAC: enhancing audio captioning with paraphrasing augmentation and clap-refine through llms,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aafbbb77-e291-4a10-8d5a-46e020e7f20d · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Drcap: Decoding CLAP latents with retrieval-augmented generation for zero-shot audio captioning,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda596e6-7081-46cb-ba46-9774f1433403 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Unveiling Encoder-Free Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b72544-169a-4e8e-8433-1332e12f5897 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Available: https://api.semanticscholar.org/CorpusID: 270559398
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 499168d0-b18a-4fe7-b6a6-0fa015b37759 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Vision as LoRA
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c220655-7736-4013-8ad6-ac3f89bf8121 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Aura: Internalizing audio understanding into llms as lora,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e0f931e-e155-478c-b1d0-61010534587b · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2113221-944c-495b-b382-6029eca27b7b · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Available: https://api.semanticscholar.org/CorpusID: 270621008
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 751c1cd3-301b-4b3d-a1f6-ff316d9dcd59 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning CED: consistent ensemble distillation for audio tagging,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86073d4e-5cd6-4472-b4bf-9b10a357464b · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Llm can read spectrogram: Encoder-free speech-language modeling,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0b5ddf1-c83f-4874-b9d5-cfae8cf34cde · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Available: https://api.semanticscholar.org/CorpusID: 289132537
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb30c1b-0193-43e1-90c7-d1fc821cd507 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Fuyu-8B: A multimodal architecture for AI agents,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 478ee192-c146-47cc-90c4-25f65decfe5d · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre- training,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c2f2e9-3202-4c57-bd4b-ce561a047d67 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Gemma 4 model card,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724a040a-aa46-4b84-b4b0-af86109996ae · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Distilling the Knowledge in a Neural Network
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89725ebb-405d-47e0-a269-d9ffd01abb86 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Minillm: On-policy distillation of large language models,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 567b2f1a-6bce-4c59-b669-576331c06f00 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Qwen3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ba4318d-f899-4257-8af5-3b8ad590d284 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Lora: Low-rank adaptation of large language models,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb08dd0a-306f-469d-b534-e96f607325cf · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ec322f-199f-44c5-a93d-efafc958acee · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42d4d08-501d-49d1-989c-da3bc21bd33c · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Swin transformer: Hierarchical vision transformer using shifted windows,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbf35b7c-daea-48e4-b024-75f4d6e949ca · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d0bf0d-195b-4907-a672-0c3d7ba10e90 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Available: https://doi.org/10.1109/TASLP.2024.3419446
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90467b0c-3fd6-40bb-8a27-5d448510226c · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Auto-acd: A large-scale dataset for audio-language representation learning,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0304028a-0b4c-4b99-8830-58044652787a · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Macs - multi-annotator captioned soundscapes,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a166d7e1-8ea9-4a93-b035-0d25475d890b · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df0554b9-7867-45d2-a805-fc19e7265e23 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Cider: Consensus-based image description evaluation,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe9598e-098c-4ea9-a1db-e947fd3bbcc5 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Improved image captioning via policy gradient optimization of spider,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e7b5af90-8d46-410d-8992-333c54062544 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning SPICE: semantic propositional image caption evaluation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4fd7b5bb-e1e5-40c8-b2ec-7d21efb79f85 · outbound
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning METEOR: an automatic metric for MT evaluation with improved correlation with human judgments,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.