Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T05:57:09.345415Z
Paper Citation Record · LEDGER
As of 25 July 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2605.29628.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T05:57:09.345415Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-24T06:31:00.690269+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bf76a0b9-a4c2-41ac-a666-a51fe07e10ac · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Learning transferable visual models from natural language supervision,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf90d5c1-6494-4fb1-9e32-91fb3086d426 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Natural language supervision for general-purpose audio representations,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba783362-e1b2-421e-8a32-2b5c9326cd14 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80c9bca2-1158-4312-ae47-5a8ac9944ab7 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio- language multimodal research,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4005e1c8-976b-44d3-8a32-15db5fa0a721 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Auto-acd: A large-scale dataset for audio-language representation learning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bac5fcd-947d-4c25-85d2-cbb94a2516b2 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Audiosetcaps: An enriched audio-caption dataset using automated gener- ation pipeline with large audio and language models,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a32f3683-0ed4-40b7-b62b-175adab839d7 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings A knowledge distillation approach to im- proving language-based audio retrieval models,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e90fdce7-0e45-4ea7-9ebc-b5ef48b75733 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Aistat lab system for dcase 2025 task6: Language-based audio retreival,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a2bdaf-b58f-4c48-9d74-14c602bb9587 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Recap: Retrieval-augmented audio captioning,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd402f97-9875-47d3-8617-653c8ab47446 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diffusion-based diverse audio captioning with retrieval-guided langevin dynamics,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a1c4316-182a-424f-8aea-467bf608a4c3 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Zero-shot diverse audio captioning with diffusion models,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e10c007-48f9-4dbe-b198-a8bef43d222a · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Audioldm: text-to-audio generation with latent diffusion models,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad1a90de-32d4-4aec-8ca0-ce3d60bac504 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5f6122b-f453-4d12-a221-57290e772d20 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Weakly-supervised automated audio captioning via text only training,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c80b2f-542d-4ae2-b716-31f0052450ba · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Training audio captioning models without audio,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 705d6885-3fa2-4e53-ba52-40dc94a1b6b3 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Zero-shot audio captioning using soft and hard prompts,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517d37dc-d995-4b18-9050-d58d15f201f3 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Drcap: Decoding clap latents with retrieval-augmented generation for zero-shot audio captioning,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3405abec-0aef-4e87-b326-398c3540b467 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9467263a-50c7-4a30-a2af-b1f1ec2e94ef · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Two effects, one trigger: On the modality gap, object bias, and information imbalance in contrastive vision-language models,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b121d8e5-0998-4b21-af9c-b7c55158ceb5 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Decipher the modality gap in multimodal contrastive learning: From convergent representations to pairwise alignment.arXiv preprint arXiv:2510.03268
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation d9f79375-a353-42c1-9e42-f2e43fc89d16 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Decap: Decoding CLIP latents for zero-shot captioning via text-only training,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8bc2571-5742-47e3-b795-747b74dc22d2 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diffgap: A lightweight diffusion module in contrastive space for bridging cross-model gap,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aebe9d8-22a5-4281-9766-c294568d773c · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diffusion-link: Diffusion probabilistic model for bridging the audio-text modality gap,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 2abe19dc-5c38-4b5c-afc7-fccac2fb3c0a · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Pull it together: Reducing the modality gap in contrastive learning,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54a2e44b-1341-4541-8f30-8ad2955ac96b · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Interpreting clip with sparse linear concept embeddings (splice),
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd08f819-fcaa-467b-a8d8-3e96c3809afc · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation ad996e2e-43f7-4f05-ae86-3585194ea51f · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Scocca: Multi-modal sparse concept decomposition via canonical correlation analysis,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 4da3e290-4b67-44a2-89ca-d6f74da0be72 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Momentum contrast for unsupervised visual representation learning,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 475ea1a4-0854-46e7-8870-78bb50e2c591 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Pengi: An audio language model for audio tasks,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b05a1972-270b-4202-9cb2-4fb9dafd1a4c · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 833b8a30-19f8-4a9f-ba06-48d2ce7367a5 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings A holistic approach to unifying automatic concept extraction and concept importance estimation,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9cd8797-0d9f-4b94-a66d-ca639d40d7fe · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav),
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2fcb08-1e0e-4218-ac1f-fd56129ea48c · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Craft: Concept recursive activation factorization for explainability,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a63104-7bc1-4639-9236-52274dc23bee · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Multi-dimensional concept discovery (mcd): a unifying framework with completeness guarantees,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26bfd95d-d137-4e49-b0c2-40e4839c881b · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Concept bottleneck models,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0446b9f3-4703-4683-824f-d11d1cd78ded · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Label-free concept bottleneck models,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dfd76a3-7016-4a37-aa6a-0b1305a3bc06 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Post-hoc concept bottleneck models,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb7b42c9-3a10-4abf-877f-d21751c8924c · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Transformation of audio embeddings into interpretable, concept-based representations,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ea3e1a-cb9b-41d3-8908-b33ac34e4a29 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Pre-trained vision-language models learn discoverable visual concepts,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f2570ed-f8f8-4d36-a6fa-40e79b8fc1c4 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diagnosing and rectifying vision models using language,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ad5a0f-563c-4a19-b6f4-9729f40ce458 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Towards understanding the modality gap in clip,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4233b3-6db7-402a-a92a-8511f24b90a4 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Connect, collapse, corrupt: Learning cross-modal tasks with uni-modal data,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73851fe5-8e83-4ff8-a1bc-dcea43511a14 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Closing the modality gap for mixed modality search,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef71513-dca9-4e70-b460-07b04c34407d · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Explaining and mitigating the modality gap in contrastive multimodal learning,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b7f815-7f4b-4b2c-a518-66246661df5c · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Closing the modality gap enables novel multimodal learning applications,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a49361d7-a67b-47cd-b02d-e9b55aecd243 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Accept the modality gap: An exploration in the hyperbolic space,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095f858a-977d-41b5-a884-82c73fa1597d · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 000e53ae-7bb6-4eb8-9572-d56466ec55b3 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Understanding contrastive representation learning through alignment and uniformity on the hypersphere,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18969ec8-56ac-4c81-9eb2-70180cf56ebb · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings I0t: Embedding standard- ization method towards zero modality gap,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a659d79-3bd4-4f65-b54e-bbccc4a8234f · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Mitigate the gap: Improving cross-modal alignment in clip,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91826cf2-389b-49cc-816c-6b9b2729b342 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diffusion bridge: leveraging diffusion model to reduce the modality gap between text and vision for zero-shot image captioning,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37c0ad62-10ac-41fb-9714-702e18ca2f81 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Text-only training for image captioning using noise-injected clip,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddfa88a2-fb46-4d02-a494-89b72c930154 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings I can’t believe there’s no images! learning visual tasks using only language supervision,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5e0f4e-f693-49de-86aa-7aaa162dd587 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Zero-shot audio cap- tioning with audio-language model guidance and audio context keywords,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747cfef5-1816-4fc2-b7b5-0c2c4ff288a2 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f36373dd-d80d-4431-9350-ee07b5b97023 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Zero-Shot Audio Captioning via Audibility Guidance
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 4c7c599f-ad27-4bd3-8e18-a15e5170f782 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Clotho: An audio caption- ing dataset,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4faa5d1c-f4fd-4235-9302-be6dcdd01529 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Audiocaps: Generating captions for audios in the wild,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f0c301b-1a3e-4849-9cbf-fb22df780c68 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Sound- vecaps: Improving audio generation with visually enhanced captions,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0b85f9-6c33-4ccc-a111-c3e92e7af6ba · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Language models are unsupervised multitask learners,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a35603-cffa-40f4-bbc7-efca01fa73aa · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings LoRA: Low-rank adaptation of large language models,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0512697c-fceb-4c5c-9d30-cff754efb79b · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Judging llm-as-a-judge with mt-bench and chatbot arena,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476834c8-0eb4-4527-8896-c91a0022b29f · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.
Observation 3e8b5a78-a490-4e51-b9b0-446119adac02 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Minimum bayes-risk decoding for statistical machine translation,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04e30bf6-da41-4fe0-8399-f93dcc6b0ece · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 014ba787-2807-48e7-80a7-fedf2e8a7d72 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings BLEU: a method for automatic evaluation of machine translation,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b05a7edc-830e-44c8-8514-07331e3ac093 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings METEOR: An automatic metric for mt evalu- ation with improved correlation with human judgments,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18ffae39-cf05-409b-8be9-8f3b17fd2bd0 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings ROUGE: A package for automatic evaluation of summaries,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36c29d42-09ff-496f-8ed6-0f3cbd3e2dd2 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings CIDEr: Consensus- based image description evaluation,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee97865-41ce-49c8-b69c-d956dca43a94 · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Spice: Semantic propositional image caption evaluation,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85e1deeb-d75c-4cb7-ad86-e0adf153b31e · outbound
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Improved image captioning via policy gradient optimization of spider,
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.