Pith. sign in

Paper Citation Record · LEDGER

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings

As of 25 July 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2605.29628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.29628 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T05:57:09.345415Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-24T06:31:00.690269+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf76a0b9-a4c2-41ac-a666-a51fe07e10ac · outbound

This paper cites Learning transferable visual models from natural language supervision,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Learning transferable visual models from natural language supervision,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:7e793ed89f9ca96ef05bf07ee2d5ad9e3973e11e28ecba68ccc7d69b154691a6

Observation bf90d5c1-6494-4fb1-9e32-91fb3086d426 · outbound

This paper cites Natural language supervision for general-purpose audio representations,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Natural language supervision for general-purpose audio representations,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:5f56496e138eff787ea8981210fd1545b851783b773f5f0ac6f6fbb94bc87dca

Observation ba783362-e1b2-421e-8a32-2b5c9326cd14 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:0c1b73eab43f1f81e421a460071ccdf4be96e9a614f38bbcea592685bb6bf75a

Observation 80c9bca2-1158-4312-ae47-5a8ac9944ab7 · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio- language multimodal research,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio- language multimodal research,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:a8eabcbd0742b6f57a3b291e3bda112abacb8c779d9356208b55757508842fa3

Observation 4005e1c8-976b-44d3-8a32-15db5fa0a721 · outbound

This paper cites Auto-acd: A large-scale dataset for audio-language representation learning,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Auto-acd: A large-scale dataset for audio-language representation learning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:b29f26f59d605b6f3f63b4add66a2dfeeaaa875b47eff5e1d917e86da2b30d9d

Observation 2bac5fcd-947d-4c25-85d2-cbb94a2516b2 · outbound

This paper cites Audiosetcaps: An enriched audio-caption dataset using automated gener- ation pipeline with large audio and language models,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Audiosetcaps: An enriched audio-caption dataset using automated gener- ation pipeline with large audio and language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:cb8d27b4b211af8ca0f4388aa2d74a7ee34eaa4fdd8c86c4a901ff33e25ea9e1

Observation a32f3683-0ed4-40b7-b62b-175adab839d7 · outbound

This paper cites A knowledge distillation approach to im- proving language-based audio retrieval models,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings A knowledge distillation approach to im- proving language-based audio retrieval models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:334861ff76601bc0666e3812cf7decad593d8e047fadf14ffef8c36e757902a1

Observation e90fdce7-0e45-4ea7-9ebc-b5ef48b75733 · outbound

This paper cites Aistat lab system for dcase 2025 task6: Language-based audio retreival,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Aistat lab system for dcase 2025 task6: Language-based audio retreival,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:f8123ce96597ccfc04cc538bf896cb79ff5ba13744624166d6fbc99f81892c00

Observation 12a2bdaf-b58f-4c48-9d74-14c602bb9587 · outbound

This paper cites Recap: Retrieval-augmented audio captioning,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Recap: Retrieval-augmented audio captioning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:f0ccf6c1ef74329c8ba7cbd463b14b9d3ec84f7eff4c83ea4e06f2e0f4108d9a

Observation cd402f97-9875-47d3-8617-653c8ab47446 · outbound

This paper cites Diffusion-based diverse audio captioning with retrieval-guided langevin dynamics,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diffusion-based diverse audio captioning with retrieval-guided langevin dynamics,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:edad0cbaf3c57589f60a6984ba64f55f0b8e136b8ff9bd60f90c1ba7360046a3

Observation 0a1c4316-182a-424f-8aea-467bf608a4c3 · outbound

This paper cites Zero-shot diverse audio captioning with diffusion models,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Zero-shot diverse audio captioning with diffusion models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:665a10afa5e8a68a61fef741ab21a91e97c6e6526f1923b2f1f8ba1d0543794d

Observation 9e10c007-48f9-4dbe-b198-a8bef43d222a · outbound

This paper cites Audioldm: text-to-audio generation with latent diffusion models,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Audioldm: text-to-audio generation with latent diffusion models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:5e027ea7712ab0a7aac2a3e111a559ea258c19423bb65065fc9224526eb1e63c

Observation ad1a90de-32d4-4aec-8ca0-ce3d60bac504 · outbound

This paper cites AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings AudioLDM 2: Learning holistic audio generation with self-supervised pretraining,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:d652e35031381cfa229f214340675d393a7865bf59ca82096941164df7750aea

Observation b5f6122b-f453-4d12-a221-57290e772d20 · outbound

This paper cites Weakly-supervised automated audio captioning via text only training,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Weakly-supervised automated audio captioning via text only training,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:d0cae3e4779bc4b24cb77cce95beed94c9a0bdb3a96425234e3a0dac2286d271

Observation 40c80b2f-542d-4ae2-b716-31f0052450ba · outbound

This paper cites Training audio captioning models without audio,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Training audio captioning models without audio,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:3d39d75d07f43fc34639a12830d05bc716c5635e543f04a21ba71564314fc85e

Observation 705d6885-3fa2-4e53-ba52-40dc94a1b6b3 · outbound

This paper cites Zero-shot audio captioning using soft and hard prompts,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Zero-shot audio captioning using soft and hard prompts,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:5fb2c04e726579d8c4d68123d48a0e127cc57adfa7d39b65af1a9cd8b5515e8d

Observation 517d37dc-d995-4b18-9050-d58d15f201f3 · outbound

This paper cites Drcap: Decoding clap latents with retrieval-augmented generation for zero-shot audio captioning,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Drcap: Decoding clap latents with retrieval-augmented generation for zero-shot audio captioning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:619202cb0ae72c030b5196059e89a50bbc149e8df129496c532165acd45a2c2b

Observation 3405abec-0aef-4e87-b326-398c3540b467 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:f51c88cfb1aff8150c3712a7f9d5f671ae08f4f2c506f94faac16fc30010c7c2

Observation 9467263a-50c7-4a30-a2af-b1f1ec2e94ef · outbound

This paper cites Two effects, one trigger: On the modality gap, object bias, and information imbalance in contrastive vision-language models,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Two effects, one trigger: On the modality gap, object bias, and information imbalance in contrastive vision-language models,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:32d3bae55cc883dc662744cbf460ef210fdd548fb9acc738eb9731a8e28bd2d5

Observation b121d8e5-0998-4b21-af9c-b7c55158ceb5 · outbound

This paper cites Decipher the modality gap in multimodal contrastive learning: From convergent representations to pairwise alignment.arXiv preprint arXiv:2510.03268.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Decipher the modality gap in multimodal contrastive learning: From convergent representations to pairwise alignment.arXiv preprint arXiv:2510.03268

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:03:08.636879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:283e5bef4bb82a0fda3af261757a4a3b3ece7e831a6fb46d51752cddb556e9a5

Observation d9f79375-a353-42c1-9e42-f2e43fc89d16 · outbound

This paper cites Decap: Decoding CLIP latents for zero-shot captioning via text-only training,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Decap: Decoding CLIP latents for zero-shot captioning via text-only training,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:54157735e8879893b6d5fc8eefd97f9f1d7de693b4fd9e71805f298e4b31a4dd

Observation f8bc2571-5742-47e3-b795-747b74dc22d2 · outbound

This paper cites Diffgap: A lightweight diffusion module in contrastive space for bridging cross-model gap,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diffgap: A lightweight diffusion module in contrastive space for bridging cross-model gap,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:2a02763b4b095b214a6589173cd98f1142e385a7998adbb307dac2e472418f94

Observation 6aebe9d8-22a5-4281-9766-c294568d773c · outbound

This paper cites Diffusion-link: Diffusion probabilistic model for bridging the audio-text modality gap,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diffusion-link: Diffusion probabilistic model for bridging the audio-text modality gap,

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:03:08.639371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:56af260cf3300c29133a0da14c1a4efe2a333ffff540c00537d62157a1746a55

Observation 2abe19dc-5c38-4b5c-afc7-fccac2fb3c0a · outbound

This paper cites Pull it together: Reducing the modality gap in contrastive learning,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Pull it together: Reducing the modality gap in contrastive learning,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:b1260d55c6a5bf3e671ed8b0f74e1a08273db6d901f56400e4c32daa918265e7

Observation 54a2e44b-1341-4541-8f30-8ad2955ac96b · outbound

This paper cites Interpreting clip with sparse linear concept embeddings (splice),.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Interpreting clip with sparse linear concept embeddings (splice),

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:17198d458deb7eff4ea3035d7aa8b16619b3f2e05a6a48639e5a2efc33987e80

Observation cd08f819-fcaa-467b-a8d8-3e96c3809afc · outbound

This paper cites Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:03:08.631498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:3681dd375d0010af6f899871d05162d4947d5a797c987d53e1306e5e75454b3f

Observation ad996e2e-43f7-4f05-ae86-3585194ea51f · outbound

This paper cites Scocca: Multi-modal sparse concept decomposition via canonical correlation analysis,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Scocca: Multi-modal sparse concept decomposition via canonical correlation analysis,

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:03:08.636518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:0aee1e0d5d6b82de3aa7214154df7877f4f7714334eb78d34b1f39f33781d2e8

Observation 4da3e290-4b67-44a2-89ca-d6f74da0be72 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Momentum contrast for unsupervised visual representation learning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:615be4a1749bdab1c5ef011bfbe2cd9e3d817da19d295e178b69154b6ce84f0e

Observation 475ea1a4-0854-46e7-8870-78bb50e2c591 · outbound

This paper cites Pengi: An audio language model for audio tasks,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Pengi: An audio language model for audio tasks,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:b307caad8fd6846bc48a79f1090846fb8db612e08e47d2448ea6fd98d9a76921

Observation b05a1972-270b-4202-9cb2-4fb9dafd1a4c · outbound

This paper cites RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T06:03:08.638708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:3194c6d39fec13307124002be54e7a271a347b30f57d5c2769384f7a8e1ebe96

Observation 833b8a30-19f8-4a9f-ba06-48d2ce7367a5 · outbound

This paper cites A holistic approach to unifying automatic concept extraction and concept importance estimation,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings A holistic approach to unifying automatic concept extraction and concept importance estimation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:d5380229e93b972469f4de25d38934a20f434b43d11b63dd02f216d4dc02e542

Observation d9cd8797-0d9f-4b94-a66d-ca639d40d7fe · outbound

This paper cites Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav),.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav),

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:6914623182bf2e34be94b204efbfbd26ac009fa0efe39a00ed4777d15976cccb

Observation 2f2fcb08-1e0e-4218-ac1f-fd56129ea48c · outbound

This paper cites Craft: Concept recursive activation factorization for explainability,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Craft: Concept recursive activation factorization for explainability,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:419452f9f10067694e9c359fea9f539a21efa4b203c1e12734955dc388b79138

Observation d7a63104-7bc1-4639-9236-52274dc23bee · outbound

This paper cites Multi-dimensional concept discovery (mcd): a unifying framework with completeness guarantees,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Multi-dimensional concept discovery (mcd): a unifying framework with completeness guarantees,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:5da0b6734072e6a6f132f049eae25bb6f3bb92083e90cfc0f2bdf30935cdabe6

Observation 26bfd95d-d137-4e49-b0c2-40e4839c881b · outbound

This paper cites Concept bottleneck models,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Concept bottleneck models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:67c296b50e8d3b1f42483a2b78e70804cadbbb50ae6487dcf96865dcb5827378

Observation 0446b9f3-4703-4683-824f-d11d1cd78ded · outbound

This paper cites Label-free concept bottleneck models,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Label-free concept bottleneck models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:dca033498b4e6059e91ca84b9c488feee35243875bd83b2c990b4e0e5b30d1c5

Observation 4dfd76a3-7016-4a37-aa6a-0b1305a3bc06 · outbound

This paper cites Post-hoc concept bottleneck models,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Post-hoc concept bottleneck models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:6dd35958651ae89dc3d8aa7d2d450c474b116ba578383b2c2d1aa34e1f62111e

Observation cb7b42c9-3a10-4abf-877f-d21751c8924c · outbound

This paper cites Transformation of audio embeddings into interpretable, concept-based representations,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Transformation of audio embeddings into interpretable, concept-based representations,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:c738c3429e5fc8850b9345f10101a7b8ed44f7bb677a95b086b906e1d6a59021

Observation 95ea3e1a-cb9b-41d3-8908-b33ac34e4a29 · outbound

This paper cites Pre-trained vision-language models learn discoverable visual concepts,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Pre-trained vision-language models learn discoverable visual concepts,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:c308b825dd4cd8ca02d9a3c5ab546fc078ecd71e31a5d5ee3cd02253bab11e47

Observation 5f2570ed-f8f8-4d36-a6fa-40e79b8fc1c4 · outbound

This paper cites Diagnosing and rectifying vision models using language,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diagnosing and rectifying vision models using language,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:0d86b80b9b09fde7f3c4f51a89205b25b47aac0cf36245386a342b5cbd22a0c0

Observation f2ad5a0f-563c-4a19-b6f4-9729f40ce458 · outbound

This paper cites Towards understanding the modality gap in clip,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Towards understanding the modality gap in clip,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:db53c386e9e37db11d6a7c9a94d7d66b66ad3166cedbfb3f3820e4a037091b15

Observation 9a4233b3-6db7-402a-a92a-8511f24b90a4 · outbound

This paper cites Connect, collapse, corrupt: Learning cross-modal tasks with uni-modal data,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Connect, collapse, corrupt: Learning cross-modal tasks with uni-modal data,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:7a18970da269e02de9819fac4e3f3b631a06fae140ce10c3c30b4b1ee59f13c0

Observation 73851fe5-8e83-4ff8-a1bc-dcea43511a14 · outbound

This paper cites Closing the modality gap for mixed modality search,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Closing the modality gap for mixed modality search,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:2dbc26089e91b5cb94b757b453e01787ea39eb561020c00bc05277f29cc094ca

Observation 3ef71513-dca9-4e70-b460-07b04c34407d · outbound

This paper cites Explaining and mitigating the modality gap in contrastive multimodal learning,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Explaining and mitigating the modality gap in contrastive multimodal learning,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:c47a83bde4ea98613c94615519b96d829ce7478b4b0cd9864be8bd56270e7ff0

Observation 23b7f815-7f4b-4b2c-a518-66246661df5c · outbound

This paper cites Closing the modality gap enables novel multimodal learning applications,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Closing the modality gap enables novel multimodal learning applications,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:c500404ca65baba0a2c5ee2dba138439a16b28bc6e0d807879d8503f9560a522

Observation a49361d7-a67b-47cd-b02d-e9b55aecd243 · outbound

This paper cites Accept the modality gap: An exploration in the hyperbolic space,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Accept the modality gap: An exploration in the hyperbolic space,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:ad02a8e4c0b482cb54946bd7a825ef9efbdff265b5cafadd43163b91146f27f6

Observation 095f858a-977d-41b5-a884-82c73fa1597d · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:03:08.641708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:0a1d79f7f76725eb7f2eec62c7f6a08caadbd21c0ee4fd2c067ccc861e92477b

Observation 000e53ae-7bb6-4eb8-9572-d56466ec55b3 · outbound

This paper cites Understanding contrastive representation learning through alignment and uniformity on the hypersphere,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Understanding contrastive representation learning through alignment and uniformity on the hypersphere,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:c6ec3683648645770e779b7fe9af76ed8a40ce6c9f0e8f88569da3f4ab35748f

Observation 18969ec8-56ac-4c81-9eb2-70180cf56ebb · outbound

This paper cites I0t: Embedding standard- ization method towards zero modality gap,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings I0t: Embedding standard- ization method towards zero modality gap,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:cdd01af1cdda200902fc282525de71df8354f74d4bce76f585ff28682c75235f

Observation 4a659d79-3bd4-4f65-b54e-bbccc4a8234f · outbound

This paper cites Mitigate the gap: Improving cross-modal alignment in clip,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Mitigate the gap: Improving cross-modal alignment in clip,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:ec9b604fd394b76beae98e2fdc372274e839ff127589c337b0e1aa2e928c4ce4

Observation 91826cf2-389b-49cc-816c-6b9b2729b342 · outbound

This paper cites Diffusion bridge: leveraging diffusion model to reduce the modality gap between text and vision for zero-shot image captioning,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Diffusion bridge: leveraging diffusion model to reduce the modality gap between text and vision for zero-shot image captioning,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:92c748849deb8d8e00a1a5ea99a8e04fd6a6d69505942cb00cc38a896e034787

Observation 37c0ad62-10ac-41fb-9714-702e18ca2f81 · outbound

This paper cites Text-only training for image captioning using noise-injected clip,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Text-only training for image captioning using noise-injected clip,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:b87cc435cb5bb740bfb8bb87eb76b097ccd15dd4b137a4f2869ea055321d951a

Observation ddfa88a2-fb46-4d02-a494-89b72c930154 · outbound

This paper cites I can’t believe there’s no images! learning visual tasks using only language supervision,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings I can’t believe there’s no images! learning visual tasks using only language supervision,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:6cc775943ab434a24bcf0a435d391d70bf4e2c1011a77a9ddda1c5558eb9ee45

Observation 6d5e0f4e-f693-49de-86aa-7aaa162dd587 · outbound

This paper cites Zero-shot audio cap- tioning with audio-language model guidance and audio context keywords,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Zero-shot audio cap- tioning with audio-language model guidance and audio context keywords,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:6aa17f44c01427a3f12541be2597dc84b7b3a419b2b724dc3afcf9bd6b55884d

Observation 747cfef5-1816-4fc2-b7b5-0c2c4ff288a2 · outbound

This paper cites an unresolved cited work.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:ce44a6f6b5427045af5641b871364a2dc2e3d10bff146ada52a57c923268cfea

Observation f36373dd-d80d-4431-9350-ee07b5b97023 · outbound

This paper cites Zero-Shot Audio Captioning via Audibility Guidance.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Zero-Shot Audio Captioning via Audibility Guidance

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:03:08.644737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:50cdc0dd97511293b0c8b953f42c281480613e76dbc14fc17f7b44fb13953dc3

Observation 4c7c599f-ad27-4bd3-8e18-a15e5170f782 · outbound

This paper cites Clotho: An audio caption- ing dataset,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Clotho: An audio caption- ing dataset,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:2f430516f25afdd3ae9434b81393a8853db3a983fe59bc7cce759a1bb0a32499

Observation 4faa5d1c-f4fd-4235-9302-be6dcdd01529 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Audiocaps: Generating captions for audios in the wild,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:77a6ffdc7eecfff676fbf71a43fbde5802f55ec3be499927f8aac5d6e2e6f363

Observation 6f0c301b-1a3e-4849-9cbf-fb22df780c68 · outbound

This paper cites Sound- vecaps: Improving audio generation with visually enhanced captions,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Sound- vecaps: Improving audio generation with visually enhanced captions,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:54809c3d25fbe692e398b2b35ae766db5b64f4089c5b6aa63bf041910f6bd744

Observation 6b0b85f9-6c33-4ccc-a111-c3e92e7af6ba · outbound

This paper cites Language models are unsupervised multitask learners,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Language models are unsupervised multitask learners,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:89191539f0117cee69a51d8c13359c025385b4e40bcb9591f36caca9aae8d8d7

Observation 34a35603-cffa-40f4-bbc7-efca01fa73aa · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings LoRA: Low-rank adaptation of large language models,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:ff61a688ff2005be57dde1eb4ae89df26170cf0c95f449aeeaad00d16dc907f1

Observation 0512697c-fceb-4c5c-9d30-cff754efb79b · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:f58e64f8245b7d5e93af5a9ecd7b6868cf0a4ab26ebc3d4faf14775aade647da

Observation 476834c8-0eb4-4527-8896-c91a0022b29f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-06-29T06:03:08.647152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:eff0e05a899b541bf8a9cf55b9857cb55494f269203a94cb5305477794a31f83

Observation 3e8b5a78-a490-4e51-b9b0-446119adac02 · outbound

This paper cites Minimum bayes-risk decoding for statistical machine translation,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Minimum bayes-risk decoding for statistical machine translation,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:4a5ea24fd966194422918f949cbabc861ceaa978b9048a8f70caca0cf902e8c1

Observation 04e30bf6-da41-4fe0-8399-f93dcc6b0ece · outbound

This paper cites HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:76a776c39058c60015a1587aca64fdc27e5a4e64be669489daa9bb720e6b336f

Observation 014ba787-2807-48e7-80a7-fedf2e8a7d72 · outbound

This paper cites BLEU: a method for automatic evaluation of machine translation,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings BLEU: a method for automatic evaluation of machine translation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:263b8e94cf8778604bf1a9a750149caa36cbbb4ae63848e9abb591ae1c680ecf

Observation b05a7edc-830e-44c8-8514-07331e3ac093 · outbound

This paper cites METEOR: An automatic metric for mt evalu- ation with improved correlation with human judgments,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings METEOR: An automatic metric for mt evalu- ation with improved correlation with human judgments,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:194d44ef16e707ebab9ae37ca958629747a32f538105cbb14c92c4eaebbc8ad4

Observation 18ffae39-cf05-409b-8be9-8f3b17fd2bd0 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings ROUGE: A package for automatic evaluation of summaries,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:dd285e154eccc9898b9d54ba658b9c5f59a7a4f5dd4370cb1dbd16cdcec289b1

Observation 36c29d42-09ff-496f-8ed6-0f3cbd3e2dd2 · outbound

This paper cites CIDEr: Consensus- based image description evaluation,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings CIDEr: Consensus- based image description evaluation,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:007aa385c36ea0699d2bf55f28ed1c212c67fab398b6fe71e3561431ccd614a1

Observation 7ee97865-41ce-49c8-b69c-d956dca43a94 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Spice: Semantic propositional image caption evaluation,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:19b4533284ad0c5eb63cbbe875cdda0992509db6d5b58945176448199b3c85bd

Observation 85e1deeb-d75c-4cb7-ad86-e0adf153b31e · outbound

This paper cites Improved image captioning via policy gradient optimization of spider,.

COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings Improved image captioning via policy gradient optimization of spider,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-29T05:57:09.345415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T05:57:09.345415Z digest=sha256:8c9f2ddd8ac538bbcea529dddf416e2db20a8b7b2252829ebe805dc5ccfab297

Pith citing papers

No inbound Pith citation observations are available.