Pith. sign in

Paper Citation Record · LEDGER

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2505.14562.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14562 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:15.094780Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.551896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:41:14.663492Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99e7f0db-e75a-41eb-b099-380885e3ee43 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning transferable visual models from natural language supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.549401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:13.452064Z digest=sha256:1e11e14b05754ece4fa9a1ca1491f61603c7e7c7492975ef847f1f4995c86bd0

Observation 2b93de3d-88e2-43dd-a87b-ba987bbb8709 · outbound

This paper cites CLAP learning audio concepts from natural lan- guage supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities CLAP learning audio concepts from natural lan- guage supervision,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.408674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:13.496702Z digest=sha256:26d95caa907a9fddf4490000466d7060b87efd69ac66886b19c8be9412bb37ca

Observation fabf8d23-e5c5-46de-8acd-fa912c52bd6b · outbound

This paper cites Multimodal learn- ing with deep boltzmann machines,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Multimodal learn- ing with deep boltzmann machines,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.235540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:13.595143Z digest=sha256:2898cbc24a9c367afbd4d37db4a90f8ea0b29620635b4f160ab354348149687b

Observation e96eb8f0-4007-4a5a-aa42-9441bc759fe2 · outbound

This paper cites Learning visual features from large weakly supervised data,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning visual features from large weakly supervised data,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.052093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:13.679426Z digest=sha256:f752353e4b35b3f88974f2ba2973631f7a9881776a0bdabbaa5459960f5ae908

Observation db6b673f-9dbc-45ab-9532-31d92af84d9d · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.912027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:13.759296Z digest=sha256:50421ecd428fef2416a8eb499312ef4eda56b100c4a33b6dd883ad7b65f5ab47

Observation 32e3baa8-88bc-4130-b9d7-aaa8079621cf · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Florence: A New Foundation Model for Computer Vision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:13.877031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:13.877031Z digest=sha256:b39ba9bc1f8031257ca1283a6582257d9cf0ca740df4c235c7bb1ee279cc4b07

Observation bb57bd20-a237-4aa2-84a2-75e9ed2c7cdd · outbound

This paper cites Wav2CLIP: Learning robust audio representa- tions from CLIP,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Wav2CLIP: Learning robust audio representa- tions from CLIP,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.768700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:13.958279Z digest=sha256:01c9e4dc54397dde91ae1d5258fbea5f112c8d74086b49d077f946fc36ce9181

Observation 396192f0-8fa6-4f63-a623-313544346d70 · outbound

This paper cites Microsoft COCO: Common objects in context,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Microsoft COCO: Common objects in context,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.591992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:14.079730Z digest=sha256:1fdd703671dc74a1eefc5071f77be4a60ec09aa3050a98189f67470cd7683351

Observation 64a99d06-0983-431a-90d9-1fd6fb60e89d · outbound

This paper cites Visual genome: Connecting lan- guage and vision using crowdsourced dense image annotations,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Visual genome: Connecting lan- guage and vision using crowdsourced dense image annotations,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.430857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:14.223807Z digest=sha256:dc75f4a0a2d6fedb64937a8816ee27d7ed16ba81a817b0096ec5d6299fd6f687

Observation 07f760c5-d9c6-4ada-86c1-3a2436579f30 · outbound

This paper cites YFCC100M: The new data in multimedia research,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities YFCC100M: The new data in multimedia research,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.243583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:14.325160Z digest=sha256:c0a6dc227edaabd5d22a50dd274b241cd06b47928bf865c6d2de63a6a9f373ea

Observation 9f471ed8-3cb0-4790-a9ae-65f3e2d60f93 · outbound

This paper cites FSD50K: An open dataset of human-labeled sound events,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities FSD50K: An open dataset of human-labeled sound events,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.106720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:14.486000Z digest=sha256:f85191dd2a8239a773cbbe6db0ae0f5c59a3f15ba509b3030dcb44d10d7a39fd

Observation 68537183-119f-44b4-955e-a282dceffadf · outbound

This paper cites Clotho: An audio captioning dataset,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Clotho: An audio captioning dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.929498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:14.547487Z digest=sha256:2e790543823931bf197e4eae388b3728c724358124c9c8d8e94e754a232fd3b3

Observation 67bd587a-e2c3-4fa3-9363-a1501bb6a625 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities AudioCaps: Generating captions for audios in the wild,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.814031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:14.665385Z digest=sha256:9094bd47d8eb841f8a44ac3b5688768506d90dc107731cf8b2cb4641a6d40ce5

Observation 0aeba42d-176e-414e-9584-807e916a9439 · outbound

This paper cites What is the ground truth? reliability of multi-annotator data for audio tag- ging,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities What is the ground truth? reliability of multi-annotator data for audio tag- ging,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.632294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:14.761764Z digest=sha256:453d316dd0b5a225170d68b6845d31314f5d85127905ce9009ec65050c9aae82

Observation d814e34b-9f97-4733-98df-56ee70e7696b · outbound

This paper cites Audio- CLIP: Extending CLIP to image, text and audio,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio- CLIP: Extending CLIP to image, text and audio,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.484053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:14.816806Z digest=sha256:84554262444400798d3ba78e0835e306db4557c750c90481844f9358a9009a9b

Observation 2ece6d4f-4a63-40ea-a543-ceba1cabcdaf · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio Set: An ontology and human-labeled dataset for audio events,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.380850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:14.933049Z digest=sha256:acb8865a4e909f023410a16421f74ee6bc5dbf5de453b9dd8446b9029193a6f0

Observation 3dbc9405-fb1f-4ae3-a221-602119900054 · outbound

This paper cites Sudarsanam, I.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Sudarsanam, I

Reference 17

Resolution
verified exact
doi, observed 2026-08-07T15:36:15.236385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:36:15.021689Z digest=sha256:9717e5189d1b8afb8d283feaad4abc28145aa2581043cf15ef079f8798a479d6

Observation 7a0b3938-31fc-49ce-ae4f-2a5ff4c5c74e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Representation Learning with Contrastive Predictive Coding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:15.094780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:15.094780Z digest=sha256:773d9869d040448bf9bfe4dc6533508155c5ba9cb196b36cbc93612a4c5ead98

Pith citing papers

Observation 367aff17-6b04-4fb6-9b8f-f32f317f4eed · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.551896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.551896Z digest=sha256:097c7e60dad49e6f7eff44229a8c49601095c38d129f5c90b4aaf3e9be6c32e3

Observation 3301051a-947f-4ff9-90c4-2bd38cb9021d · inbound

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception cites this paper.

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:41:14.716289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:41:11.839628Z digest=sha256:05e9222e4c8175be7345d4f853d783e70615a25e5e88c08ea5c034547823484f