Pith. sign in

Paper Citation Record · LEDGER

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 2 inbound Pith citation observations for arXiv:2505.14562.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14562 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:15.094780Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.551896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:41:14.663492Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99e7f0db-e75a-41eb-b099-380885e3ee43 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning transferable visual models from natural language supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.549401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:13.452064Z digest=sha256:3c89799f104fdeecc3eb43f06898e34badd2b0be0d01d0d3e066c12d240ebd2e

Observation 2b93de3d-88e2-43dd-a87b-ba987bbb8709 · outbound

This paper cites CLAP learning audio concepts from natural lan- guage supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities CLAP learning audio concepts from natural lan- guage supervision,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.408674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:13.496702Z digest=sha256:7f9dbbddf7f4a1650f70ad734059788bd01d3ae5228bd277c4f1f18d8b29e945

Observation fabf8d23-e5c5-46de-8acd-fa912c52bd6b · outbound

This paper cites Multimodal learn- ing with deep boltzmann machines,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Multimodal learn- ing with deep boltzmann machines,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.235540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:13.595143Z digest=sha256:3a082914056f9cf45e11c6a7c9ff5ff5f3b4e99b1b9bb3567f333d78b8f0ef98

Observation e96eb8f0-4007-4a5a-aa42-9441bc759fe2 · outbound

This paper cites Learning visual features from large weakly supervised data,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Learning visual features from large weakly supervised data,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:17.052093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:13.679426Z digest=sha256:2266a91346998c07fe308159a717ad995ad4deb27ae86b015957bb36faa1a0d8

Observation db6b673f-9dbc-45ab-9532-31d92af84d9d · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.912027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:13.759296Z digest=sha256:2edb6b1130b2da312f02789704b47f7346bb0dba238e0016f5501b65ba990ffb

Observation 32e3baa8-88bc-4130-b9d7-aaa8079621cf · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Florence: A New Foundation Model for Computer Vision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:13.877031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:13.877031Z digest=sha256:b39ba9bc1f8031257ca1283a6582257d9cf0ca740df4c235c7bb1ee279cc4b07

Observation bb57bd20-a237-4aa2-84a2-75e9ed2c7cdd · outbound

This paper cites Wav2CLIP: Learning robust audio representa- tions from CLIP,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Wav2CLIP: Learning robust audio representa- tions from CLIP,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.768700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:13.958279Z digest=sha256:061b6d261ff3235e770bf870d67b8ab32da4550cc6ed6157b09f007418bcefeb

Observation 396192f0-8fa6-4f63-a623-313544346d70 · outbound

This paper cites Microsoft COCO: Common objects in context,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Microsoft COCO: Common objects in context,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.591992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:14.079730Z digest=sha256:5acafb6030c60fcd4f34b28b074698e730322405ac6b1850ab1ff2538510a20c

Observation 64a99d06-0983-431a-90d9-1fd6fb60e89d · outbound

This paper cites Visual genome: Connecting lan- guage and vision using crowdsourced dense image annotations,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Visual genome: Connecting lan- guage and vision using crowdsourced dense image annotations,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.430857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:14.223807Z digest=sha256:24a4a032bb4655dcb410c87b4a1d8a0ecb1181d126440de67ff21e0668e5f769

Observation 07f760c5-d9c6-4ada-86c1-3a2436579f30 · outbound

This paper cites YFCC100M: The new data in multimedia research,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities YFCC100M: The new data in multimedia research,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.243583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:14.325160Z digest=sha256:1eb45502b22b8d78903401fd668adc389c8a203e15296c9c0a5caf14572a8e6f

Observation 9f471ed8-3cb0-4790-a9ae-65f3e2d60f93 · outbound

This paper cites FSD50K: An open dataset of human-labeled sound events,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities FSD50K: An open dataset of human-labeled sound events,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:16.106720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:14.486000Z digest=sha256:eb890477c65cb32d6c794e1bd1518b4debca976990bac23ef5b777bf746cb515

Observation 68537183-119f-44b4-955e-a282dceffadf · outbound

This paper cites Clotho: An audio captioning dataset,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Clotho: An audio captioning dataset,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.929498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:14.547487Z digest=sha256:f4e62ead56b577728526ea8f3bd55df473bd2c76f3bad1562b3507a95e621960

Observation 67bd587a-e2c3-4fa3-9363-a1501bb6a625 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities AudioCaps: Generating captions for audios in the wild,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.814031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:14.665385Z digest=sha256:46ff1a77ec7d5d5d356635f7bc8398ef1f88fbea4464c74801e1ebe9d53c6260

Observation 0aeba42d-176e-414e-9584-807e916a9439 · outbound

This paper cites What is the ground truth? reliability of multi-annotator data for audio tag- ging,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities What is the ground truth? reliability of multi-annotator data for audio tag- ging,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.632294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:14.761764Z digest=sha256:40fa09af004ead2ddd9a3359cdfa69cdd94c9763c243940b53c05557131754a5

Observation d814e34b-9f97-4733-98df-56ee70e7696b · outbound

This paper cites Audio- CLIP: Extending CLIP to image, text and audio,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio- CLIP: Extending CLIP to image, text and audio,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.484053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:14.816806Z digest=sha256:8f194abf62f2ccf6a73178a08f8f367365f3026de27c36dd46647bee77442564

Observation 2ece6d4f-4a63-40ea-a543-ceba1cabcdaf · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Audio Set: An ontology and human-labeled dataset for audio events,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:15.380850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:14.933049Z digest=sha256:88bfd25c31051a010ed459d2818ca94aaec7f2664e782bee705a55f408db646c

Observation 3dbc9405-fb1f-4ae3-a221-602119900054 · outbound

This paper cites Sudarsanam, I.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Sudarsanam, I

Reference 17

Resolution
verified exact
doi, observed 2026-08-07T15:36:15.236385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:36:15.021689Z digest=sha256:e0ef75bb23d29824eaf92af34b4b1d46d831d53b629852b8c2dda6a427e317ae

Observation 7a0b3938-31fc-49ce-ae4f-2a5ff4c5c74e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities Representation Learning with Contrastive Predictive Coding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:15.094780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:15.094780Z digest=sha256:773d9869d040448bf9bfe4dc6533508155c5ba9cb196b36cbc93612a4c5ead98

Pith citing papers

Observation 367aff17-6b04-4fb6-9b8f-f32f317f4eed · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.551896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.551896Z digest=sha256:86193018d104b5ceb97145d39c5e30e352ce2882ebc01a13d274033b56f6669c

Observation 3301051a-947f-4ff9-90c4-2bd38cb9021d · inbound

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception cites this paper.

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:41:14.716289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:41:11.839628Z digest=sha256:4449589d2b829c6a1a771eb2de0945e90e3cec3dde14ce617ce1d7f7e29b75ad