Pith. sign in

Paper Citation Record · LEDGER

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval

As of 14 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2412.13071.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13071 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:30:38.344969Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:35:56.585075Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T23:35:57.101917Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy37
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 747e71d5-249e-4927-85c5-e98aef9b494a · outbound

This paper cites Multimodal learning and reasoning,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Multimodal learning and reasoning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.373544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.071595Z digest=sha256:4280f159bf9e596fa124f5614c6637647c0f0d2c9991b74b8b9d949617177a18

Observation 3452085f-9aab-4c45-9191-f92a0e913ed2 · outbound

This paper cites Multimodal machine learning: Integrating language, vision and speech,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Multimodal machine learning: Integrating language, vision and speech,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.355785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.077910Z digest=sha256:23d344d73348b06d056ac2fc73c05238f1b0f880c76b5356e819beeb7e738925

Observation 14e9557a-17bb-4e54-9ba4-ce74251acce4 · outbound

This paper cites Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.338312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.083838Z digest=sha256:e2f259ee3790bb69a93c356b3a084347d81d73af26a2596908e5901223552687

Observation f2450719-807a-4864-a9f3-2f5e71fb6ca4 · outbound

This paper cites Multilingual multimodal pre-training for zero-shot cross-lingual transfer of vision-language models,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Multilingual multimodal pre-training for zero-shot cross-lingual transfer of vision-language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.320120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.090251Z digest=sha256:da2bce5a24785ceb4bd9880a7625841531c74febd203328e20997841b6fc3d39

Observation 9257c914-18b8-4033-bb4f-afb9852592b1 · outbound

This paper cites Few-shot joint multimodal aspect-sentiment analysis based on generative multimodal prompt,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Few-shot joint multimodal aspect-sentiment analysis based on generative multimodal prompt,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.298939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.096466Z digest=sha256:a4bbcc79000cb8414c612be647135e7fd6705d450b6a7b4a8917e90f8c0faca1

Observation e8d3e159-1e6a-41d9-8b49-f46aeae424ae · outbound

This paper cites MuRAG: Multimodal retrieval-augmented generator for open question answering over images and text,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval MuRAG: Multimodal retrieval-augmented generator for open question answering over images and text,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.279754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.102926Z digest=sha256:fe41814b53f80ae05b9b2ab535b02f71ab36d9b7cd502d535348eff28b7f96fd

Observation 05379c55-8c5f-4578-bd4e-5d46e9359ef2 · outbound

This paper cites Large-scale multi-modal pre-trained models: A comprehensive survey,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Large-scale multi-modal pre-trained models: A comprehensive survey,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.257869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.109912Z digest=sha256:a4f65d7e0aa7c7d2fc549a5e89ccd86f76cae33e01ab40b501c66fc3be58bd71

Observation d55188bb-c7e4-4bbe-ae5c-35ef79a49317 · outbound

This paper cites Must-c: A multilingual corpus for end-to-end speech translation,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Must-c: A multilingual corpus for end-to-end speech translation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.239274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.115258Z digest=sha256:9c7d3297c7c411181f98933d132ec322899eab90cf09e460f2bfee8b9cb79a1d

Observation e61a75c9-20eb-4c5f-976b-3b45962e5078 · outbound

This paper cites Reformulating information retrieval from speech and text as a detection problem,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Reformulating information retrieval from speech and text as a detection problem,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.221034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.121205Z digest=sha256:326749a55efbbc313092228c2e501d35f4ab5f8e1686645d16ef5aa58f8918f7

Observation af3d3df8-b2d7-4e27-b62d-c36ec0c13efa · outbound

This paper cites Abstractive spoken document summarization using hierarchical model with multi-stage attention diversity optimization.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Abstractive spoken document summarization using hierarchical model with multi-stage attention diversity optimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.202726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.126875Z digest=sha256:a2122194e78334844da3d27ceed4ce0ef6d8a54d1dc7b7dba09fc4090ea401a2

Observation c3c32554-fa15-4a13-81d6-77cb1c43015e · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:38.132954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:38.132954Z digest=sha256:b52ee42644fcc775b66c6d08b294d5c5d93db73e9ce76f669ff85157dfc0cf9c

Observation b7f7b294-a33f-48cf-b462-86cabe3b5111 · outbound

This paper cites Self-supervised speech representation learning: A review,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Self-supervised speech representation learning: A review,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.184198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.138537Z digest=sha256:0f40b675927a1b8b7ee8f81e888e3c157b5b9409225f3ecf9d436b9a0070b91e

Observation 214c3fd3-77c1-44df-9c38-c6919b1484f7 · outbound

This paper cites Audio self-supervised learning: A survey,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Audio self-supervised learning: A survey,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:38.143898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:38.143898Z digest=sha256:8ebbd8a6e41df1a497c552f1309d02a20abc38b6a4e897f1eff24c0d5377b1d5

Observation fa1bceb3-40eb-4abd-aee5-490eb1ac5150 · outbound

This paper cites Cross-modal contrastive learning for speech translation,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Cross-modal contrastive learning for speech translation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.151569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.149465Z digest=sha256:c789aa9761739ec0c9999c602652cf1043ddec41ebb929cc1a35457d3855a153

Observation f97b62ae-c145-459a-923e-04ac79a10c54 · outbound

This paper cites Comparison of pretrained embeddings to identify hate speech in indian code-mixed text,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Comparison of pretrained embeddings to identify hate speech in indian code-mixed text,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.130162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.155951Z digest=sha256:6d1e7250583e5322daff0af076f3b7362e3609ccf0de8f014e49a1412bdd624d

Observation fff67f02-4775-4e6a-bf8a-e707542b4e05 · outbound

This paper cites Speech emotion recognition with multi-task learning.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Speech emotion recognition with multi-task learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.108765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.162395Z digest=sha256:73cdaa589afcb33e872db9f237c519bf6d33f1b7a31a40ac353ce6d0d00a2926

Observation 42bdf282-c2a8-4ff6-9c61-d33b71b47dd1 · outbound

This paper cites Adaptive knowledge distillation between text and speech pre-trained models,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Adaptive knowledge distillation between text and speech pre-trained models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.085980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.167782Z digest=sha256:97ea2cec465b7dbe25ea9ebd3e39945c933afdb62df5c138f76ee5e6b8067972

Observation 372ad29b-fb2e-40bc-a37f-1dcb750b28e7 · outbound

This paper cites Look&listen: Multi-modal correlation learning for active speaker detection and speech enhancement,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Look&listen: Multi-modal correlation learning for active speaker detection and speech enhancement,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.067471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.173001Z digest=sha256:c8611648777f5bbdfd08c700398649d6a78dda4e44f4dd88f109c08df56e9561

Observation 80b6d3a5-8e08-4ac5-bdc2-14c6643b1ec1 · outbound

This paper cites Multi-modal speech emotion recognition using self-attention mechanism and multi-scale fusion framework,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Multi-modal speech emotion recognition using self-attention mechanism and multi-scale fusion framework,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.042117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.179254Z digest=sha256:3ae8e4d0e51c85bfee47c9f6da4adc64c55ca15f23e0a4324dfe60620678aa8e

Observation 670923cf-b514-4192-b39b-ffdb633d8429 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Learning transferable visual models from natural language supervision,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:38.184360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:38.184360Z digest=sha256:00769dfe5d45e7371269fdb321a896061371df8963aad47a6954775d19a1d741

Observation c058ee1a-fee0-4cad-96e4-515fc7b547b7 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Flamingo: a visual language model for few-shot learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:39.009654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.189305Z digest=sha256:f66e87c15fe19ead9f30bfd5f24b470daa334755c4a1e3b135d3fe5a780832fe

Observation 243c1be6-0e9e-4526-b9dc-e21422a6d307 · outbound

This paper cites Optimizing alignment of speech and language latent spaces for end-to-end speech recognition and understanding,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Optimizing alignment of speech and language latent spaces for end-to-end speech recognition and understanding,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.991048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.194970Z digest=sha256:adc96c43b1292cc03e0e7f14ed8f797d9d07677ca6eadb5a21c445dafd72c1e3

Observation 696a6ad0-ed42-421a-8db0-c3ad73b2f7e7 · outbound

This paper cites Samu-xlsr: Semantically-aligned multimodal utterance- level cross-lingual speech representation,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Samu-xlsr: Semantically-aligned multimodal utterance- level cross-lingual speech representation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.972386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.200231Z digest=sha256:18214a666400849157ced3d2f1d8b7028cd616455a729230f993092198a881db

Observation 05f3944a-4f52-4eb0-af3a-647e9aa7f040 · outbound

This paper cites MAESTRO: Matched Speech Text Representations through Modality Matching,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval MAESTRO: Matched Speech Text Representations through Modality Matching,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.951810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.207943Z digest=sha256:b5827007c3d23e27a365271775b65ff283b38cfe5a6031d0424e2b0513232383

Observation bdb2fe73-7651-4345-93a7-fdd0372d791c · outbound

This paper cites Unsupervised cross-modal alignment of speech and text embedding spaces,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Unsupervised cross-modal alignment of speech and text embedding spaces,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.931698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.214480Z digest=sha256:1b4964b2b061226f932ec6ff619c838bbfb6530931d7e81258c4df99ccd79bb0

Observation ebe8e8d6-2486-4f62-a81d-206358c03a4e · outbound

This paper cites SpeechUT: Bridging speech and text with hidden-unit for encoder-decoder based speech-text pre-training,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval SpeechUT: Bridging speech and text with hidden-unit for encoder-decoder based speech-text pre-training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.909334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.219525Z digest=sha256:26e4ea2d5a63d0fd1c147d2f89e971e9dee4a7ded974ea360cb42bfa534306a3

Observation 574785d2-9648-493f-8425-d793d19d70c5 · outbound

This paper cites SpeechBERT: An Audio-and-Text Jointly Learned Language Model for End-to-End Spoken Question Answering,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval SpeechBERT: An Audio-and-Text Jointly Learned Language Model for End-to-End Spoken Question Answering,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.888274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.226348Z digest=sha256:1ef5de9b234a62e64bb2095869570d9c303434472add88fe5a613b14d97a7f09

Observation c33a2b1c-2e04-44e4-979b-b3a3c024c060 · outbound

This paper cites Speechclip: Inte- grating speech with pre-trained vision and language model,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Speechclip: Inte- grating speech with pre-trained vision and language model,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.870328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.235302Z digest=sha256:c6d5b524f69ed57765f3ca67d3415c38fde76a86ad94bc5fa036c151f9bfe698

Observation c3494941-6bd2-4839-8a03-c36a5b8ba767 · outbound

This paper cites Acoustic Neighbor Embeddings.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Acoustic Neighbor Embeddings

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:30:38.432722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.243634Z digest=sha256:d81ec5de71226c2f8b8db8f4a5f32c5583edd9d00e810b6f1566ed48cf25c5e9

Observation db6fea38-e2a1-4fe4-93e1-40bc0a4499a1 · outbound

This paper cites On generative spoken language modeling from raw audio,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval On generative spoken language modeling from raw audio,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.850792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.252818Z digest=sha256:5e9de7483ddcf60f41b3b8395300a9a0cb3bc9af1286a09faaf243e1824bedbc

Observation b160593c-5e0f-47ae-9d81-4ac3f32638ad · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Common voice: A massively-multilingual speech corpus,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:38.265190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:38.265190Z digest=sha256:52de72e6ad0b76997dec7e0fd85efc472964db5ba7ff2a21540120c8c1aacf92

Observation 6510c1d6-04e5-4723-befd-8dc0545aded2 · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Fleurs: Few-shot learning evaluation of universal representations of speech,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.786786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.271364Z digest=sha256:7b55fa9b829d5bbfadd16bcebee1871c97daaf716ad22969a77636c2eb50e24c

Observation d6dc24fb-e7b8-432f-ae69-3cab0d18cf9b · outbound

This paper cites Brown corpus manual,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Brown corpus manual,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.763161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.277201Z digest=sha256:ecc030aca3a09d4fe245d8b8c9620e1ab2d35f1ecca5ca859929e6c130ce58c3

Observation d99b2fbd-0c31-4076-b380-df43353294a3 · outbound

This paper cites Extracting and evaluating general world knowledge from the brown corpus,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Extracting and evaluating general world knowledge from the brown corpus,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.739669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.282702Z digest=sha256:c9aa1b7d5a969ec9a1c7af3a83183cea53a54301cab8062bc79eed55c1832165

Observation 8de4d2bc-682a-4287-a363-f381b900766e · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.714064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.289711Z digest=sha256:5fb08320591a85ad72642733b0a1f9f4ca9de5e50070aa956ce6e5251d5fc7be

Observation a507930a-af68-42df-a495-30a822481e4a · outbound

This paper cites Frodo: From detections to 3d objects,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Frodo: From detections to 3d objects,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.691899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.295516Z digest=sha256:b8d482b401c10f26e615539ced958455ebffb8761873f81e9853492af546ea17

Observation 7b29a640-51fe-4b1c-855b-6394210f1893 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:38.302130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:38.302130Z digest=sha256:86a5078b2e55aa37b464070dc639d998683b7292180e81665a19c963596e9f60

Observation ce5ef1fd-c179-43a0-9624-31194d93c5f0 · outbound

This paper cites EfficientNet: Rethinking model scaling for convolutional neural networks,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval EfficientNet: Rethinking model scaling for convolutional neural networks,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.647187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.308127Z digest=sha256:7e44a68bbe83784c8d2677768a5ec51c4301c2d9ca78517282abdb50a844623b

Observation 22c594d5-9bde-4e0a-975c-759a71d17faf · outbound

This paper cites Long short-term memory,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Long short-term memory,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:38.314362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:38.314362Z digest=sha256:1463c3383db333bc7c16c2c7d94fc597217259c050d398f7b07570a42b99bb4b

Observation 8e7915b6-2426-47a9-b91c-624167a6d8df · outbound

This paper cites Unsupervised cross-lingual representation learning at scale,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Unsupervised cross-lingual representation learning at scale,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.626027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.320958Z digest=sha256:2c4103cfac3f8a4b6630494b7b9c394531b17791e87a0720655394d35ea7aea5

Observation 505bfbb3-e787-4066-8414-79046eae7cd0 · outbound

This paper cites Language-agnostic BERT sentence embedding,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Language-agnostic BERT sentence embedding,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.602020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.326252Z digest=sha256:65f4abc81131d2031d0bc40a9149376cfe41d794acd065d9c78f911b24627af7

Observation 9f77f6f7-c627-44be-b422-e4d7439c8f44 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Adam: A Method for Stochastic Optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T13:30:38.331807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:30:38.331807Z digest=sha256:e28f692095db7ea6fc68ff13130dd0699012ffd117e53853e6c675a0a8f46da0

Observation c829c89d-fa03-47e3-a84e-faaf63364e85 · outbound

This paper cites SMaLL-100: Introducing shallow multilingual machine translation model for low-resource languages,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval SMaLL-100: Introducing shallow multilingual machine translation model for low-resource languages,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.580492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.339261Z digest=sha256:5c524aa64cb426e406db15332da83ec4776c4dd71d0ae1f77bb62088c97a7ef3

Observation a6fa30c9-5528-4f85-80c3-470f498123d2 · outbound

This paper cites Visualizing data using t-SNE,.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Visualizing data using t-SNE,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.554532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.344969Z digest=sha256:ba1c654e2dfa91b2ece56e935a07449f410af315ad4c67c335108acecfc716f9

Observation 573bd04a-e36a-4dcd-b1c9-61e912e57a65 · outbound

This paper cites Available: https://aclanthology.org/2021.tacl-1.79/.

CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval Available: https://aclanthology.org/2021.tacl-1.79/

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:30:38.828910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:30:38.259743Z digest=sha256:cd9b4ccc4614b09e44b29c0cb48d1f1118b9a6ef323dec139836d124dc1cf78d

Pith citing papers

Observation c2cad6b0-d883-48aa-87bf-d32062f03bd2 · inbound

Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation cites this paper.

Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:35:57.105502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T23:35:56.585075Z digest=sha256:1d1e0792adddfa44fa1ee0fe626eaf63e22cad77de1dbd0f587a702aebdfd80c