Pith. sign in

Paper Citation Record · LEDGER

mSLAM: Massively multilingual joint pre-training for speech and text

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2202.01374.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2202.01374 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:52:49.835256Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T21:51:17.752402Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0a27e7da-de70-4ff1-aa85-a9c294b12dcf · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language mSLAM: Massively multilingual joint pre-training for speech and text

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.775223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:fc4b957ac33deaba79f86b70d445bccc5ce9a71284f0aacc27d3673a9dd67c27

Observation 46595cb1-d75f-46f7-8141-624fc939a5c5 · inbound

AudioPaLM: A Large Language Model That Can Speak and Listen cites this paper.

AudioPaLM: A Large Language Model That Can Speak and Listen mSLAM: Massively multilingual joint pre-training for speech and text

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:07:57.851683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T07:07:57.800866Z digest=sha256:652fe462802fe03e36f961fcf2f0a7f902a7d4b5a2d3fc17fc4ced674445e75e

Observation beb416f4-df7a-44e0-8960-ad83b6756b07 · inbound

STORM: Strategic Orchestration of Modalities for Rare Event Classification cites this paper.

STORM: Strategic Orchestration of Modalities for Rare Event Classification mSLAM: Massively multilingual joint pre-training for speech and text

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:09:14.792158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:09:14.792158Z digest=sha256:05732e605e2a434484920bbf1807e59de74e675443adbf547d3b803b736c906f

Observation ee707da0-40d7-407b-a982-46099aa5bfa7 · inbound

Large Concept Models: Language Modeling in a Sentence Representation Space cites this paper.

Large Concept Models: Language Modeling in a Sentence Representation Space mSLAM: Massively multilingual joint pre-training for speech and text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:01.021194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:37:01.021194Z digest=sha256:622ec90de1b8f37e506570d51c9ce4ee3be8ee919dc1c2d1fadfd1f1205cda13

Observation 3ffc3bae-0de9-4ff4-8b5e-d835bb75984d · inbound

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning cites this paper.

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning mSLAM: Massively multilingual joint pre-training for speech and text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:22.389557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:27:22.389557Z digest=sha256:0db3bc3075b6196250f82a5d66aeaa9b2e67792b0f14ff6f5f2c82d229b6e141

Observation d78d5ae2-659d-4f9f-ac33-a6669d093d25 · inbound

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators cites this paper.

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators mSLAM: Massively multilingual joint pre-training for speech and text

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T12:30:52.016897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T12:30:52.016897Z digest=sha256:7d80d77c4ce1cdd76277e780263da2cbdbb425e3d310ac33c838037535c11279

Observation 073b2990-7cce-476c-be3a-76ec126d729b · inbound

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment cites this paper.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment mSLAM: Massively multilingual joint pre-training for speech and text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.929162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.929162Z digest=sha256:6cab182662d8a236f75a580154543ad36cfec61bb928f2fcf2bb5487c3a49c13

Observation 21895f18-f986-4ed7-aacf-4a66dc443483 · inbound

Self-Improvement for Audio Large Language Model using Unlabeled Speech cites this paper.

Self-Improvement for Audio Large Language Model using Unlabeled Speech mSLAM: Massively multilingual joint pre-training for speech and text

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:49.835256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:52:49.835256Z digest=sha256:f7b6010821a1b38bfd8eede686ceaac6387a485a1b521bd0f589588dc3f0e47a

Observation d12b6ab0-06fa-4b7e-8094-9df5092d3073 · inbound

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs cites this paper.

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs mSLAM: Massively multilingual joint pre-training for speech and text

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:51:17.754099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T21:49:21.785096Z digest=sha256:d17eaffbc43236a3ec718bccc8d5bc2cdd3974fc7b28521be988f28a4ca45fa7

Observation 78517cd0-5e7d-41bc-be3c-0882b3a65c71 · inbound

Text-Utilization for Encoder-dominated Speech Recognition Models cites this paper.

Text-Utilization for Encoder-dominated Speech Recognition Models mSLAM: Massively multilingual joint pre-training for speech and text

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:51:25.132172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T13:37:09.614784Z digest=sha256:8d734d735e6909fb0993d45421f76ea9185859d7eb51a5c1dfc088886e4617fe