Pith. sign in

Paper Citation Record · LEDGER

MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2204.08582.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.08582 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:27:59.642737Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

22
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 47ccb78d-96ef-4450-adc3-19ab7e6cf041 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:11.333078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:f08ebfa85a5345f7cb17662ea975cba24ffe32bcb87baf1f856e34bca9857198

Observation 20727a5d-14d2-4df3-8cc3-6f1e8531f64b · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 231

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:11.553728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:ad21147d74d24f016f8cd3c1a465f3e06c819e6d592b3d596d267f16b56dea44

Observation 9798b72b-7f95-4db8-96bf-ae1066c9892c · inbound

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings cites this paper.

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:14:16.595157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-24T09:11:37.220487Z digest=sha256:7c8a6546dbb11aea4c46db808e073d119f7ca7383c6dc707ec3759cb723c01e5

Observation 4b91aacd-75a5-49de-abe8-58c9751b9ca4 · inbound

NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models cites this paper.

NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:15:16.351510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T21:15:16.112918Z digest=sha256:a5b4bfadfa6306cae9e33e59f699c989e5b059d8974a73cb38b8348efec9d7c4

Observation 01a2f8e3-12b6-4798-9e04-0a2ddf49ac5a · inbound

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification cites this paper.

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 1955

Resolution
unresolved
no resolver link, observed 2026-08-08T17:27:59.642737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:27:59.642737Z digest=sha256:ce51a062622f47579ea86aa3163f0665b448d02577e7e06e15a0cdb146a56a32

Observation 0a611fcd-294e-4569-b711-e8989aa50400 · inbound

Intent Classification on Low-Resource Languages with Query Similarity Search cites this paper.

Intent Classification on Low-Resource Languages with Query Similarity Search MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:08.066823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:08.066823Z digest=sha256:22b94a48ac488cc51486af78b78e24f718846b4a24ee0506a2f5e4bd7936ce0d

Observation c44afdc6-6ba8-410d-b38b-c7bc0dac977e · inbound

LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation cites this paper.

LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:48:32.084674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:48:32.084674Z digest=sha256:61a2da6856321ad09f0ba03d2327165dde370dd1db73df3993199adf4d6c73b7

Observation 483159a3-476e-4f3d-81bb-ea033f65b641 · inbound

GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface cites this paper.

GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:54.168354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:35:54.168354Z digest=sha256:1d273cdaa74a96bd299d20e0d5e230edc6372ee59ac1b0f769f877ce60c9014c

Observation 7fa525ac-2fc8-42a4-9445-f3721c7d5541 · inbound

Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty cites this paper.

Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:00:09.681882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:00:07.732291Z digest=sha256:dcc36de88e4f924b9d56106c2652253c12c4a391e20230c1ab28b8b1a8f0c128

Observation b2cb6743-6a96-4f4a-ab19-02d76309edc9 · inbound

Uncovering the Latent Potential of Deep Intermediate Representations cites this paper.

Uncovering the Latent Potential of Deep Intermediate Representations MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.299164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T05:36:24.743558Z digest=sha256:12b4ad5d3dfdb90f1055e8ad90556607ee8a2aac4880edb693d48bd375172989

Observation 6066659e-09f6-433a-ab74-53ebec17c718 · inbound

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models cites this paper.

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T01:27:13.123451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:27:13.123451Z digest=sha256:283e1ae5e46c3ce2b6b7c28e540ca63e0c534d50a1247db8c880e4d777aa7a57