Pith. sign in

Paper Citation Record · LEDGER

MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2204.08582.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.08582 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:27:59.642737Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

22
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 47ccb78d-96ef-4450-adc3-19ab7e6cf041 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:11.333078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:d2a9ac473c056908d34650e86ed8effb2908d8ba67a849484c57a9e012aadceb

Observation 20727a5d-14d2-4df3-8cc3-6f1e8531f64b · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 231

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:11.553728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:db90610633891c54dccff27bccd3571218e5d4126195812081fe719752d277d2

Observation 9798b72b-7f95-4db8-96bf-ae1066c9892c · inbound

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings cites this paper.

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:14:16.595157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-24T09:11:37.220487Z digest=sha256:47b8997468c1c89f147c20511f3e47412755e06f5848b64bf79a4a9b46759ed9

Observation 4b91aacd-75a5-49de-abe8-58c9751b9ca4 · inbound

NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models cites this paper.

NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:15:16.351510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T21:15:16.112918Z digest=sha256:b59d08a484f4cba24381fbca1beb5191de5f3de6a56d1d90bcc480dd83252bb1

Observation 01a2f8e3-12b6-4798-9e04-0a2ddf49ac5a · inbound

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification cites this paper.

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 1955

Resolution
unresolved
no resolver link, observed 2026-08-08T17:27:59.642737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:27:59.642737Z digest=sha256:35070c70397c734f4b3530da47fae0d12f71d3d288ca524022813a636efd4f37

Observation 0a611fcd-294e-4569-b711-e8989aa50400 · inbound

Intent Classification on Low-Resource Languages with Query Similarity Search cites this paper.

Intent Classification on Low-Resource Languages with Query Similarity Search MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:08.066823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:08.066823Z digest=sha256:146ad2cd70d826b793807792599d59e4a8fed903c4d38b144c80b63d2d957479

Observation c44afdc6-6ba8-410d-b38b-c7bc0dac977e · inbound

LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation cites this paper.

LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:48:32.084674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:48:32.084674Z digest=sha256:c3c88a1e7149ba3499187ff8175bb63695031c77db50a4afee45d2e4a6a8ebf0

Observation 483159a3-476e-4f3d-81bb-ea033f65b641 · inbound

GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface cites this paper.

GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:54.168354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:35:54.168354Z digest=sha256:6bb360d9c95925cf331939eb93650c01f7656b5ff043657b7c52cd10174fa561

Observation 7fa525ac-2fc8-42a4-9445-f3721c7d5541 · inbound

Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty cites this paper.

Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:00:09.681882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:00:07.732291Z digest=sha256:869596276393c14f93f66f924c93e0d05a3e7796a74ab6f55d0fa34efe0564f5

Observation b2cb6743-6a96-4f4a-ab19-02d76309edc9 · inbound

Uncovering the Latent Potential of Deep Intermediate Representations cites this paper.

Uncovering the Latent Potential of Deep Intermediate Representations MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.299164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T05:36:24.743558Z digest=sha256:72621186e70971ea01ef94391d214e9616b03f0471f00dadae3be475f926d17b

Observation 6066659e-09f6-433a-ab74-53ebec17c718 · inbound

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models cites this paper.

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T01:27:13.123451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:27:13.123451Z digest=sha256:5600502a58c26c3fa2684b90acaf015d2b511c6ac543a5c8058ce7a91459c2ac