Pith. sign in

Paper Citation Record · LEDGER

MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2204.08582.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.08582 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:25:47.821364Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

22
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 47ccb78d-96ef-4450-adc3-19ab7e6cf041 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 156

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:11.333078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:5a8e19c006432d84ff1960224ac992bb402bdbc704fa0791096608ab3ae311bb

Observation 20727a5d-14d2-4df3-8cc3-6f1e8531f64b · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 231

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:11.553728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:0984e0d472710242bd988039f24135386d65eb2da340519b90a744add47add9e

Observation 9798b72b-7f95-4db8-96bf-ae1066c9892c · inbound

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings cites this paper.

Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:14:16.595157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-24T09:11:37.220487Z digest=sha256:f32bfbdc03ee1c8ff7291f624242fdc7d89d726a16694186dcb0ca80f8394964

Observation 4b91aacd-75a5-49de-abe8-58c9751b9ca4 · inbound

NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models cites this paper.

NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 158

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:15:16.351510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T21:15:16.112918Z digest=sha256:4aec72250c4ece1a3114d4ed18186e6ccc0c34918412fda1b7c021199a3e3ce4

Observation eb67db71-e711-4d04-b2a5-2b3e52d8a8e1 · inbound

IntentGPT: Few-shot Intent Discovery with Large Language Models cites this paper.

IntentGPT: Few-shot Intent Discovery with Large Language Models MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:31:34.301497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:31:34.301497Z digest=sha256:ba044222d6c16fe3243d31e8714a4660a840665837dcb1475c277c45aabd3c77

Observation d3274685-f536-40c1-b4de-3a25cffe551c · inbound

Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding cites this paper.

Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:56:41.849928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:56:41.849928Z digest=sha256:0448440f33072978fcdfe9f0a9e6a50d5fd92a60fe23f4f538f38f5561acc7cc

Observation 10aaec19-5e62-473e-8116-da35bd2b89c1 · inbound

Advancing Single and Multi-task Text Classification through Large Language Model Fine-tuning cites this paper.

Advancing Single and Multi-task Text Classification through Large Language Model Fine-tuning MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T17:48:13.474348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:48:13.474348Z digest=sha256:411fbc73481261ee6d1df88af7c8d1bf0dbe7f29709560985728a4aebbea24f6

Observation e55749b8-eea0-42a8-9ed0-4fd059cca110 · inbound

Jasper and Stella: distillation of SOTA embedding models cites this paper.

Jasper and Stella: distillation of SOTA embedding models MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T01:03:28.098358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T01:03:28.098358Z digest=sha256:646d94f1b9e06fa3586e42c4a10f0cc48b19001ba5ac7a355ecc11a7e01d69ed

Observation e5227fe1-2784-42c6-9d26-7760e7b5979f · inbound

A Multi-Encoder Frozen-Decoder Approach for Fine-Tuning Large Language Models cites this paper.

A Multi-Encoder Frozen-Decoder Approach for Fine-Tuning Large Language Models MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:38:20.517252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:38:20.517252Z digest=sha256:6cfb899be3d52f3fd7fba21d2484690ac8f8fae393f95204b3de399b639e3fe3

Observation 51707239-d9b9-43b1-83d1-f424685cf9da · inbound

Automatic Labelling with Open-source LLMs using Dynamic Label Schema Integration cites this paper.

Automatic Labelling with Open-source LLMs using Dynamic Label Schema Integration MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:20:32.312626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:20:32.312626Z digest=sha256:0d0a3d96b114570057afbe0dd91f69b57b2458fa895612e34290320abb278106

Observation 377af1b3-9d64-40b3-9ba9-d445bb5c67a1 · inbound

Analysis of Indic Language Capabilities in LLMs cites this paper.

Analysis of Indic Language Capabilities in LLMs MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:10.566979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:10.566979Z digest=sha256:dca8c4a64e5353c6758d31fb6d779c8020b857c913f1c80620bf28691347d9b4

Observation 01a2f8e3-12b6-4798-9e04-0a2ddf49ac5a · inbound

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification cites this paper.

ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 1955

Resolution
unresolved
no resolver link, observed 2026-08-08T17:27:59.642737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:27:59.642737Z digest=sha256:c163328d85ad52de3732074e00ab28ca01824186acb71ebab5fbb6b425560064

Observation e13a291e-6ed4-4871-9d8a-a0131878a111 · inbound

Cequel: Cost-Effective Querying of Large Language Models for Text Clustering cites this paper.

Cequel: Cost-Effective Querying of Large Language Models for Text Clustering MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T11:25:47.821364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:25:47.821364Z digest=sha256:5ea55bda164a858ba44d1be362b06cb6f89961c0980a1c3605892bd3063e925d

Observation 1c198a14-319d-48b6-9bbf-373959701937 · inbound

A Survey on Parameter-Efficient Fine-Tuning for Foundation Models in Federated Learning cites this paper.

A Survey on Parameter-Efficient Fine-Tuning for Foundation Models in Federated Learning MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T05:16:30.652995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:16:30.652995Z digest=sha256:dfa9458621a6faa0d94c047105fbd8728064abd8ebfd700916539f57ed9b9cf9

Observation 0a611fcd-294e-4569-b711-e8989aa50400 · inbound

Intent Classification on Low-Resource Languages with Query Similarity Search cites this paper.

Intent Classification on Low-Resource Languages with Query Similarity Search MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:08.066823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:08.066823Z digest=sha256:46a55f38d4d98271d937a24461138745cab33756e4cd7ccc38429ef205c9a7c0

Observation c44afdc6-6ba8-410d-b38b-c7bc0dac977e · inbound

LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation cites this paper.

LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:48:32.084674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:48:32.084674Z digest=sha256:40fc49ad4c9e3e23fd91b4af5aa117ed12ec7e5678ddc2c050d4f9e80dfd4705

Observation 483159a3-476e-4f3d-81bb-ea033f65b641 · inbound

GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface cites this paper.

GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:54.168354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:35:54.168354Z digest=sha256:bf71c65d1e7e9888be872bf3cec72ee1a7372094714a15a57eefbf0b611f258c

Observation 7fa525ac-2fc8-42a4-9445-f3721c7d5541 · inbound

Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty cites this paper.

Trust the uncertain teacher: distilling dark knowledge via calibrated uncertainty MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:00:09.681882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T13:00:07.732291Z digest=sha256:1196423cf259733560d0052c058ffcd5d61435a8fc9f4bab878cd7d7c4f5493f

Observation b2cb6743-6a96-4f4a-ab19-02d76309edc9 · inbound

Uncovering the Latent Potential of Deep Intermediate Representations cites this paper.

Uncovering the Latent Potential of Deep Intermediate Representations MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.299164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-25T05:36:24.743558Z digest=sha256:0eaf1dabf380057126eaa9217578ebd642671c71deb9ce29f0dfa4d3935c155f

Observation 6066659e-09f6-433a-ab74-53ebec17c718 · inbound

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models cites this paper.

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T01:27:13.123451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:27:13.123451Z digest=sha256:8905a3060a4fb4607bf905c06d1fcd735bef07364d0791538279123558ada182

Observation 6c44d333-42d4-43a6-b91d-f13504dad467 · inbound

The Embedder's Dilemma: LLMs Are Better, but at What Cost? cites this paper.

The Embedder's Dilemma: LLMs Are Better, but at What Cost? MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T21:37:01.423713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:37:01.423713Z digest=sha256:3bc5e7afb9268758e08ac17c2bb2ea964ca85c0991efd9255f2e7865dbd67017