Pith. sign in

REVIEW 4 cited by

MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.08582 v2 pith:OQXKNQ6I submitted 2022-04-18 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords massivedatasetlanguagesaccuracyassistantclassificationintentmodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the MASSIVE dataset--Multilingual Amazon Slu resource package (SLURP) for Slot-filling, Intent classification, and Virtual assistant Evaluation. MASSIVE contains 1M realistic, parallel, labeled virtual assistant utterances spanning 51 languages, 18 domains, 60 intents, and 55 slots. MASSIVE was created by tasking professional translators to localize the English-only SLURP dataset into 50 typologically diverse languages from 29 genera. We also present modeling results on XLM-R and mT5, including exact match accuracy, intent classification accuracy, and slot-filling F1 score. We have released our dataset, modeling code, and models publicly.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 22 citations worldwide. Full citation record

  1. Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Across eight zero-shot intent datasets, instruction-tuned ~3B open-weight models can match or beat larger base models, top systems are statistically tied on MASSIVE, and SNIPS is saturated.

  2. GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A single 205M-parameter encoder model unifies named entity recognition, text classification, and hierarchical structured extraction through declarative schemas.

  3. LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A multilingual visual question-answering benchmark across 11 languages and 5 social attributes, evaluated on 7 large multimodal models.

  4. Intent Classification on Low-Resource Languages with Query Similarity Search

    cs.IR 2025-05 conditional novelty 4.0 of 10

    A k-nearest-neighbor search over multilingual query embeddings provides zero-shot intent classification for low-resource languages, with accuracy below translation-based and supervised baselines.

Pith tools