Pith. sign in

REVIEW 2 cited by

ORCA: Interpreting Prompted Language Models via Locating Supporting Data Evidence in the Ocean of Pretraining Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.12600 v1 pith:HLKOVFTT submitted 2022-05-25 cs.CL cs.LG

classification cs.CLcs.LG
keywords datapretrainingevidencelanguagemodelmodelssupportingtask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large pretrained language models have been performing increasingly well in a variety of downstream tasks via prompting. However, it remains unclear from where the model learns the task-specific knowledge, especially in a zero-shot setup. In this work, we want to find evidence of the model's task-specific competence from pretraining and are specifically interested in locating a very small subset of pretraining data that directly supports the model in the task. We call such a subset supporting data evidence and propose a novel method ORCA to effectively identify it, by iteratively using gradient information related to the downstream task. This supporting data evidence offers interesting insights about the prompted language models: in the tasks of sentiment analysis and textual entailment, BERT shows a substantial reliance on BookCorpus, the smaller corpus of BERT's two pretraining corpora, as well as on pretraining examples that mask out synonyms to the task verbalizers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Derivational Morphology Reveals Analogical Generalization in Large Language Models

    cs.CL 2024-11 conditional novelty 7.0 of 10

    GPT-J's adjective nominalization behavior is better explained by an exemplar-based analogical model than by a rule-based model, with word frequency effects even for regular forms.

  2. Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Cross-lingual retrieval-augmented fine-tuning beats target-only training on average in eight hate speech detection languages, with peak performance near 2,000 retrieved instances.

Pith tools