Pith. sign in

REVIEW 4 cited by

Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.00719 v3 pith:XFOB7WVQ submitted 2020-05-02 cs.CL

classification cs.CL
keywords probingtasktasksmodelspropertiesencodeevenlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although neural models have achieved impressive results on several NLP benchmarks, little is understood about the mechanisms they use to perform language tasks. Thus, much recent attention has been devoted to analyzing the sentence representations learned by neural encoders, through the lens of `probing' tasks. However, to what extent was the information encoded in sentence representations, as discovered through a probe, actually used by the model to perform its task? In this work, we examine this probing paradigm through a case study in Natural Language Inference, showing that models can learn to encode linguistic properties even if they are not needed for the task on which the model was trained. We further identify that pretrained word embeddings play a considerable role in encoding these properties rather than the training task itself, highlighting the importance of careful controls when designing probing experiments. Finally, through a set of controlled synthetic tasks, we demonstrate models can encode these properties considerably above chance-level even when distributed in the data as random noise, calling into question the interpretation of absolute claims on probing tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  2. LLMs are Bayesian, In Expectation, Not in Realization

    stat.ML 2025-07 conditional novelty 6.0 of 10

    Transformers can be Bayes-competitive in prequential log loss even when their predictive distributions are not invariant to example order, provided the cumulative predictive KL to the Bayesian reference stays small.

  3. SemSketches-2021: experimenting with the machine processing of the pilot semantic sketches corpus

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The authors release a pilot corpus of 915 semantic sketches for Russian and report that the best automatic matcher achieves only 0.277 accuracy, far below human-level performance.

  4. Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A survey showing that common systematic generalization benchmarks measure behavioural systematicity, not the representational systematicity that Fodor and Pylyshyn's challenge requires, and mapping them onto Hadley's ...

Pith tools