Fragility, the activation noise level causing probe accuracy collapse, reveals evolving lexical-to-compositional moral encoding, layer robustness gradients, and fine-tuning differences invisible to saturated probing accuracy.
Wojcik and Peter H
3 Pith papers cite this work, alongside 1,654 external citations. Polarity classification is still indexing.
fields
cs.CL 3years
2026 3representative citing papers
AMALIA matches larger models on agreement with human coders for moral-authority annotation, yet its recovery gap shows most of that performance is not reproduced by the theory-defined route—revealing a validity shortfall invisible to agreement metrics.
Large-scale statistical analysis of four harmful language datasets reveals that interactions between annotator characteristics and linguistic cues drive annotation variation, with lexical features and attitudes prominent but patterns varying by dataset.
citing papers explorer
-
When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis
Fragility, the activation noise level causing probe accuracy collapse, reveals evolving lexical-to-compositional moral encoding, layer robustness gradients, and fine-tuning differences invisible to saturated probing accuracy.
-
Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
AMALIA matches larger models on agreement with human coders for moral-authority annotation, yet its recovery gap shows most of that performance is not reproduced by the theory-defined route—revealing a validity shortfall invisible to agreement metrics.
-
Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation
Large-scale statistical analysis of four harmful language datasets reveals that interactions between annotator characteristics and linguistic cues drive annotation variation, with lexical features and attitudes prominent but patterns varying by dataset.