REVIEW 6 cited by
Robust Lottery Tickets for Pre-trained Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent works on Lottery Ticket Hypothesis have shown that pre-trained language models (PLMs) contain smaller matching subnetworks(winning tickets) which are capable of reaching accuracy comparable to the original models. However, these tickets are proved to be notrobust to adversarial examples, and even worse than their PLM counterparts. To address this problem, we propose a novel method based on learning binary weight masks to identify robust tickets hidden in the original PLMs. Since the loss is not differentiable for the binary mask, we assign the hard concrete distribution to the masks and encourage their sparsity using a smoothing approximation of L0 regularization.Furthermore, we design an adversarial loss objective to guide the search for robust tickets and ensure that the tickets perform well bothin accuracy and robustness. Experimental results show the significant improvement of the proposed method over previous work on adversarial robustness evaluation.
Forward citations
Cited by 6 Pith papers
-
NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling
Dynamic sparse training of multiple heads on a shared backbone outperforms full dense ensembles on ImageNet and C4 while using less compute.
-
StructCoh: Structured Contrastive Learning for Context-Aware Text Semantic Matching
StructCoh, a graph-enhanced contrastive learning framework for text semantic matching, reportedly outperforms prior methods on legal and plagiarism benchmarks, but the reported results are not reproducible from the pa...
-
Using External knowledge to Enhanced PLM for Semantic Matching
A BERT-based model that injects WordNet lexical-relation signals into attention and adaptively fuses them claims consistent accuracy improvements on 10 semantic matching datasets.
-
Comateformer: Combined Attention Transformer for Semantic Sentence Matching
Comateformer replaces softmax attention with a product of tanh similarity and sigmoid dissimilarity scores, and reports consistent gains on ten semantic matching datasets.
-
Multi-Granularity Reasoning for Natural Language Inference
Stacking element-wise multi-layer BERT interactions and DenseNet yields modest NLI gains over BERT/RoBERTa baselines on standard benchmarks.
-
Boosting Neural Language Inference via Cascaded Interactive Reasoning
A new feature-extraction module for NLI combines all BERT layers with element-wise sentence-pair interactions and a DenseNet, reporting average gains of roughly one point on ten benchmarks, though evaluation inconsist...
Discussion (0). Continue with ORCID to comment.