ATLAS traces RLVR data to 20 atomic sources, most datasets are variants, and DAPO++ curated with SCA improves RLVR performance while Q predicts training effectiveness.
Understanding black-box predictions via influence functions,
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
Susceptibilities applied to regret in deep RL agents reveal stagewise internal development in parameter space of a gridworld model that policy inspection alone cannot detect, validated via activation steering.
A framework based on linear response and influence functions maps data sensitivities in global QCD analyses to show how experiments determine central values, uncertainties, and correlations of non-perturbative functions.
BERT stores relational knowledge extractable via cloze queries without fine-tuning and matches supervised baselines on open-domain QA tasks.
Introduces an instance-level influence-based XAI method for dysarthria severity assessment that explains predictions by computing per-utterance influence scores from training samples and validates them via controlled deletion experiments.
idSCD uses semantic correlation descriptors to perform dataset membership inference by comparing learned semantic structures, outperforming baselines in NLI, emotion, and medical text experiments.
PRISM weights target examples by model preference to build an improved direction for influence-based data selection in LLM fine-tuning.
Influence scoring can use only forward passes: CountSketch-compressed outer products of the LM-head residual and final hidden state give accurate attribution and valuation from 14M to 32B parameters.
Watermark-based dataset inference achieves membership detection performance comparable to loss-based methods when subset exposure is high, under alternate assumptions.
Two methods using multiple unsupervised rankers for soft labels and filtering harmful examples allow learning-to-rank to beat BM25 with far fewer training examples.
citing papers explorer
-
RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data
ATLAS traces RLVR data to 20 atomic sources, most datasets are variants, and DAPO++ curated with SCA improves RLVR performance while Q predicts training effectiveness.
-
Interpreting Reinforcement Learning Agents with Susceptibilities
Susceptibilities applied to regret in deep RL agents reveal stagewise internal development in parameter space of a gridworld model that policy inspection alone cannot detect, validated via activation steering.
-
Mapping data sensitivities in global QCD analysis with linear response and influence functions
A framework based on linear response and influence functions maps data sensitivities in global QCD analyses to show how experiments determine central values, uncertainties, and correlations of non-perturbative functions.
-
Language Models as Knowledge Bases?
BERT stores relational knowledge extractable via cloze queries without fine-tuning and matches supervised baselines on open-domain QA tasks.
-
Towards Dys-XAI: Influence-Based Explanations for Dysarthria Severity Assessment
Introduces an instance-level influence-based XAI method for dysarthria severity assessment that explains predictions by computing per-utterance influence scores from training samples and validates them via controlled deletion experiments.
-
idSCD: Identifying Training Datasets through Semantic Correlation Descriptors
idSCD uses semantic correlation descriptors to perform dataset membership inference by comparing learned semantic structures, outperforming baselines in NLI, emotion, and medical text experiments.
-
PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning
PRISM weights target examples by model preference to build an improved direction for influence-based data selection in LLM fine-tuning.
-
Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation
Influence scoring can use only forward passes: CountSketch-compressed outer products of the LM-head residual and final hidden state give accurate attribution and valuation from 14M to 32B parameters.
-
Watermarking for Proprietary Dataset Protection
Watermark-based dataset inference achieves membership detection performance comparable to loss-based methods when subset exposure is high, under alternate assumptions.
-
Learning More From Less: Towards Strengthening Weak Supervision for Ad-Hoc Retrieval
Two methods using multiple unsupervised rankers for soft labels and filtering harmful examples allow learning-to-rank to beat BM25 with far fewer training examples.