REVIEW 3 cited by
Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent model editing techniques promise to mitigate the problem of memorizing false or outdated associations during LLM training. However, we show that these techniques can introduce large unwanted side effects which are not detected by existing specificity benchmarks. We extend the existing CounterFact benchmark to include a dynamic component and dub our benchmark CounterFact+. Additionally, we extend the metrics used for measuring specificity by a principled KL divergence-based metric. We use this improved benchmark to evaluate recent model editing techniques and find that they suffer from low specificity. Our findings highlight the need for improved specificity benchmarks that identify and prevent unwanted side effects.
Forward citations
Cited by 3 Pith papers
-
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
Small changes in data or settings used to find a circuit in a language model often produce very different circuits: under bootstrap resampling, average pairwise overlap of EAP-IG circuits across tasks and models is on...
-
Linear Correlation in LM's Compositional Generalization and Hallucination
Language models' next-token predictions for related knowledge are connected by near-linear transformations that persist through fine-tuning, explaining both compositional generalization and hallucination.
-
A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models
A survey claims proactive defenses against LLM misinformation outperform post-hoc detection by up to 63%, but no meta-analysis details are provided to support the claim.
Discussion (0). Continue with ORCID to comment.