GAR, a new synthetic benchmark for compositional relational reasoning, reveals poor LLM performance and yields the discovery of 'True/False' attention heads that encode truthfulness in Vicuna-33B.
heads vary with different rretrieve, because they retrieve different attributes according to the semantic re- lation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Benchmarking and Understanding Compositional Relational Reasoning of LLMs
GAR, a new synthetic benchmark for compositional relational reasoning, reveals poor LLM performance and yields the discovery of 'True/False' attention heads that encode truthfulness in Vicuna-33B.