Pith. sign in

The Importance of Being Recurrent for Modeling Hierarchical Structure

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Recent work has shown that recurrent neural networks (RNNs) can implicitly capture and exploit hierarchical information when trained to solve common natural language processing tasks such as language modeling (Linzen et al., 2016) and neural machine translation (Shi et al., 2016). In contrast, the ability to model structured data with non-recurrent neural networks has received little attention despite their success in many NLP tasks (Gehring et al., 2017; Vaswani et al., 2017). In this work, we compare the two architectures---recurrent versus non-recurrent---with respect to their ability to model hierarchical structure and find that recurrency is indeed important for this purpose.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Can Interpretation Predict Behavior on Unseen Data?

cs.LG · 2025-07-08 · conditional · novelty 6.0

Presence of hierarchical attention heads on in-distribution data predicts hierarchical out-of-distribution generalization across 270 small transformers, independent of causal support.

citing papers explorer

Showing 1 of 1 citing paper.

  • Can Interpretation Predict Behavior on Unseen Data? cs.LG · 2025-07-08 · conditional · none · ref 38 · internal anchor

    Presence of hierarchical attention heads on in-distribution data predicts hierarchical out-of-distribution generalization across 270 small transformers, independent of causal support.