Pith. sign in

REVIEW 2 major objections 4 minor 98 references

Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey

T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Systematic behaviour on a benchmark does not show that a model has systematic internal representations, and most current benchmarks do not require the strong form of systematicity that comes closest to human generalization.

desk verdict A useful survey that cleanly separates behavioural from representational systematicity, but its classification of visual benchmarks under a syntax-free Hadley taxonomy is underdetermined. read the letter →

arxiv 2506.04461 v1 pith:IMUHCEJ4 submitted 2025-06-04 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords systematicitycompositionalitybehaviouralrepresentationalHadley'staxonomymechanisticinterpretabilitysystematicgeneralizationFodorandPylyshyn
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Machine-learning papers routinely cite Fodor and Pylyshyn's challenge as a reason to build or test compositional models, but this survey argues that the community is answering a different question. Fodor and Pylyshyn asked for systematic mental representations; the paper contends that nearly all current benchmarks and training studies measure systematic behaviour only. Using Hadley's three-level scale of weak, quasi, and strong systematicity, the authors classify representative language and vision benchmarks and conclude that many never require strong systematicity, the level closest to human generalization. Their constructive recommendation is that behavioural success must be paired with mechanistic interpretability, evidence about the internal representations that actually drive a model's output.

What carries the argument

The load-bearing machinery is Hadley's (1994) three-level taxonomy of systematicity—weak, quasi, and strong—which classifies what a training set has already shown a learner. Weak systematicity permits novel combinations of familiar words only in familiar syntactic positions; quasi-systematicity adds recursion over structurally familiar clauses; strong systematicity requires handling words in syntactic positions never seen in training, the level Hadley ties to human generalization. The authors use this scale to operationalize Fodor and Pylyshyn's representational systematicity, and they pair it with the competence/performance distinction: behavioural success is evidence, not proof, and mechanistic interpretability supplies the causal evidence needed to move from behaviour to representations.

What would settle it

Enumerate the train and test splits of a benchmark the paper classifies as testing only weak systematicity, for example SCAN Split 1, and count the test items in which some word appears in a syntactic position never attested in any training sentence; if that fraction is substantial rather than near zero, the weak-systematicity classification is false. The same position-coverage calculation, published for each benchmark, would settle which of the survey's classifications are correct.

Watch

Extended reading notes

Core claim

The central claim is that Fodor and Pylyshyn defined compositionality through systematicity of representations, not behaviour, and that the machine-learning literature invoking them has largely substituted behavioural systematicity for representational systematicity. On the authors' reading, a model can pass systematicity benchmarks through memorization or task-specific heuristics while lacking the structured internal representations Fodor and Pylyshyn argued are necessary. The survey adopts Hadley's weak/quasi/strong taxonomy as the operationalization of that distinction and uses it to classify benchmarks: SCAN's Split 1 and PCFG SET's systematicity split test weak systematicity; productivity-oriented splits require quasi-systematicity; COGS and ReCOGS-style structural generalization approach strong systematicity; most visual benchmarks do not reach it. The conclusion is that claims to have answered the Fodor–Pylyshyn challenge should be backed by mechanistic interpretability of the model's representations, not by benchmark scores alone.

Load-bearing premise

Everything downstream rests on the assumption that Hadley's 1994 three-level scale, built for connectionist language learning, transfers to modern end-to-end models and especially to visual benchmarks, where the paper concedes there is no agreed syntax of images to anchor the levels.

Editorial extensions

If this is right

  • Benchmark reports should state which level of Hadley's taxonomy the train/test split actually demands, rather than treating any compositional split as evidence of Fodor–Pylyshyn systematicity.
  • A claim that a model addresses the Fodor–Pylyshyn challenge should be accompanied by mechanistic evidence that the identified representations are causally responsible for the systematic behaviour.
  • New language benchmarks should be built to target strong systematicity by controlling which words appear in which syntactic positions during training.
  • For large pre-trained models, claims of strong behavioural systematicity are hard to support because the training distribution is not controlled.
  • Visual systematicity benchmarks cannot be fully scored on Hadley's scale until a theory of the syntax of images is available.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the authors do not develop is a quantitative distance-to-strong measure: for any train/test split, compute the fraction of test items whose words appear only in syntactic positions absent from all training sentences; benchmarks could then be ordered by how much beyond weak systematicity they demand.
  • If the survey is right, cross-benchmark model rankings should not be read as rankings of systematic generalization; the same model may exploit different shortcuts on different benchmarks, which is a direct explanation for the disagreement across benchmarks the paper cites.
  • The paper's position implies a testable criterion for representational systematicity: causal interventions on internal representations, for example swapping role or binding vectors, should change the model's output exactly as the compositional operation predicts; where they do not, the behavioural success is probably shortcut-driven.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper argues that much of the machine-learning literature on systematicity measures behavioural systematicity, whereas Fodor and Pylyshyn's original challenge concerned systematicity of representations. The authors adopt Hadley's (1994) weak/quasi/strong taxonomy as an operational framework, apply it to a selection of language and vision benchmarks, and conclude that many current benchmarks fall short of strong systematicity, the level closest to human systematic generalization. They then review evidence for and against representational systematicity in end-to-end Transformer models, drawing on mechanistic interpretability, and close with recommendations for evaluating systematicity claims.

Significance. If the central distinction is accepted, the paper provides an important corrective to the field's frequent invocation of Fodor and Pylyshyn. Its strengths are the clear separation of behavioural from representational systematicity, the concrete examples illustrating the dangers of conflating the two (e.g., Bastings et al.'s SCAN shortcuts and ReCOGS's impact on COGS results), and the constructive emphasis on mechanistic interpretability as a complement to behavioural benchmarks. The paper is an opinionated survey, so its soundness rests on the accuracy and internal consistency of its interpretive claims; the main argument is well supported by the historical discussion and the case studies in language, though the visual-benchmark analysis is less secure.

major comments (2)
  1. [§4.2] The classification of visual benchmarks under Hadley's taxonomy is underdetermined because the paper itself acknowledges that no strong theory of image syntax exists. The text states, 'Lacking a strong theory of the syntax of images, it is unclear how Hadley's levels of systematicity apply to visual concepts,' yet it immediately assigns 'weak systematicity' to disentanglement datasets on the grounds that they 'lack a hierarchical structure' and do not place concepts in novel contexts. Hierarchical structure, syntactic position, and novel context are not defined for image contents. As a result, the weak/quasi/strong labels for visual benchmarks are not well grounded, and the Section 6 conclusion that 'many current benchmarks fall short of testing strong systematicity' is not supported for the visual half of the survey. I recommend either providing a working visual syntax (e.g., scene graphs or a grammar over generative factors) or explicitly limiting the Hadley-based classification to linguistic benchmarks and presenting the visual discussion as exploratory.
  2. [§2.3 and §4.1] The classification of productivity splits as requiring at least quasi-systematicity rests on an unproven conjecture. The paper asserts that weak systematicity yields no productivity because a sentence of unseen length 'will necessarily' contain at least one word in a novel syntactic position. This is not established for arbitrary phrase-structure grammars; under recursive grammars, arguments may occupy the same positions at all depths, and whether a word is in a novel position depends on the grammar's definition of positions. The classification of the PCFG SET Productivity split and SCAN's Split 2 depends on this step. Please either prove the claim under explicit grammar assumptions or soften the classification to 'may require at least quasi-systematicity'.
minor comments (4)
  1. [§3.1] The informal phrase 'Mike Young and colleagues' should be replaced with a formal citation, since the text later cites Young and Wasserman (1997, 2001) by name.
  2. [References] Several reference entries contain formatting errors in author names (e.g., 'DeV os', 'V on Kügelgen'); these should be corrected before publication.
  3. [§4.2] The sentence 'ARC may constitute a test for strong compositionality but cannot be judged on an objective basis' is ambiguous; please separate the hedged claim from the justification that the construction process is not systematically described.
  4. [§4.1] For SCAN Split 1, the classification as weak systematicity relies on the statistical claim that even 2% of the training data exhibits all possible commands in all possible positions; please state explicitly that this is an empirical judgment about the dataset rather than a formal property of the split.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's central distinction and benchmark classifications are imported from external sources (Fodor & Pylyshyn; Hadley), not derived from its own conclusion.

full rationale

This is an opinionated survey rather than a derivation, and its central claims are backed by external, non-self-authored sources. The behavioural-vs-representational distinction is attributed to Fodor and Pylyshyn (1988), and the weak/quasi/strong taxonomy is Hadley's (1994); the paper explicitly adopts these frameworks rather than defining them in terms of its own target conclusion. The benchmark classifications in Sections 4.1 and 4.2 are applications of Hadley's externally defined levels to specific datasets, and the observation that many benchmarks 'fall short of testing strong systematicity' follows from inspecting those datasets under Hadley's definitions; it is a classification, not a forced equivalence. The paper itself flags the main threat to this application in Section 4.2: 'lacking a strong theory of the syntax of images, it is unclear how Hadley's levels of systematicity apply to visual concepts,' which is an acknowledged limitation on the validity of the visual extension, not a circular step. Self-citations (Lewis et al. 2024; Doumas et al. 2022) appear only as examples or as out-of-scope footnotes and are not load-bearing; they do not supply the framework or the uniqueness of the taxonomy. No fitted parameters are renamed as predictions, and no uniqueness theorem from the authors' prior work is invoked. The productivity conjectures in Section 2.3 are consequences of the definitions of weak/quasi/strong systematicity and unseen-length sentences, not circular inputs. The least secure part of the paper is the transfer of Hadley's linguistic taxonomy to visual benchmarks, but that is an underdetermination/correctness concern, not circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper's central claim rests on four main interpretive assumptions: the representational reading of Fodor and Pylyshyn, the validity of Hadley's taxonomy, the requirement of controlled operationalisation, and the conditional inference from valid behavioural success to representational systematicity. There are no free parameters or invented entities. The assumptions are largely explicit, but the strongest one, the validity of Hadley's taxonomy, receives less critical scrutiny than it needs.

assumptions (4)
  • domain assumption Fodor and Pylyshyn's systematicity is fundamentally a property of representations, not behaviour.
    This is the foundational interpretive claim of the paper, stated in Section 2.1 and used throughout. It is not proven but is attributed to Fodor and Pylyshyn.
  • domain assumption Hadley's three levels of systematicity provide a valid operationalization of Fodor and Pylyshyn's representational systematicity.
    Adopted in Section 3.1 and Section 4. The paper treats Hadley's taxonomy as a robust framework, but it is a theoretical choice and the paper itself notes uncertainty for visual tasks.
  • domain assumption A valid behavioural operationalisation requires principled control over dataset syntax and semantics.
    Stated in Section 3.2, Case 1. This assumption underlies the paper's critique of benchmarks that lack such control.
  • domain assumption Behavioural success under a valid operationalisation is evidence of representational systematicity only if Fodor and Pylyshyn's position is accepted.
    Discussed in Section 3.2, Case 2. The paper notes that this is a conditional conclusion, not an independently established fact.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey." pith.science (2026). https://pith.science/paper/IMUHCEJ4

@misc{pith2026250604461,
  author       = {Pith},
  title        = {Pith review of: Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IMUHCEJ4}},
  note         = {Machine review of arXiv:2506.04461}
}
read the original abstract

A core aspect of compositionality, systematicity is a desirable property in ML models as it enables strong generalization to novel contexts. This has led to numerous studies proposing benchmarks to assess systematic generalization, as well as models and training regimes designed to enhance it. Many of these efforts are framed as addressing the challenge posed by Fodor and Pylyshyn. However, while they argue for systematicity of representations, existing benchmarks and models primarily focus on the systematicity of behaviour. We emphasize the crucial nature of this distinction. Furthermore, building on Hadley's (1994) taxonomy of systematic generalization, we analyze the extent to which behavioural systematicity is tested by key benchmarks in the literature across language and vision. Finally, we highlight ways of assessing systematicity of representations in ML models as practiced in the field of mechanistic interpretability.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

98 extracted references · 31 canonical work pages

  1. [1]

    Carvalho, and André Freitas

    Nura Aljaafari, Danilo S. Carvalho, and André Freitas. 2025. https://doi.org/10.48550/arXiv.2410.12924 Interpreting token compositionality in LLMs : A robustness analysis . arXiv preprint. ArXiv:2410.12924 [cs] version: 2

  2. [2]

    Jacob Andreas. 2019. https://doi.org/10.48550/arXiv.1902.07181 Measuring Compositionality in Representation Learning . arXiv preprint. ArXiv:1902.07181 [cs]

  3. [3]

    Rim Assouel, Pietro Astolfi, Florian Bordes, Michal Drozdzal, and Adriana Romero-Soriano. 2024. Oc-clip: Object-centric binding in contrastive language-image pretraining. In NeurIPS 2024 Workshop on Compositional Learning: Perspectives, Methods, and Paths Forward

  4. [4]

    Rim Assouel, Pau Rodriguez, Perouz Taslakian, David Vazquez, and Yoshua Bengio. 2022. Object-centric compositional imagination for visual abstract reasoning. In ICLR2022 Workshop on the Elements of Reasoning: Objects, Structure and Causality

  5. [5]

    Samy Badreddine, Artur d'Avila Garcez, Luciano Serafini, and Michael Spranger. 2022. https://doi.org/10.1016/j.artint.2021.103649 Logic Tensor Networks . Artificial Intelligence, 303:103649. ArXiv:2012.13635 [cs]

  6. [6]

    Renée Baillargeon and Julie DeVos. 1991. https://doi.org/10.1111/j.1467-8624.1991.tb01602.x Object Permanence in Young Infants : Further Evidence . Child Development, 62(6):1227--1246. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1467-8624.1991.tb01602.x

  7. [7]

    Jasmijn Bastings, Marco Baroni, Jason Weston, Kyunghyun Cho, and Douwe Kiela. 2018. https://doi.org/10.18653/v1/W18-5407 Jump to better conclusions: SCAN both left and right . In Proceedings of the 2018 EMNLP Workshop BlackboxNLP : Analyzing and Interpreting Neural Networks for NLP , pages 47--55, Brussels, Belgium. Association for Computational Linguistics

  8. [8]

    Yonatan Belinkov. 2022. https://doi.org/10.1162/coli_a_00422 Probing Classifiers : Promises , Shortcomings , and Advances . Computational Linguistics, 48(1):207--219

Show all 98 references
  1. [9]

    Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen,...

  2. [10]

    Elisabeth Camp. 2007. https://doi.org/10.1111/j.1520-8583.2007.00124.x Thinking with Maps . Philosophical Perspectives, 21(1):145--182. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1520-8583.2007.00124.x

  3. [11]

    Patrick Cavanagh. 2021. https://doi.org/10.1177/0301006621991491 The Language of Vision * . Perception, 50(3):195--215. Publisher: SAGE Publications Ltd STM

  4. [12]

    Fran c ois Chollet. 2019. On the measure of intelligence. arXiv preprint arXiv:1911.01547

  5. [13]

    Noam Chomsky. 1965. https://www.jstor.org/stable/j.ctt17kk81z Aspects of the Theory of Syntax , 50 edition. The MIT Press

  6. [14]

    Bob Coecke, Mehrnoosh Sadrzadeh, and Stephen Clark. 2010. https://doi.org/10.48550/arXiv.1003.4394 Mathematical Foundations for a Compositional Distributional Model of Meaning . arXiv preprint. ArXiv:1003.4394 [cs]

  7. [15]

    Róbert Csordás, Kazuki Irie, and Juergen Schmidhuber. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.49 The Devil is in the Detail : Simple Tricks Improve Systematic Generalization of Transformers . In Proceedings of the 2021 Conference on Empirical Methods in Natural Langu...

  8. [16]

    Robert Cummins. 1996. Systematicity. The Journal of Philosophy, 93(12):591--614

  9. [17]

    Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2023. https://doi.org/10.48550/arXiv.2309.08600 Sparse Autoencoders Find Highly Interpretable Features in Language Models . arXiv preprint. ArXiv:2309.08600 [cs]

  10. [18]

    Verna Dankers, Elia Bruni, and Dieuwke Hupkes. 2022. https://doi.org/10.48550/arXiv.2108.05885 The paradox of the compositionality of natural language: a neural machine translation case study . arXiv preprint. ArXiv:2108.05885 [cs]

  11. [19]

    Aniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Charles Blundell, Philippe Beaudoin, Nicolas Heess, Michael C Mozer, and Yoshua Bengio. 2021. Neural production systems. Advances in Neural Information Processing Systems, 34:25673--25687

  12. [20]

    Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, and Kyle Mahowald. 2022. Why is winoground hard? investigating failures in visuolinguistic compositionality. arXiv preprint arXiv:2211.00768

  13. [21]

    Leonidas A. A. Doumas, Guillermo Puebla, Andrea E. Martin, and John E. Hummel. 2022. https://doi.org/10.1037/rev0000346 A theory of relation learning and cross-domain generalization . Psychological Review, 129:999--1041. Place: US Publisher: American Psychological Association

  14. [22]

    Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022. Toy models of sup...

  15. [23]

    Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dari...

  16. [24]

    Joël Fagot and Carole Parron. 2010. https://doi.org/10.1037/a0017169 Relational matching in baboons ( Papio papio) with reduced grouping requirements . Journal of Experimental Psychology: Animal Behavior Processes, 36(2):184--193. Place: US Publisher: American Psychological As...

  17. [25]

    Wasserman, and Michael E

    Joël Fagot, Edward A. Wasserman, and Michael E. Young. 2001. https://doi.org/10.1037/0097-7403.27.4.316 Discriminating the relation between relations: The role of entropy in abstract conceptualization by baboons ( Papio papio) and humans ( Homo sapiens) . Journal of Experiment...

  18. [26]

    Jerome Feldman. 2013. https://doi.org/10.1007/s11571-012-9219-8 The neural binding problem(s) . Cognitive Neurodynamics, 7(1):1--11

  19. [27]

    Jiahai Feng, Stuart Russell, and Jacob Steinhardt. 2024. https://openreview.net/forum?id=0yvZm2AjUr Monitoring Latent World States in Language Models with Propositional Probes

  20. [28]

    Jiahai Feng and Jacob Steinhardt. 2024. https://doi.org/10.48550/arXiv.2310.17191 How do Language Models Bind Entities in Context ? arXiv preprint. ArXiv:2310.17191 [cs]

  21. [29]

    McLaughlin

    Jerry Fodor and Brian P. McLaughlin. 1990. https://doi.org/10.1016/0010-0277(90)90014-B Connectionism and the problem of systematicity: Why Smolensky 's solution doesn't work . Cognition, 35(2):183--204

  22. [30]

    Fodor and Zenon W

    Jerry A. Fodor and Zenon W. Pylyshyn. 1988. https://doi.org/10.1016/0010-0277(88)90031-5 Connectionism and cognitive architecture: A critical analysis . Cognition, 28(1):3--71

  23. [31]

    Gottlob Frege. 1892. Über Sinn und Bedeutung , 1. auflage edition. Zeitschrift für Philosophie und philosophische Kritik , Neue Folge . Pfeffer, Leipzig

  24. [32]

    Pullum, and Ivan A

    Gerald Gazdar, Evan Klein, Geoffrey K. Pullum, and Ivan A. Sag. 1985. https://www.hup.harvard.edu/books/9780674344563 Generalized Phrase Structure Grammar . Harvard University Press

  25. [33]

    Wichmann

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020. https://doi.org/10.1038/s42256-020-00257-z Shortcut learning in deep neural networks . Nature Machine Intelligence, 2(11):665--673. Publisher:...

  26. [34]

    Adele Goldberg. 1995. Constructions: A construction grammar approach to argument structure . University of Chicago Press, Chicago ; London. Series Title: Cognitive theory of language and culture

  27. [35]

    Klaus Greff, Sjoerd van Steenkiste, and Jürgen Schmidhuber. 2020. https://doi.org/10.48550/arXiv.2012.05208 On the Binding Problem in Artificial Neural Networks . arXiv preprint. ArXiv:2012.05208 [cs]

  28. [36]

    Robert F. Hadley. 1994. https://doi.org/10.1111/j.1468-0017.1994.tb00225.x Systematicity in Connectionist Language Learning . Mind & Language, 9(3):247--272. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1468-0017.1994.tb00225.x

  29. [37]

    Per-Kristian Halvorsen and William A. Ladusaw. 1979. https://doi.org/10.1007/BF00126510 Montague's ‘universal grammar’: An introduction for the linguist . Linguistics and Philosophy, 3(2):185--223

  30. [38]

    Jaakko Hintikka. 1979. https://doi.org/10.1007/978-1-4020-4108-2_1 Language- Games . In Esa Saarinen, editor, Game- Theoretical Semantics : Essays on Semantics by Hintikka , Carlson , Peacocke , Rantala , and Saarinen , pages 1--26. Springer Netherlands, Dordrecht

  31. [39]

    Wilhelm von Humboldt. 1836. On Language : On the Diversity of Human Language Construction and Its Influence on the Mental Development of the Human Species . Google-Books-ID: \_UODbGlD4WUC

  32. [40]

    Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni. 2020. https://doi.org/10.48550/arXiv.1908.08351 Compositionality decomposed: how do neural networks generalise? arXiv preprint. ArXiv:1908.08351

  33. [41]

    Theo M. V. Janssen and Barbara H. Partee. 1997. https://doi.org/10.1016/B978-044481714-3/50011-4 Chapter 7 - Compositionality . In Johan van Benthem and Alice ter Meulen, editors, Handbook of Logic and Language , pages 417--473. North-Holland, Amsterdam

  34. [42]

    Lawrence Zitnick, and Ross Girshick

    Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. https://doi.org/10.1109/CVPR.2017.215 CLEVR : A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning . In 2017 IEEE Conference on Compu...

  35. [43]

    Aravind K. Joshi. 2005. https://doi.org/10.1093/oxfordhb/9780199276349.013.0026 Tree- Adjoining Grammars . In Ruslan Mitkov, editor, The Oxford Handbook of Computational Linguistics , page 0. Oxford University Press

  36. [44]

    jylin04 , JackS, Adam Karvonen, and Can . 2024. https://www.alignmentforum.org/posts/gcpNuEZnxAPayaKBY/othellogpt-learned-a-bag-of-heuristics-1 OthelloGPT learned a bag of heuristics

  37. [46]

    Najoung Kim and Tal Linzen. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.731 COGS : A Compositional Generalization Challenge Based on Semantic Interpretation . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 9087...

  38. [47]

    Yeongbin Kim, Gautam Singh, Junyeong Park, Caglar Gulcehre, and Sungjin Ahn. 2023. Imagine the unseen world: a benchmark for systematic generalization in visual world models. Advances in Neural Information Processing Systems, 36:27880--27896

  39. [48]

    Seijin Kobayashi, Simon Schug, Yassir Akram, Florian Redhardt, Johannes von Oswald, Razvan Pascanu, Guillaume Lajoie, and João Sacramento. 2024. https://doi.org/10.48550/arXiv.2407.12275 When can transformers compositionally generalize in-context? arXiv preprint. ArXiv:2407.12275 [cs]

  40. [49]

    Brenden Lake and Marco Baroni. 2018. https://proceedings.mlr.press/v80/lake18a.html Generalization without Systematicity : On the Compositional Skills of Sequence -to- Sequence Recurrent Networks . In Proceedings of the 35th International Conference on Machine Learning , pages...

  41. [50]

    Kevin J. Lande. 2021. https://doi.org/10.1111/nous.12324 Mental structures . Noûs, 55(3):649--677. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/nous.12324

  42. [51]

    Kevin J. Lande. 2024. https://doi.org/10.1002/wcs.1691 Compositionality in perception: A framework . WIREs Cognitive Science, 15(6):e1691. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/wcs.1691

  43. [52]

    Martha Lewis, Nihal Nayak, Peilin Yu, Jack Merullo, Qinan Yu, Stephen Bach, and Ellie Pavlick. 2024. https://aclanthology.org/2024.findings-eacl.101/ Does CLIP Bind Concepts ? Probing Compositionality in Large Image Models . In Findings of the Association for Computational Lin...

  44. [53]

    Bingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen, Yuekun Yao, and Najoung Kim. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.194 SLOG : A Structural Generalization Benchmark for Semantic Parsing . In Proceedings of the 2023 Conference on Empirical Methods in Natur...

  45. [54]

    Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg

    Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2024. https://doi.org/10.48550/arXiv.2210.13382 Emergent World Representations : Exploring a Sequence Model Trained on a Synthetic Task . arXiv preprint. ArXiv:2210.13382 [cs]

  46. [55]

    McClelland

    Yuxuan Li and James L. McClelland. 2022. https://doi.org/10.48550/arXiv.2210.00400 Systematic Generalization and Emergent Structures in Transformers Trained on Structured Tasks . arXiv preprint. ArXiv:2210.00400 [cs]

  47. [56]

    Weiduo Liao, Ying Wei, Mingchen Jiang, Qingfu Zhang, and Hisao Ishibuchi. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/hash/6a42b45af2b72e6e5b5e3a6fe695809f-Abstract-Datasets_and_Benchmarks.html Does Continual Learning Meet Compositionality ? New Benchmarks and ...

  48. [57]

    Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. 2020. Object-centric learning with slot attention. Advances in neural information processing systems, 33:11525--11538

  49. [58]

    R Duncan Luce. 1996. The ongoing dialog between empirical science and measurement theory. journal of mathematical psychology, 40(1):78--98

  50. [59]

    Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna. 2023. https://doi.org/10.1109/CVPR52729.2023.01050 CREPE : Can Vision - Language Foundation Models Reason Compositionally ? In 2023 IEEE / CVF Conference on Computer Vision and Pattern Recogni...

  51. [60]

    Tenenbaum, and Jiajun Wu

    Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum, and Jiajun Wu. 2019. https://doi.org/10.48550/arXiv.1904.12584 The Neuro - Symbolic Concept Learner : Interpreting Scenes , Words , and Sentences From Natural Supervision . arXiv preprint. ArXiv:1904.12584 [cs]

  52. [61]

    Thomas McCoy, Tal Linzen, Ewan Dunbar, and Paul Smolensky

    R. Thomas McCoy, Tal Linzen, Ewan Dunbar, and Paul Smolensky. 2020. https://aclanthology.org/2020.scil-1.34 Tensor Product Decomposition Networks : Uncovering Representations of Structure Learned by Neural Networks . In Proceedings of the Society for Computation in Linguistics...

  53. [62]

    Kate McCurdy, Paul Soulos, Paul Smolensky, Roland Fernandez, and Jianfeng Gao. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.524 Toward Compositional Behavior in Neural Models : A Survey of Current Views . In Proceedings of the 2024 Conference on Empirical Methods in Natur...

  54. [63]

    Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. 2024. https://doi.org/10.18653/v1/2024.naacl-long.281 Language Models Implement Simple Word2Vec -style Vector Arithmetic . In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computationa...

  55. [64]

    Arseny Moskvichev, Victor Vikram Odouard, and Melanie Mitchell. 2023. The conceptarc benchmark: Evaluating understanding and generalization in the arc domain. arXiv preprint arXiv:2305.07141

  56. [65]

    Neel Nanda, Andrew Lee, and Martin Wattenberg. 2023. https://doi.org/10.48550/arXiv.2309.00941 Emergent Linear Representations in World Models of Self - Supervised Sequence Models . arXiv preprint. ArXiv:2309.00941 [cs]

  57. [66]

    nostalgebraist . 2020. https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens interpreting GPT : the logit lens

  58. [67]

    Victor Vikram Odouard and Melanie Mitchell. 2022. Evaluating understanding on conceptual abstraction benchmarks. arXiv preprint arXiv:2206.14187

  59. [68]

    Dick, and Hidenori Tanaka

    Maya Okawa, Ekdeep Singh Lubana, Robert P. Dick, and Hidenori Tanaka. 2023. https://openreview.net/forum?id=ZXH8KUgFx3#all Compositional Abilities Emerge Multiplicatively : Exploring Diffusion Models on a Synthetic Task

  60. [69]

    Chris Olah. 2023. https://transformer-circuits.pub/2023/superposition-composition/index.html Distributed Representations : Composition & Superposition

  61. [70]

    Stevenson

    Gustaw Opiełka, Hannes Rosenbusch, and Claire E. Stevenson. 2025. https://doi.org/10.48550/arXiv.2503.03666 Analogical Reasoning Inside Large Language Models : Concept Vectors and the Limits of Abstraction . arXiv preprint. ArXiv:2503.03666 [cs]

  62. [71]

    Kiho Park, Yo Joong Choe, and Victor Veitch. 2024. The linear representation hypothesis and the geometry of large language models. In Proceedings of the 41st International Conference on Machine Learning , volume 235 of ICML '24 , pages 39643--39666, Vienna, Austria. JMLR.org

  63. [72]

    Barbara H. Partee. 2004. https://doi.org/10.1002/9780470751305.ch7 Compositionality . In Compositionality in Formal Semantics , pages 153--181. John Wiley & Sons, Ltd. Section: 7 \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/9780470751305.ch7

  64. [73]

    Ellie Pavlick. 2023. https://doi.org/10.1098/rsta.2022.0041 Symbols and grounding in large language models . Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 381(2251):20220041. Publisher: Royal Society

  65. [74]

    Jean Piaget. 2013. https://doi.org/10.4324/9781315009650 The Construction Of Reality In The Child . Routledge, London

  66. [75]

    David Premack and Guy Woodruff. 1978. https://doi.org/10.1017/S0140525X00076512 Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4):515--526

  67. [76]

    Jake Quilty-Dunn, Nicolas Porot, and Eric Mandelbaum. 2023. https://doi.org/10.1017/S0140525X22002849 The best game in town: The reemergence of the language-of-thought hypothesis across the cognitive sciences . Behavioral and Brain Sciences, 46:e261

  68. [77]

    Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2021. https://doi.org/10.48550/arXiv.2005.00719 Probing the Probing Paradigm : Does Probing Accuracy Entail Task Relevance ? arXiv preprint. ArXiv:2005.00719 [cs]

  69. [78]

    Williams, and Lotem Elber-Dorozko

    Jacob Russin, Sam Whitman McGrath, Danielle J. Williams, and Lotem Elber-Dorozko. 2024. https://doi.org/10.48550/arXiv.2405.15164 From Frege to chatGPT : Compositionality in language, cognition, and deep neural networks . arXiv preprint. ArXiv:2405.15164 [cs]

  70. [79]

    Lukas Schott, Julius Von Kügelgen, Frederik Träuble, Peter Vincent Gehler, Chris Russell, Matthias Bethge, Bernhard Schölkopf, Francesco Locatello, and Wieland Brendel. 2021. https://openreview.net/forum?id=9RUHPlladgh Visual Representation Learning Does Not Generalize Strongl...

  71. [80]

    Prithviraj Sen, Breno WSR de Carvalho, Ryan Riegel, and Alexander Gray. 2022. Neuro-symbolic inductive logic programming with logical neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 36, pages 8212--8219

  72. [81]

    Michaud, Stephen Casper, Max Tegmark, William Saunders, David Bau, Eric Todd, Atticus Geiger, Mor Geva, Jesse Hoogland, Daniel Murfet, and Tom McGrath

    Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, Stella Biderman, Adria Garriga-Alonso, Arthur Conmy, Neel Nanda, Jessica Rumbelow, Martin Wattenberg, Nandi Schoots, ...

  73. [82]

    Paul Soulos, Henry Conklin, Mattia Opper, Paul Smolensky, Jianfeng Gao, and Roland Fernandez. 2024. Compositional generalization across distributional shifts with sparse tree operations. arXiv preprint arXiv:2412.14076

  74. [83]

    Mark Steedman. 2019. https://doi.org/10.1515/9783110540253-014 Combinatory Categorial Grammar . In Current Approaches to Syntax . Publication Title: Current Approaches to Syntax

  75. [84]

    Stanley Smith Stevens. 1946. On the theory of scales of measurement. Science, 103(2684):677--680

  76. [85]

    Kaiser Sun, Adina Williams, and Dieuwke Hupkes. 2023. https://doi.org/10.18653/v1/2023.conll-1.19 The Validity of Evaluation Results : Assessing Concurrence Across Compositionality Benchmarks . In Proceedings of the 27th Conference on Computational Natural Language Learning ( ...

  77. [86]

    Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 a . https://doi.org/10.18653/v1/P19-1452 BERT Rediscovers the Classical NLP Pipeline . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4593--4601, Florence, Italy. Association ...

  78. [87]

    Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R

    Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 b . https://doi.org/10.48550/arXiv.1905.06316 What do you learn from context? Probing for sentence structure ...

  79. [88]

    Jonathan Thomm, Giacomo Camposampiero, Aleksandar Terzic, Michael Hersche, Bernhard Sch \"o lkopf, and Abbas Rahimi. 2024. Limits of transformer language models on learning to compose algorithms. In The Thirty-eighth Annual Conference on Neural Information Processing Systems

  80. [89]

    Roger K. R. Thompson, David L. Oden, and Sarah T. Boysen. 1997. https://doi.org/10.1037/0097-7403.23.1.31 Language-naive chimpanzees ( Pan troglodytes) judge relations between relations in a conceptual matching-to-sample task . Journal of Experimental Psychology: Animal Behavi...

  81. [90]

    Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. 2022. https://doi.org/10.1109/CVPR52688.2022.00517 Winoground: Probing Vision and Language Models for Visio - Linguistic Compositionality . In 2022 IEEE / CVF Conference on...

  82. [91]

    Li, Arnab Sen Sharma, Aaron Mueller, Byron C

    Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. 2024. https://doi.org/10.48550/arXiv.2310.15213 Function Vectors in Large Language Models . arXiv preprint. ArXiv:2310.15213 [cs]

  83. [92]

    Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan

    Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. 2024. https://doi.org/10.48550/arXiv.2406.03689 Evaluating the World Model Implicit in a Generative Model . arXiv preprint. ArXiv:2406.03689 [cs]

  84. [93]

    Martin Wattenberg and Fernanda B. Viégas. 2024. https://doi.org/10.48550/arXiv.2407.14662 Relational Composition in Neural Networks : A Survey and Call to Action . arXiv preprint. ArXiv:2407.14662 [cs]

  85. [94]

    Dag Westerståhl. 1998. https://www.jstor.org/stable/25001726 On Mathematical Proofs of the Vacuity of Compositionality . Linguistics and Philosophy, 21(6):635--643. Publisher: Springer

  86. [95]

    Manning, and Christopher Potts

    Zhengxuan Wu, Christopher D. Manning, and Christopher Potts. 2023. https://doi.org/10.1162/tacl_a_00623 ReCOGS : How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation . Transactions of the Association for Computational Linguistics, 11:1719--1733

  87. [96]

    Young and Edward A

    Michael E. Young and Edward A. Wasserman. 1997. https://doi.org/10.1037/0097-7403.23.2.157 Entropy detection by pigeons: Response to mixed visual displays after same–different discrimination training . Journal of Experimental Psychology: Animal Behavior Processes, 23(2):157--1...

  88. [97]

    Young and Edward A

    Michael E. Young and Edward A. Wasserman. 2001. https://doi.org/10.1037/0278-7393.27.1.278 Entropy and variability discrimination . Journal of Experimental Psychology: Learning, Memory, and Cognition, 27(1):278--293. Place: US Publisher: American Psychological Association

  89. [98]

    Aimen Zerroug, Mohit Vaishnav, Julien Colin, Sebastian Musslick, and Thomas Serre. 2022. A benchmark for compositional visual reasoning. Advances in neural information processing systems, 35:29776--29788

  90. [99]

    Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, and Song-Chun Zhu. 2019. Raven: A dataset for relational and analogical visual reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5317--5327

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.