REVIEW 2 cited by
Structural generalization is hard for sequence-to-sequence models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sequence-to-sequence (seq2seq) models have been successful across many NLP tasks, including ones that require predicting linguistic structure. However, recent work on compositional generalization has shown that seq2seq models achieve very low accuracy in generalizing to linguistic structures that were not seen in training. We present new evidence that this is a general limitation of seq2seq models that is present not just in semantic parsing, but also in syntactic parsing and in text-to-text tasks, and that this limitation can often be overcome by neurosymbolic models that have linguistic knowledge built in. We further report on some experiments that give initial answers on the reasons for these limitations.
Forward citations
Cited by 2 Pith papers
-
Randomly Sampled Language Reasoning Problems Elucidate Limitations of In-Context Learning
On randomly sampled 3-state DFA language tasks, foundation LLMs underperform n-gram baselines under pure in-context-learning prompts.
-
Infusing Prompts with Syntax and Semantics
Appending syntactic and semantic analyses to prompts improves text-to-SQL accuracy in four low-resource languages and speeds fine-tuning.
Discussion (0). Continue with ORCID to comment.