REVIEW 2 cited by
ABCNN: Attention-Based Convolutional Neural Network for Modeling Sentence Pairs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
How to model a pair of sentences is a critical issue in many NLP tasks such as answer selection (AS), paraphrase identification (PI) and textual entailment (TE). Most prior work (i) deals with one individual task by fine-tuning a specific system; (ii) models each sentence's representation separately, rarely considering the impact of the other sentence; or (iii) relies fully on manually designed, task-specific linguistic features. This work presents a general Attention Based Convolutional Neural Network (ABCNN) for modeling a pair of sentences. We make three contributions. (i) ABCNN can be applied to a wide variety of tasks that require modeling of sentence pairs. (ii) We propose three attention schemes that integrate mutual influence between sentences into CNN; thus, the representation of each sentence takes into consideration its counterpart. These interdependent sentence pair representations are more powerful than isolated sentence representations. (iii) ABCNN achieves state-of-the-art performance on AS, PI and TE tasks.
Forward citations
Cited by 2 Pith papers
-
Representing text as abstract images enables image classifiers to also simultaneously classify text
Converting text pairs into abstract RGB images lets an image classifier perform inventor name disambiguation, achieving F1 of 99.09% on the IS and E&S benchmark datasets.
-
A Sensitivity Analysis of Attention-Gated Convolutional Neural Networks for Sentence Classification
A hyperparameter sensitivity study of AGCNN for sentence classification, with tuned settings improving accuracy by roughly 0.45 to 0.81 percentage points on the same six datasets used for tuning.
Discussion (0). Continue with ORCID to comment.