Pith. sign in

REVIEW 3 major objections 4 minor 235 references

Reasoning-Driven Question-Answering for Natural Language Understanding

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read QA systems that reason over semantic abstractions beat retrieval and neural baselines on science and biology exams, and a formal model shows why multi-step reasoning has limits.

desk verdict A solid compilation of previously published empirical work whose only new piece—the formal theory chapter—is missing from the supplied text, leaving the strongest claim unverified. read the letter →

arxiv 1908.04926 v1 pith:YUMXHJPX submitted 2019-08-14 cs.CL cs.AI

classification cs.CLcs.AI
keywords naturallanguageunderstandingquestionansweringabductivereasoningsupportgraphoptimizationintegerlinearprogrammingsemanticabstractionstemporalcommonsenselimitations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This thesis tries to establish that the bottleneck in natural language understanding (NLU) is not data volume but the ability to reason over abstract representations of meaning. It defends the claim with three moves: constrained-optimization QA systems that chain facts across tables and semantic graphs outperform retrieval and neural baselines on science and biology questions; two new datasets force multi-sentence and temporal-common-sense reasoning and expose large gaps to human performance; and a formal model of the meaning-symbol interface proves that multi-step reasoning algorithms face fundamental limits when symbols are noisy, incomplete, and ambiguous. If the thesis is right, progress in NLU comes from building reasoning and grounding into QA systems rather than from scaling corpora or models alone.

What carries the argument

The load-bearing object is the support graph: a subgraph of an augmented graph whose nodes are question constituents, answer options, and knowledge units (table cells or semantic-graph nodes), with edges weighted by entailment or similarity. An ILP formulation selects the support graph that maximizes weighted alignments while enforcing connectivity, evidence-chaining, and semantic-relation constraints; the same template powers the two system chapters, over tables and over semantic graphs. For the limitations result, the carrying object is the meaning-symbol interface: a two-layer model in which a clean, unique meaning space is observed only through a noisy, incomplete, variable symbol space. The proof of limits uses a cut-based construction that separates meaning pairs that are connected from those that are disconnected, showing that any algorithm relying on local symbol-graph distances must fail in the noisy regime.

What would settle it

One concrete check is to assemble a benchmark whose instances are independently verified to require multi-sentence chaining and temporal common sense, and then show that a model with no explicit reasoning component, trained on enough data, matches or exceeds human accuracy; that outcome would contradict the thesis's claim that reasoning over abstractions and world knowledge is needed for QA progress.

Watch

Extended reading notes

Core claim

On the author's own terms, the central claim is that question answering should be treated as abductive reasoning: the system must find the best support graph connecting a question to exactly one answer through available knowledge, where 'best' is defined by structural constraints and soft preferences over alignments. Casting this search as an integer linear program lets the same machinery operate over curated tables, relation-extraction tuples, and multi-view semantic graphs, and it beats the retrieval and neural baselines on unseen exam questions. The theoretical companion claim is that reasoning over natural language happens in a noisy symbol space that only approximates a clean meaning space; the thesis constructs a connectivity-reasoning model and proves both when accurate recovery of meaning-space connectivity is possible and when it is provably impossible. This is offered as the first formal framework for multi-step reasoning algorithms under incompleteness, ambiguity, and variability.

Load-bearing premise

The load-bearing premise is that accuracy on static question-answering benchmarks measures real progress toward natural language understanding, even though the thesis itself concedes that such benchmarks are skewed toward simplicity and give a biased estimate of the space of questions.

Editorial extensions

If this is right

  • A QA system that explicitly chains evidence through structured abstractions can outperform broad-coverage retrieval and a specialized neural reader on small-data reasoning domains, with 2–6 percent absolute gains on science exams and near-parity with a domain-specific biology system.
  • Because the same optimization template is applied to tables, tuples, and semantic graphs, new semantic abstractions can be added to the framework without changing the reasoning machinery.
  • Forcing solvers to use question terms that are predicted to be essential makes them more robust to distractors; the thesis reports up to 5 percent absolute gains for a retrieval solver and a 41.7 percent error reduction on a curated hard set.
  • The proposed multi-sentence and temporal-common-sense benchmarks imply that existing datasets understate the difficulty of NLU, and systems trained on current benchmarks should show a large gap to human performance on these new instances.
  • The formal limitation result implies that brittleness in multi-step reasoning over language is not only an engineering problem: within the model's assumptions, no symbol-space algorithm can always recover meaning-space connectivity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the formal limitation suggests a lower-bound-style claim—within the model, more training data alone cannot remove the ambiguity injected by the symbol space, so progress will require grounding or additional structured world knowledge.
  • Editorial inference: the support-graph/ILP formulation can be hybridized with modern neural models by using neural similarity scores as edge weights and keeping the ILP as a trainable inference layer; the thesis provides a clean interface for that combination.
  • Editorial inference: the essential-terms study points to a cheaper supervision signal—annotating which terms matter instead of full answers—that could transfer to other NLU tasks and could be used to audit neural attention mechanisms.
  • Editorial inference: the thesis's own warning that static benchmarks are biased toward simplicity implies a testable research program: build evaluation sets by construction, verifying that each instance requires multi-step chaining, and use those sets to measure progress rather than relying on sampled corpora.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This dissertation-style manuscript investigates natural language understanding (NLU) through the question answering (QA) task and is organized into three parts. Part I develops reasoning-driven QA systems: TableILP, which casts QA as an integer linear program over semi-structured tabular knowledge and is evaluated on elementary science exams; SemanticILP, which generalizes this formulation to raw text by reasoning over semantic abstractions from off-the-shelf NLP tools; and a supervised essential-term classifier that identifies critical question words and is shown to improve an IR-based solver. Part II introduces two challenge datasets: MultiRC, requiring multi-sentence reading comprehension, and TacoQA, requiring temporal commonsense reasoning. Part III announces a formal framework for multi-step reasoning algorithms and claims to prove fundamental limitations of such algorithms under properties of language use such as incompleteness and ambiguity. The abstract and introduction present this theoretical contribution as a central claim, but the corresponding Chapter 8 is absent from the provided text; only its table-of-contents entry and references to it in earlier chapters are visible. The empirical chapters present clear experimental designs with baselines, ablations, and significance tests, and several datasets and code bases are released publicly.

Significance. If the results hold, the thesis offers a coherent and useful body of work: an ILP-based reasoning framework that outperforms structured baselines on science QA with limited training data; a crowd-sourced dataset and classifier for essential question terms; two QA datasets that push beyond single-sentence and static-benchmark settings; and, potentially, a formal analysis of the limits of reasoning algorithms in natural language. The empirical chapters are carefully presented: they include baselines, ablations, statistical significance tests, hidden test sets, and public releases of code and datasets, which are strengths that should be acknowledged. The most distinctive claim, however, is the theoretical framework of Chapter 8, which cannot be inspected in the current manuscript because the chapter body and Appendix A.2 are missing. Since Part III is the only part of the thesis that is not already published as peer-reviewed empirical work, the headline theoretical claim is unverified as submitted, and the overall significance of the thesis cannot be fully assessed without it.

major comments (3)
  1. [Chapter 8 / Appendix A.2] The central theoretical claim of the dissertation—presenting 'the first formal framework for multi-step reasoning algorithms' and proving 'fundamental limitations' for reasoning algorithms—is stated in the abstract and introduction, and Chapters 2 and 3 explicitly defer to Chapter 8 (e.g., Section 3.5 says the brittleness of multi-step reasoning is studied in Chapter 8). However, the submitted text contains only the table-of-contents entry for Chapter 8; the chapter body and its supplementary appendix (A.2) are not included. I therefore cannot inspect the meaning-symbol interface, the noise model (epsilon+, p-), Definition 9, or the proof in Section 8.6. This is a load-bearing omission because Part III is the only part of the thesis that is not already published as peer-reviewed empirical work. The authors must supply the complete text of Chapter 8 and Appendix A.2 (or the full content of the corresponding publication) so that the theoretical claims can be evaluated.
  2. [Section 5.3.1 / Table 18] The text states that the ET classifier 'has a 5% higher AUC (area under the curve)' relative to baselines, but Table 18 reports AUC of 0.79 for the ET Classifier, equal to PropSurf and lower than PropLem's 0.80. This numerical inconsistency contradicts the table as printed. The sentence must be corrected and the AUC claim reworded to match the reported data; the F1 and MAP advantages remain supported by the tables, but the AUC statement is not.
  3. [Section 5.4.2] The demonstration that essentiality information improves the TableILP solver is based on only 12 curated questions (the QR set). The reported '41.7% error reduction' corresponds to correcting 5 of 12 errors made by vanilla TableILP. This sample is too small to support the general conclusion that the ET cascade helps TableILP cope with distracting terms. The section should either be expanded with a larger evaluation or explicitly framed as a small case study with limited statistical power, so that readers are not misled about the strength of the evidence.
minor comments (4)
  1. [Section 1.4 and cross-references] Several internal cross-references are inconsistent with the actual chapter numbering. The thesis outline says Chapter 5 presents MultiRC and Chapter 6 presents TacoQA, but the actual chapters are 6 and 7 respectively; Section 2.4.5 refers to 'Chapter 6' for temporal commonsense reasoning, which is actually Chapter 7. These should be corrected throughout.
  2. [Section 2.3.1] The sentence 'In Chapter 2, 3 we use elementary-school science tests' appears to contain a typo; it should presumably read 'In Chapters 3 and 4, we use elementary-school science tests.'
  3. [Section 4.5] There are grammatical errors in the final paragraph: 'a major portion of our understanding come is only implied from text' and 'lack explicit explicit attention' contain typos and should be proofread.
  4. [Section 5.3.1] The phrase 'Binomial 10 exact test' appears to be a typo; it should read 'binomial exact test.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical claims are benchmark-tested and self-contained; Chapter 8 is missing but not circular.

full rationale

The derivation chain in this thesis is empirical rather than deductive. The systems (TableILP, SemanticILP, ET classifier) are evaluated on held-out standardized exams, ProcessBank, and crowd-annotated essentiality data, with public release of datasets and code; ablations identify which components matter. No prediction is obtained by fitting a parameter to the same quantity it is said to predict: the ET classifier is trained on human annotations and evaluated on disjoint test terms, and the IR+ET threshold is tuned on training questions while reported scores are on test sets. The MultiRC and TacoQA datasets are constructed with explicit verification steps (multi-sentence validation via crowd workers) rather than being derived from the systems being benchmarked. The only potentially load-bearing self-referential element is Chapter 8's formal theory, which is absent from the supplied text; however, the abstract's claim is not shown to reduce to an equation in the input, and absence of the chapter is a completeness issue, not circularity. Self-citations in the text are provenance for previously published chapters and toolkits (e.g., CogCompNLP), and are not used to justify the empirical conclusions, which are reproduced in the thesis itself. Therefore no significant circularity is established.

Assumptions & free parameters 6 free parameters · 4 assumptions · 2 invented entities

The central claims of the thesis rest on hand-set optimization weights, thresholds, and model parameters that are not fully specified in the provided text. The empirical evaluations rely on QA benchmarks as the measurement protocol and on the reliability of NLP annotation tools, both of which are acknowledged to be imperfect. The theoretical chapter introduces a new conceptual model with its own assumptions. These are listed above as free parameters, axioms, and invented entities.

free parameters (6)
  • TableILP objective weights = not stated in excerpt
    The ILP objective in Chapter 3 sums weighted variables; weights are tuned on the development set (Tables 27-29 in appendix).
  • Alignment thresholds (e.g., MinCellCellAlignment) = not stated in excerpt
    Thresholds control which pairwise variables are created; listed in Table 28 of the appendix.
  • Knowledge filtering counts (top 7 tables, 20 rows) = 7 tables, 20 rows
    Hand-chosen filtering constants in Chapter 3.4.1.
  • ET classifier threshold xi = 0.36 for IR+ET; cascades (0.4, 0.6, 0.8, 1.0)
    Selected by optimizing end-to-end performance on training/dev sets (Chapter 5.4).
  • SemanticILP ensemble weights = not stated
    Weights for combining simplified solver scores are trained on the union of training data (Chapter 4.4.2).
  • Noise parameters in Chapter 8 model (epsilon+, p-, lambda) = epsilon+ = 0.7, lambda = 3 used in Figure 29
    Parameters of the meaning-symbol interface model; appear in the theoretical analysis.
assumptions (4)
  • domain assumption QA over static datasets is a valid proxy for NLU progress
    The thesis uses QA benchmarks as the measurement protocol (Section 2.3.1) while acknowledging datasets are biased samples of the universal question space.
  • domain assumption NLP semantic annotators (SRL, coreference, etc.) produce sufficiently accurate abstractions for reasoning
    Chapter 4 builds the semantic graph representation on these tools and assumes their outputs are reliable enough to support QA.
  • standard math ILP with industrial solvers is a practical optimization framework for NLP
    Chapter 2.6.4 states ILP is NP-hard in general but solvers are fast in practice; the thesis relies on SCIP.
  • ad hoc to paper The meaning-symbol interface model in Chapter 8 captures essential properties of language (incompleteness, ambiguity)
    The theoretical framework is introduced in the thesis; its assumptions are not fully visible in the provided excerpt.
invented entities (2)
  • Essential question terms independent evidence
    purpose: A token-level importance label used to filter queries and focus reasoning in QA systems
    Introduced in Chapter 5 with a crowd-sourced dataset of 19k annotated terms and public classifier; human agreement and QA improvements provide external handles.
  • Support graph
    purpose: A subgraph that connects the question to one answer via aligned knowledge, used to define the reasoning objective
    Defined in Chapters 3 and 4; purely an internal optimization construct with no external falsifiable handle.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Reasoning-Driven Question-Answering for Natural Language Understanding." pith.science (2026). https://pith.science/paper/YUMXHJPX

@misc{pith2026190804926,
  author       = {Pith},
  title        = {Pith review of: Reasoning-Driven Question-Answering for Natural Language Understanding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YUMXHJPX}},
  note         = {Machine review of arXiv:1908.04926}
}
read the original abstract

Natural language understanding (NLU) of text is a fundamental challenge in AI, and it has received significant attention throughout the history of NLP research. This primary goal has been studied under different tasks, such as Question Answering (QA) and Textual Entailment (TE). In this thesis, we investigate the NLU problem through the QA task and focus on the aspects that make it a challenge for the current state-of-the-art technology. This thesis is organized into three main parts: In the first part, we explore multiple formalisms to improve existing machine comprehension systems. We propose a formulation for abductive reasoning in natural language and show its effectiveness, especially in domains with limited training data. Additionally, to help reasoning systems cope with irrelevant or redundant information, we create a supervised approach to learn and detect the essential terms in questions. In the second part, we propose two new challenge datasets. In particular, we create two datasets of natural language questions where (i) the first one requires reasoning over multiple sentences; (ii) the second one requires temporal common sense reasoning. We hope that the two proposed datasets will motivate the field to address more complex problems. In the final part, we present the first formal framework for multi-step reasoning algorithms, in the presence of a few important properties of language use, such as incompleteness, ambiguity, etc. We apply this framework to prove fundamental limitations for reasoning algorithms. These theoretical results provide extra intuition into the existing empirical evidence in the field.

Figures

Figures reproduced from arXiv: 1908.04926 by the authors.

Figure 1
Figure 1. FIGURE 1 : A sample story appeared on the New York Times (taken from Mc [PITH_FULL_IMAGE:figures/full_fig_p015_1.png] view at source ↗
Figure 6
Figure 6. FIGURE 6 : A hypothetical manifold of all the NLU instances. Static datasets [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 11
Figure 11. FIGURE 11 [PITH_FULL_IMAGE:figures/full_fig_p016_11.png] view at source ↗
Figures from the paper (37 more)
Figure 12
Figure 12. Figure 12: FIGURE 12 : Overlap of the predictions of [PITH_FULL_IMAGE:figures/full_fig_p016_12.png]
Figure 18
Figure 18. Figure 18: FIGURE 18 [PITH_FULL_IMAGE:figures/full_fig_p017_18.png]
Figure 20
Figure 20. Figure 20: FIGURE 20 : Pipeline of our dataset construction. . . . . . . . . . . . . . . . . . 92 [PITH_FULL_IMAGE:figures/full_fig_p017_20.png]
Figure 28
Figure 28. Figure 28: FIGURE 28 [PITH_FULL_IMAGE:figures/full_fig_p018_28.png]
Figure 30
Figure 30. Figure 30: FIGURE 30 : Notation for the ILP formulation. . . . . . . . . . . . . . . . . . . 138 [PITH_FULL_IMAGE:figures/full_fig_p018_30.png]
Figure 1
Figure 1. Figure 1: A sample story appeared on the New York Times (taken from McCarthy (1976)). [PITH_FULL_IMAGE:figures/full_fig_p021_1.png]
Figure 2
Figure 2. Figure 2: Ambiguity (left) appears when mapping a raw string to its actual meaning; Variability (right) is having many ways of referring to the same meaning. 2 [PITH_FULL_IMAGE:figures/full_fig_p021_2.png]
Figure 3
Figure 3. Figure 3: Visualization of two semantic tasks for the given story in Figure 1. Top fig [PITH_FULL_IMAGE:figures/full_fig_p024_3.png]
Figure 4
Figure 4. Figure 4: An overview of the contributions and challenges addressed in each chapter of this [PITH_FULL_IMAGE:figures/full_fig_p026_4.png]
Figure 5
Figure 5. Figure 5: Each highlight is color-coded to indicate its contribution type. In the following [PITH_FULL_IMAGE:figures/full_fig_p027_5.png]
Figure 5
Figure 5. Figure 5: Major highlights of NLU in the past 50 years (within the AI community). For [PITH_FULL_IMAGE:figures/full_fig_p028_5.png]
Figure 6
Figure 6. Figure 6: A hypothetical manifold of all the NLU instances. Static datasets make it easy [PITH_FULL_IMAGE:figures/full_fig_p031_6.png]
Figure 7
Figure 7. Figure 7: Example frames used in this work. Generic basic science frames (left), used in [PITH_FULL_IMAGE:figures/full_fig_p034_7.png]
Figure 8
Figure 8. Figure 8: Brief definitions for popular reasoning classes and their examples. [PITH_FULL_IMAGE:figures/full_fig_p039_8.png]
Figure 9
Figure 9. Figure 9: TableILP searches for the best support graph (chains of rea￾soning) connecting the question to an answer, in this case June. Con￾straints on the graph define what constitutes valid support and how to score it (Section 3.3.3). Further, we would like the system to be ro￾…
Figure 10
Figure 10. Figure 10: Depiction of SemanticILP reasoning for the example paragraph given in the text. Semantic abstractions of the question, answers, knowledge snippet are shown in different colored boxes (blue, green, and yellow, resp.). Red nodes and edges are the elements that are align…
Figure 11
Figure 11. Figure 11: Knowledge Representation used in our formulation. Raw text is associated with a collection of Semantic￾Graphs, which convey certain informa￾tion about the text. There are implicit similarity edges among the nodes of the connected components of the graphs, and from nod…
Figure 12
Figure 12. Figure 12: Overlap of the predictions of Se￾manticILP and IR on 50 randomly-chosen questions from AI2Public 4th. Complementarity to IR. Given that in the science domain the input snippets fed to SemanticILP are retrieved through a process similar to the IR solver, one might natu…
Figure 13
Figure 13. Figure 13: Performance change for varying knowledge length. [PITH_FULL_IMAGE:figures/full_fig_p084_13.png]
Figure 14
Figure 14. Figure 14: Essentiality scores generated by our system, which assigns high essentiality to “drop” and “temperature”. Towards this goal, we propose a system that can assign an essentiality score to each term in the question. For the above example, our sys￾tem generates the scores…
Figure 15
Figure 15. Figure 15: Crowd-sourcing interface for annotating essential terms in a question, including [PITH_FULL_IMAGE:figures/full_fig_p091_15.png]
Figure 16
Figure 16. Figure 16: Crowd-sourcing interface for verifying the validity of essentiality annotations [PITH_FULL_IMAGE:figures/full_fig_p092_16.png]
Figure 17
Figure 17. Figure 17: The relationship between the fraction of question words dropped and the fraction of the questions attempted (fraction of the questions workers felt comfortable answering). Dropping most essential terms (blue lines) results in very few questions remain￾ing answerable, …
Figure 18
Figure 18. Figure 18: Precision-recall trade-off for various classifiers as the threshold is varied. ET classifier (green) is significantly better throughout. As noted earlier, each of these essentiality identification methods are parameterized by a threshold for balancing precision and re…
Figure 19
Figure 19. Figure 19: Examples from our MultiRCcorpus. Each example shows relevant excerpts from a paragraph; multi-sentence question that can be answered by com￾bining information from multiple sentences of the para￾graph; and corresponding answer-options. The correct answer(s) is indicat…
Figure 20
Figure 20. Figure 20: gives a high-level idea of the process. The first two steps deal with creating multi-sentence questions, followed by two steps for construction of candidate answers. Step 1: generating multi-sentence questions given paragraphs Step 2: Verifying multi-sentenceness Step…
Figure 21
Figure 21. Figure 21: Distribution of (left) general phenomena; (right) variations of the “coreference” [PITH_FULL_IMAGE:figures/full_fig_p116_21.png]
Figure 22
Figure 22. Figure 22: Most frequent first chunks of the questions (counts in log scale). [PITH_FULL_IMAGE:figures/full_fig_p117_22.png]
Figure 23
Figure 23. Figure 23: PR curve for each of the baselines. There is a considerable gap with the baselines and human. To get a sense of our dataset’s hardness, we evaluate both human performance and mul￾tiple computational baselines. Each base￾line scores an answer-option with a real￾valued …
Figure 24
Figure 24. Figure 24: Five types of temporal commonsense in TacoQA. Note that a question may have multiple answers. 103 [PITH_FULL_IMAGE:figures/full_fig_p122_24.png]
Figure 25
Figure 25. Figure 25: BERT + unit normalization performance per temporal reasoning category (top), per￾formance gain over random baseline per category (bottom) We then showed that systems equipped with state-of-the-art language models such as ELMo and BERT are still far behind humans, thus…
Figure 26
Figure 26. Figure 26: The interface between meanings and symbols: each meaning (top) can be uttered in many ways into symbolic forms (bottom). While there is a rich literature on reason￾ing, there is little understanding of the na￾ture of the problem and its limitations, especially in the …
Figure 27
Figure 27. Figure 27: The meaning space contains [clean and unique] symbolic representation and the facts, while the symbol space contains [noisy, incomplete and variable] representation of the facts. We show sample meaning and symbol space nodes to answer the question: Is a metal spoon a …
Figure 28
Figure 28. Figure 28: The construction considered in Definition 9. The node-pair m-m0 is connected with distance d in GM, and dis￾connected in G0 M, after dropping the edges of a cut C. For each symbol graph, we consider it “local” Laplacian. Consider a meaning graph GM in which two nodes …
Figure 29
Figure 29. Figure 29: Various colors in the figure depict the average distance between node-pairs in the [PITH_FULL_IMAGE:figures/full_fig_p150_29.png]
Figure 30
Figure 30. Figure 30: Notation for the ILP formulation [PITH_FULL_IMAGE:figures/full_fig_p157_30.png]
Figure 31
Figure 31. Figure 31: With varied values for p− a heat map representation of the distribution of the av￾erage distances of node-pairs in symbol graph based on the distances of their corresponding meaning nodes is presented. With g being Lipschitz, one can upper-bound the variations on the …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

235 extracted references · 72 canonical work pages

  1. [1]

    Achterberg

    T. Achterberg. SCIP: solving constraint integer programs . Math. Prog. Computation, 1 0 (1): 0 1--41, 2009

  2. [2]

    Angeli and C

    G. Angeli and C. D. Manning. NaturalLI: Natural Logic Inference for Common Sense Reasoning . In Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), 2014

  3. [3]

    Arivazhagan, C

    N. Arivazhagan, C. Christodoulopoulos, and D. Roth. Labeling the semantic roles of commas. In AAAI, 2016

  4. [4]

    C. F. Baker, C. J. Fillmore, and J. B. Lowe. The berkeley framenet project . In Proc. of the Annual Meeting of the Association of Computational Linguistics (ACL), pages 86--90, 1998

  5. [5]

    Bamman, B

    D. Bamman, B. O'Connor, and N. A. Smith. Learning Latent Personas of Film Characters . In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics, ACL 2013, Volume 1: Long Papers , pages 352--361, 2013. URL http://aclweb.org/anthology/P/P13/P13-1035.pdf

  6. [6]

    Banarescu, C

    L. Banarescu, C. Bonial, S. Cai, M. Georgescu, K. Griffitt, U. Hermjakob, K. Knight, M. Palmer, and N. Schneider. Abstract meaning representation for sembanking . In Linguistic Annotation Workshop and Interoperability with Discourse, 2013

  7. [7]

    Banko, M

    M. Banko, M. J. Cafarella, S. Soderland, M. Broadhead, and O. Etzioni. Open Information Extraction from the Web . In Proc. of the International Joint Conference on Artificial Intelligence (IJCAI), 2007

  8. [8]

    Bar-Haim, I

    R. Bar-Haim, I. Dagan, and J. Berant. Knowledge-Based Textual Inference via Parse-Tree Transformations. J. Artif. Intell. Res.(JAIR), 54: 0 1--57, 2015

Show all 235 references
  1. [9]

    Bauer, Y

    L. Bauer, Y. Wang, and M. Bansal. Commonsense for Generative Multi-Hop Question Answering Tasks . In Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), pages 4220--4230, 2018

  2. [10]

    Bentivogli, P

    L. Bentivogli, P. Clark, I. Dagan, and D. Giampiccolo. The Sixth PASCAL Recognizing Textual Entailment Challenge . In TAC, 2008

  3. [11]

    Berant, I

    J. Berant, I. Dagan, and J. Goldberger. Global learning of focused entailment graphs . In Proc. of the Annual Meeting of the Association of Computational Linguistics (ACL), pages 1220--1229, 2010

  4. [12]

    Berant, V

    J. Berant, V. Srikumar, P.-C. Chen, A. V. Linden, B. Harding, B. Huang, P. Clark, and C. D. Manning. Modeling Biological Processes for Reading Comprehension. In Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), 2014

  5. [13]

    V. W. Berninger, W. Nagy, and S. Beers. Child writers’ construction and reconstruction of single sentences and construction of multi-sentence texts: Contributions of syntax and transcription to translation . Reading and writing, 24 0 (2): 0 151--182, 2011

  6. [14]

    A. M. Bisantz and K. J. Vicente. Making the abstraction hierarchy concrete . International Journal of human-computer studies, 40 0 (1): 0 83--117, 1994

  7. [15]

    D. G. Bobrow. Natural language input for a computer problem solving system . Technical report, MIT, 1964

  8. [16]

    Bollacker, C

    K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor. Freebase : A collaboratively created graph database for structuring human knowledge . In ICMD, pages 1247--1250. ACM, 2008

  9. [17]

    Brachman, D

    R. Brachman, D. Gunning, S. Bringsjord, M. Genesereth, L. Hirschman, and L. Ferro. Selected Grand Challenges in Cognitive Science . Technical report, MITRE Technical Report 05-1218, 2005

  10. [18]

    Brill, S

    E. Brill, S. Dumais, and M. Banko. An analysis of the AskMSR question-answering system . In Proceedings of EMNLP, pages 257--264, 2002

  11. [19]

    P. F. Brown, P. V. Desouza, R. L. Mercer, V. J. D. Pietra, and J. C. Lai. Class-based n-gram models of natural language . Computational linguistics, 18 0 (4): 0 467--479, 1992

  12. [20]

    J. G. Carbonell and R. D. Brown. Anaphora resolution: a multi-strategy approach . In Proceedings of the 12th conference on Computational linguistics-Volume 1, pages 96--101. Association for Computational Linguistics, 1988

  13. [21]

    Carlson, J

    A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. R. H. Jr., and T. M. Mitchell. Toward an Architecture for Never-Ending Language Learning . In Proceedings of the National Conference on Artificial Intelligence (AAAI), 2010

  14. [22]

    Chang, S

    K.-W. Chang, S. Upadhyay, M.-W. Chang, V. Srikumar, and D. Roth. Illinois-SL : A JAVA library for structured prediction . arXiv preprint arXiv:1509.07179, 2015

  15. [23]

    Chang, L

    M.-W. Chang, L. Ratinov, N. Rizzolo, and D. Roth. Learning and inference with constraints. In Proc. of the Conference on Artificial Intelligence (AAAI), 7 2008. URL http://cogcomp.org/papers/CRRR08.pdf

  16. [24]

    Chang, D

    M.-W. Chang, D. Goldwasser, D. Roth, and V. Srikumar. Discriminative Learning over Constrained Latent Representations . Proceedings of Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics (HLT 20...

  17. [25]

    Chang, L

    M.-W. Chang, L. Ratinov, and D. Roth. Structured learning with constrained conditional models. Machine Learning, 88 0 (3): 0 399--431, 6 2012. URL http://cogcomp.org/papers/ChangRaRo12.pdf

  18. [26]

    D. Chen, J. Bolton, and C. D. Manning. A Thorough Examination of the CNN/Daily Mail Reading Comprehension Task . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016, Volume 1: Long Papers , 2016. URL http://aclweb.org/anthology/...

  19. [27]

    Q. Chen, X. Zhu, Z. Ling, S. Wei, H. Jiang, and D. Inkpen. Enhanced LSTM for Natural Language Inference . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (ACL 2017), Vancouver, July 2017. ACL

  20. [28]

    Chklovski and P

    T. Chklovski and P. Pantel. VerbOcean: Mining the Web for Fine-Grained Semantic Verb Relations . In EMNLP, 2004

  21. [29]

    Chung and L

    F. Chung and L. Lu. The average distances in random graphs with given expected degrees . Proceedings of the National Academy of Sciences, 99 0 (25): 0 15879--15882, 2002

  22. [30]

    K. W. Church and P. Hanks. Word Association Norms, Mutual Information and Lexicography . In 27thACL, pages 76--83, 1989

  23. [31]

    P. Clark. Elementary School Science and Math Tests as a Driver for AI: Take the A risto Challenge! In 29th AAAI/IAAI, pages 4019--4021, Austin, TX, 2015

  24. [32]

    Clark and O

    P. Clark and O. Etzioni. My Computer is an Honor Student — but how Intelligent is it? Standardized Tests as a Measure of AI . AI Magazine , 2016. (To appear)

  25. [33]

    Clark, N

    P. Clark, N. Balasubramanian, S. Bhakthavatsalam, K. Humphreys, J. Kinkead, A. Sabharwal, and O. Tafjord. Automatic Construction of Inference-Supporting Knowledge Bases . In 4thAKBC Workshop, Montreal, Canada, 2014

  26. [34]

    Clark, O

    P. Clark, O. Etzioni, T. Khot, A. Sabharwal, O. Tafjord, P. Turney, and D. Khashabi. Combining Retrieval, Statistics, and Inference to Answer Elementary Science Questions . In Proceedings of the National Conference on Artificial Intelligence (AAAI), 2016

  27. [35]

    Clark, I

    P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge . CoRR, abs/1803.05457, 2018

  28. [36]

    Clarke and M

    J. Clarke and M. Lapata. Global inference for sentence compression: An integer linear programming approach . Journal of Artificial Intelligence Research, 31: 0 399--429, 2008

  29. [37]

    Clarke, D

    J. Clarke, D. Goldwasser, M.-W. Chang, and D. Roth. Driving semantic parsing from the world's response. In Proc. of the Conference on Computational Natural Language Learning (CoNLL), 7 2010. URL http://cogcomp.org/papers/CGCR10.pdf

  30. [38]

    Cocos, V

    A. Cocos, V. Wharton, E. Pavlick, M. Apidianaki, and C. Callison-Burch. Learning Scalar Adjective Intensity from Paraphrases . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1752--1762, 2018

  31. [39]

    T. H. Cormen, C. E. Leiserson, R. L. Rivest, and C. Stein. Introduction to algorithms . MIT press, 2009

  32. [40]

    Dagan, D

    I. Dagan, D. Roth, M. Sammons, and F. M. Zanzoto. Recognizing textual entailment: Models and applications. 7 2013

  33. [41]

    Dalvi, S

    B. Dalvi, S. Bhakthavatsalam, and P. Clark. IKE - An Interactive Tool for Knowledge Extraction . In 5thAKBC Workshop, 2016

  34. [42]

    H. T. Dang and M. Palmer. The role of semantic roles in disambiguating verb senses . In Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics, pages 42--49. Association for Computational Linguistics, 2005

  35. [43]

    H. A. Davidson. Alfarabi, Avicenna, and Averroes on intellect: their cosmologies, theories of the active intellect, and theories of human intellect . Oxford University Press, 1992

  36. [44]

    E. Davis. The Limitations of Standardized Science Tests as Benchmarks for Artificial Intelligence Research: Position Paper . CoRR, abs/1411.1629, 2014. URL http://arxiv.org/abs/1411.1629

  37. [45]

    R. Dechter. Reasoning with Probabilistic and Deterministic Graphical Models: Exact Algorithms . In Reasoning with Probabilistic and Deterministic Graphical Models: Exact Algorithms, 2013

  38. [46]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding . arXiv preprint arXiv:1810.04805, 2018

  39. [47]

    X. Ding, T. Jiang, et al. Spectral distributions of adjacency and Laplacian matrices of random graphs . The annals of applied probability, 20 0 (6): 0 2086--2117, 2010

  40. [48]

    A. N. P. DivyeKhilnani and S. B. D. Jurafsky. Using Query Patterns to Learn the Duration of Events . Computational Semantics IWCS 2011, page 145, 2011

  41. [49]

    Q. Do, Y. S. Chan, and D. Roth. Minimally supervised event causality identification. In Proc. of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Edinburgh, Scotland, 7 2011. URL http://cogcomp.org/papers/DoChaRo11.pdf

  42. [50]

    Q. Do, W. Lu, and D. Roth. Joint inference for event timeline construction. In Proc. of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2012. URL http://cogcomp.org/papers/DoLuRo12.pdf

  43. [51]

    L. Dong, F. Wei, M. Zhou, and K. Xu. Question Answering over Freebase with Multi-Column Convolutional Neural Networks . In Proc. of the Annual Meeting of the Association of Computational Linguistics (ACL), 2015

  44. [52]

    Erdos and A

    P. Erdos and A. R \'e nyi. On the evolution of random graphs . Publ. Math. Inst. Hung. Acad. Sci, 5 0 (1): 0 17--60, 1960

  45. [53]

    Etzioni, M

    O. Etzioni, M. Banko, S. Soderland, and D. Weld. Open information extraction from the web . Communications of the ACM, 51 0 (12): 0 68--74, 2008

  46. [54]

    J. S. B. Evans, S. E. Newstead, and R. M. Byrne. Human reasoning: The psychology of deduction . Psychology Press, 1993

  47. [55]

    Fader, L

    A. Fader, L. Zettlemoyer, and O. Etzioni. Open question answering over curated and extracted knowledge bases . In Proceedings of SIGKDD, pages 1156--1165, 2014

  48. [56]

    Ferrucci, E

    D. Ferrucci, E. Brown, J. Chu-Carroll, J. Fan, D. Gondek, A. A. Kalyanpur, A. Lally, J. W. Murdock, E. Nyberg, J. Prager, et al. Building W atson: An overview of the DeepQA project . AI Magazine, 31 0 (3): 0 59--79, 2010

  49. [57]

    Fikes and T

    R. Fikes and T. Kehler. The role of frame-based representation in reasoning . Communications of the ACM, 28 0 (9): 0 904--920, 1985

  50. [58]

    C. J. Fillmore. Scenes-and-frames semantics . Linguistic structures processing, 59: 0 55--88, 1977

  51. [59]

    J. L. Fleiss. Measuring nominal scale agreement among many raters. Psychological bulletin, 76 0 (5): 0 378, 1971

  52. [60]

    Forbes and Y

    M. Forbes and Y. Choi. Verb Physics: Relative Physical Knowledge of Actions and Objects . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), volume 1, pages 266--276, 2017

  53. [61]

    R. M. French. The Turing Test: the first 50 years . Trends in cognitive sciences, 4 0 (3): 0 115--122, 2000

  54. [62]

    Fried, P

    D. Fried, P. Jansen, G. Hahn-Powell, M. Surdeanu, and P. Clark. Higher-order lexical semantic models for non-factoid answer reranking . Transactions of the Association for Computational Linguistics, 3: 0 197--210, 2015

  55. [63]

    Funahashi

    K.-I. Funahashi. On the approximate realization of continuous mappings by neural networks . Neural networks, 2 0 (3): 0 183--192, 1989

  56. [64]

    Gabrilovich and S

    E. Gabrilovich and S. Markovitch. Computing semantic relatedness using wikipedia-based explicit semantic analysis. In IJcAI, volume 7, pages 1606--1611, 2007

  57. [65]

    Gardner, P

    M. Gardner, P. Talukdar, and T. Mitchell. Combining vector space embeddings with symbolic logical inference over open-domain text . In AAAI spring symposium, 2015

  58. [66]

    Gardner, J

    M. Gardner, J. Grus, M. Neumann, O. Tafjord, P. Dasigi, N. Liu, M. Peters, M. Schmitz, and L. Zettlemoyer. AllenNLP: A Deep Semantic Natural Language Processing Platform . 2018

  59. [67]

    E. N. Gilbert. Random graphs . The Annals of Mathematical Statistics, 30 0 (4): 0 1141--1144, 1959

  60. [68]

    Gildea and D

    D. Gildea and D. Jurafsky. Automatic labeling of semantic roles . Computational linguistics, 28 0 (3): 0 245--288, 2002

  61. [69]

    Goldwasser and D

    D. Goldwasser and D. Roth. Learning from natural instructions. Machine Learning, 94 0 (2): 0 205--232, 2 2014. URL http://cogcomp.org/papers/GoldwasserRo14.pdf

  62. [70]

    Granroth-Wilding and S

    M. Granroth-Wilding and S. Clark. What happens next? event prediction using a compositional neural network model . In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, pages 2727--2733. AAAI Press, 2016

  63. [71]

    Gunning, V

    D. Gunning, V. Chaudhri, P. Clark, K. Barker, J. Chaw, and M. Greaves. P roject H alo Update - Progress Toward Digital A ristotle . AI Magazine, 31 0 (3), 2010

  64. [72]

    Gururangan, S

    S. Gururangan, S. Swayamdipta, O. Levy, R. Schwartz, S. Bowman, and N. A. Smith. Annotation Artifacts in Natural Language Inference Data . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techn...

  65. [73]

    S. Harnad. The symbol grounding problem . Physica D: Nonlinear Phenomena, 42 0 (1-3): 0 335--346, 1990

  66. [74]

    S. Harnad. The Turing Test is not a trick: Turing indistinguishability is a scientific criterion . ACM SIGART Bulletin, 3 0 (4): 0 9--10, 1992

  67. [75]

    K. M. Hermann, T. Kocisk \' y , E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom. Teaching Machines to Read and Comprehend . In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, pages 1693--1...

  68. [76]

    Hernandez-Orallo

    J. Hernandez-Orallo. Beyond the Turing test . Journal of Logic, Language and Information, 9 0 (4): 0 447--466, 2000

  69. [77]

    Hirschman, M

    L. Hirschman, M. Light, E. Breck, and J. D. Burger. Deep Read: A Reading Comprehension System . In 27th Annual Meeting of the Association for Computational Linguistics, ACL 1999 , 1999. URL http://www.aclweb.org/anthology/P99-1042

  70. [78]

    J. R. Hobbs, M. E. Stickel, P. A. Martin, and D. Edwards. Interpretation as Abduction . Artif. Intell., 63: 0 69--142, 1988

  71. [79]

    J. R. Hobbs, M. E. Stickel, D. E. Appelt, and P. Martin. Interpretation as abduction . Artificial intelligence, 63 0 (1-2): 0 69--142, 1993

  72. [80]

    J. H. Holland, K. J. Holyoak, R. E. Nisbett, and P. R. Thagard. Induction: Processes of inference, learning, and discovery . MIT press, 1989

  73. [81]

    Hori and F

    C. Hori and F. Sadaoki. Speech summarization: an approach through word extraction and a method for evaluation . IEICE TRANSACTIONS on Information and Systems, 87 0 (1): 0 15--25, 2004

  74. [82]

    M. J. Hosseini, H. Hajishirzi, O. Etzioni, and N. Kushman. Learning to Solve Arithmetic Word Problems with Verb Categorization . In 2014EMNLP, pages 523--533, 2014

  75. [83]

    D. Howell. Statistical methods for psychology . Cengage Learning, 2012

  76. [84]

    M. Hu, Y. Peng, Z. Huang, X. Qiu, F. Wei, and M. Zhou. Reinforced mnemonic reader for machine reading comprehension. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, pages 4099--4106. AAAI Press, 2018

  77. [85]

    Ide and K

    N. Ide and K. Suderman. Integrating Linguistic Resources: The American National Corpus Model . In Proceedings of the Fifth International Conference on Language Resources and Evaluation, LREC 2006 , pages 621--624, 2006. URL http://www.lrec-conf.org/proceedings/lrec2006/pdf/560_pdf.pdf

  78. [86]

    N. Ide, C. F. Baker, C. Fellbaum, C. J. Fillmore, and R. J. Passonneau. MASC: the Manually Annotated Sub-Corpus of American English . In Proceedings of the International Conference on Language Resources and Evaluation, LREC 2008 , 2008. URL http://www.lrec-conf.org/proceedings...

  79. [87]

    Jansen, N

    P. Jansen, N. Balasubramanian, M. Surdeanu, and P. Clark. What's in an Explanation? Characterizing Knowledge and Inference Requirements for Elementary Science Exams . In Proc. the International Conference on Computational Linguistics (COLING), pages 2956--2965, 2016

  80. [88]

    Jansen, R

    P. Jansen, R. Sharp, M. Surdeanu, and P. Clark. Framing QA as Building and Ranking Intersentence Answer Justifications . Computational Linguistics, 2017

  81. [89]

    P. A. Jansen. A Study of Automatically Acquiring Explanatory Inference Patterns from Corpora of Explanations: Lessons from Elementary Science Exams . In AKBC, 2016

  82. [90]

    P. A. Jansen, E. Wainwright, S. Marmorstein, and C. T. Morrison. WorldTree: A Corpus of Explanation Graphs for Elementary Science Questions supporting Multi-Hop Inference . CoRR, abs/1802.03052, 2018

  83. [91]

    M. E. Janzen and K. J. Vicente. Attention allocation within the abstraction hierarchy . In Proceedings of the Human Factors and Ergonomics Society Annual Meeting, volume 41, pages 274--278. SAGE Publications, 1997

  84. [92]

    Jia and P

    P. Jia and P. Liang. Adversarial Examples for Evaluating Reading Comprehension Systems . Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), 2017

  85. [93]

    Joachims

    T. Joachims. Text categorization with support vector machines: Learning with many relevant features . Machine learning: ECML-98, pages 137--142, 1998

  86. [94]

    Johnson and R

    A. Johnson and R. W. Proctor. Attention: Theory and practice . Sage Publications, 2004

  87. [95]

    P. N. Johnson-Laird. Mental models in cognitive science . Cognitive science, 4 0 (1): 0 71--115, 1980

  88. [96]

    Joshi, E

    M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Volume 1: Long Papers , pages 160...

  89. [97]

    Kaisser and B

    M. Kaisser and B. Webber. Question answering based on semantic roles . In Proceedings of the workshop on deep linguistic processing, pages 41--48, 2007

  90. [98]

    R. M. Kaplan, J. Bresnan, et al. Lexical-functional grammar: A formal system for grammatical representation . In The Mental Representation of Grammatical Relations. The MIT Press, 1982

  91. [99]

    R. J. Kate and R. J. Mooney. Probabilistic Abduction using Markov Logic Networks . In In: IJCAI-09 Workshop on Plan, Activity, and Intent Recognition, 2009

  92. [100]

    Kaushik and Z

    D. Kaushik and Z. C. Lipton. How Much Reading Does Reading Comprehension Require? A Critical Investigation of Popular Benchmarks . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 5010--5015, 2018

  93. [101]

    Kembhavi, M

    A. Kembhavi, M. Seo, D. Schwenk, J. Choi, A. Farhadi, and H. Hajishirzi. Are You Smarter Than A Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension . The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017

  94. [102]

    Khashabi, T

    D. Khashabi, T. Khot, A. Sabharwal, P. Clark, O. Etzioni, and D. Roth. Question answering via integer programming over semi-structured knowledge. In Proc. of the International Joint Conference on Artificial Intelligence (IJCAI), 2016. URL http://cogcomp.org/papers/KKSCER16.pdf

  95. [103]

    Khashabi, T

    D. Khashabi, T. Khot, A. Sabharwal, and D. Roth. Learning what is essential in questions. In The Conference on Computational Natural Language Learning (Proc. of the Conference on Computational Natural Language Learning (CoNLL)), 2017. URL http://cogcomp.org/papers/2017_conll_e...

  96. [104]

    Khashabi, S

    D. Khashabi, S. Chaturvedi, M. Roth, S. Upadhyay, and D. Roth. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics ...

  97. [105]

    Khashabi, T

    D. Khashabi, T. Khot, A. Sabharwal, and D. Roth. Question answering as global reasoning over semantic abstractions. In Proceedings of The Conference on Artificial Intelligence (Proc. of the Conference on Artificial Intelligence (AAAI)), 2018 b . URL http://cogcomp.org/papers/2...

  98. [106]

    Khashabi, M

    D. Khashabi, M. Sammons, B. Zhou, T. Redman, C. Christodoulopoulos, V. Srikumar, N. Rizzolo, L. Ratinov, G. Luo, Q. Do, C.-T. Tsai, S. Roy, S. Mayhew, Z. Feng, J. Wieting, X. Yu, Y. Song, S. Gupta, S. Upadhyay, N. Arivazhagan, Q. Ning, S. Ling, and D. Roth. Cogcompnlp: Your sw...

  99. [107]

    Khashabi, E

    D. Khashabi, E. S. Azer, T. Khot, A. Sabharwal, and D. Roth. On the capabilities and limitations of reasoning for natural language understanding, 2019. under review

  100. [108]

    T. Khot, N. Balasubramanian, E. Gribkoff, A. Sabharwal, P. Clark, and O. Etzioni. Exploring M arkov Logic Networks for Question Answering . In 2015EMNLP, Lisbon, Portugal, 2015

  101. [109]

    T. Khot, A. Sabharwal, and P. Clark. Answering Complex Questions Using Open Information Extraction . Proc. of the Annual Meeting of the Association of Computational Linguistics (ACL), 2017

  102. [110]

    Kingsbury and M

    P. Kingsbury and M. Palmer. From TreeBank to PropBank. In LREC, pages 1989--1993, 2002

  103. [111]

    G. S. Kirk, J. E. Raven, and M. Schofield. The presocratic philosophers: A critical history with a selcetion of texts . Cambridge University Press, 1983

  104. [112]

    Knight and D

    K. Knight and D. Marcu. Summarization beyond sentence extraction: A probabilistic approach to sentence compression . Artificial Intelligence, 139 0 (1): 0 91--107, 2002

  105. [113]

    J. Ko, E. Nyberg, and L. Si. A probabilistic graphical model for joint answer ranking in question answering . In Proceedings of SIGIR, pages 343--350, 2007

  106. [114]

    Kozareva and E

    Z. Kozareva and E. Hovy. Learning temporal information for states and events . In Fifth International Conference on Semantic Computing, pages 424--429. IEEE, 2011

  107. [115]

    Krishnamurthy, O

    J. Krishnamurthy, O. Tafjord, and A. Kembhavi. Semantic parsing to probabilistic programs for situated question answering . Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), 2016

  108. [116]

    C. C. T. Kwok, O. Etzioni, and D. S. Weld. Scaling question answering to the Web . In The International World Wide Web Conference, 2001

  109. [117]

    G. Lai, Q. Xie, H. Liu, Y. Yang, and E. H. Hovy. RACE: Large-scale ReAding Comprehension Dataset From Examinations . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017 , pages 785--794, 2017. URL https://aclanthology.info/pape...

  110. [118]

    H. Lee, Y. Peirsman, A. Chang, N. Chambers, M. Surdeanu, and D. Jurafsky. Stanford's multi-pass sieve coreference resolution system at the CoNLL-2011 shared task . In CONLL Shared Task, pages 28--34, 2011

  111. [119]

    H. Lee, A. Chang, Y. Peirsman, N. Chambers, M. Surdeanu, and D. Jurafsky. Deterministic coreference resolution based on entity-centric, precision-ranked rules . Computational Linguistics, 39 0 (4): 0 885--916, 2013

  112. [120]

    K. Lee, Y. Artzi, J. Dodge, and L. Zettlemoyer. Context-dependent semantic parsing for time expressions. In ACL (1), pages 1437--1447, 2014

  113. [121]

    Leeuwenberg and M.-F

    A. Leeuwenberg and M.-F. Moens. Temporal Information Extraction by Predicting Relative Time-lines . Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), 2018

  114. [122]

    W. G. Lehnert. The Process of Question Answering. PhD thesis, Yale University, 1977

  115. [123]

    D. B. Lenat. CYC: A large-scale investment in knowledge infrastructure . Communications of the ACM, 38 0 (11): 0 33--38, 1995

  116. [124]

    Levy and Y

    O. Levy and Y. Goldberg. Linguistic regularities in sparse and explicit word representations . In Proceedings of the eighteenth conference on computational natural language learning, pages 171--180, 2014

  117. [125]

    F. Li, X. Zhang, J. Yuan, and X. Zhu. Classifying What-Type Questions by Head Noun Tagging . In Proceedings 22nd International Conference on Computational Linguistics (COLING), 2007

  118. [126]

    Li and D

    X. Li and D. Roth. Learning Question Classifiers . In Proceedings of the 19th International Conference on Computational Linguistics - Volume 1, COLING '02, pages 1--7, Stroudsburg, PA, USA, 2002. Association for Computational Linguistics

  119. [127]

    Y. Li, L. Xu, F. Tian, L. Jiang, X. Zhong, and E. Chen. Word Embedding Revisited: A New Representation Learning and Explicit Matrix Factorization Perspective. In Proc. of the International Joint Conference on Artificial Intelligence (IJCAI), pages 3650--3656, 2015

  120. [128]

    Z. Li, X. Ding, and T. Liu. Constructing Narrative Event Evolutionary Graph for Script Event Prediction . Proc. of the International Joint Conference on Artificial Intelligence (IJCAI), 2018

  121. [129]

    X. V. Lin, R. Socher, and C. Xiong. Multi-Hop Knowledge Graph Reasoning with Reward Shaping . In Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), 2018

  122. [130]

    Liu and P

    H. Liu and P. Singh. ConceptNet—a practical commonsense reasoning tool-kit . BT technology journal, 22 0 (4): 0 211--226, 2004

  123. [131]

    X. Liu, Y. Shen, K. Duh, and J. Gao. Stochastic answer networks for machine reading comprehension. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1694--1704, 2018

  124. [132]

    A. A. Mahabal, D. Roth, and S. Mittal. Robust handling of polysemy via sparse representations. In *SEM, 2018. URL http://cogcomp.org/papers/MahabalRoMi18.pdf

  125. [133]

    McCallum, A

    A. McCallum, A. Neelakantan, R. Das, and D. Belanger. Chains of Reasoning over Entities, Relations, and Text using Recurrent Neural Networks . In EACL , pages 132--141, 2017

  126. [134]

    McCarthy

    J. McCarthy. Programs with common sense . Defense Technical Information Center, 1963

  127. [135]

    McCarthy

    J. McCarthy. An example for natural language understanding and the AI problems it raises . Formalizing Common Sense: Papers by John McCarthy, 355, 1976

  128. [136]

    McCarthy and M

    J. McCarthy and M. I. Levin. LISP 1.5 programmer's manual . MIT press, 1965

  129. [137]

    McCarthy and V

    J. McCarthy and V. Lifschitz. Formalizing common sense: papers , volume 5. Intellect Books, 1990

  130. [138]

    J. F. McCarthy. Using decision trees for coreference resolution . In Proc. 14th International Joint Conf. on Artificial Intelligence (IJCAI), Quebec, Canada, Aug. 1995, 1995

  131. [139]

    Merkhofer, J

    E. Merkhofer, J. Henderson, D. Bloom, L. Strickhart, and G. Zarrella. MITRE at SemEval-2018 Task 11: Commonsense Reasoning without Commonsense Knowledge . In Proceedings of the International Workshop on Semantic Evaluation (SemEval-2018), New Orleans, LA, USA, 2018

  132. [140]

    Meyers, R

    A. Meyers, R. Reeves, C. Macleod, R. Szekely, V. Zielinska, B. Young, and R. Grishman. The NomBank project: An interim report . In HLT-NAACL 2004 workshop: Frontiers in corpus annotation, volume 24, page 31, 2004

  133. [141]

    Mihalcea and A

    R. Mihalcea and A. Csomai. Wikify!: linking documents to encyclopedic knowledge . In CIKM, pages 233--242, 2007

  134. [142]

    Mikolov, K

    T. Mikolov, K. Chen, G. Corrado, and J. Dean. Efficient estimation of word representations in vector space . arXiv preprint arXiv:1301.3781, 2013

  135. [143]

    S. Milgram. Six degrees of separation . Psychology Today, 2: 0 60--64, 1967

  136. [144]

    G. Miller. WordNet: a lexical database for English . Communications of the ACM, 38 0 (11): 0 39--41, 1995

  137. [145]

    S. Min, M. J. Seo, and H. Hajishirzi. Question Answering through Transfer Learning from Large Fine-grained Supervision Data . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Volume 2: Short Papers , pages 510--517, 2017. UR...

  138. [146]

    M. Minsky. A Framework for Representing Knowledge . Technical report, Massachusetts Institute of Technology, Cambridge, MA, USA, 1974

  139. [147]

    M. Minsky. Society of mind . Simon and Schuster, 1988

  140. [148]

    Minsky and S

    M. Minsky and S. Papert. Perceptron: an introduction to computational geometry . The MIT Press, Cambridge, expanded edition, 19: 0 88, 1969

  141. [149]

    T. M. Mitchell, J. Betteridge, A. Carlson, E. Hruschka, and R. Wang. Populating the semantic web by macro-reading internet text . In International Semantic Web Conference, pages 998--1002. Springer, 2009

  142. [150]

    Moldovan, M

    D. Moldovan, M. Pa s ca, S. Harabagiu, and M. Surdeanu. Performance issues and error analysis in an open-domain question answering system . ACM Transactions on Information Systems (TOIS), 21 0 (2): 0 133--154, 2003

  143. [151]

    Moreda, H

    P. Moreda, H. Llorens, E. S. Bor \'o , and M. Palomar. Combining semantic information in question answering systems . Inf. Process. Manage., 47: 0 870--885, 2011

  144. [152]

    Narasimhan and R

    K. Narasimhan and R. Barzilay. Machine Comprehension with Discourse Relations . In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natur...

  145. [153]

    B. K. Natarajan. Sparse approximate solutions to linear systems . SIAM journal on computing, 24 0 (2): 0 227--234, 1995

  146. [154]

    Nguyen, M

    T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, and L. Deng. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset . CoRR, abs/1611.09268, 2016. URL http://arxiv.org/abs/1611.09268

  147. [155]

    J. Ni, C. Zhu, W. Chen, and J. McAuley. Learning to attend on essential terms: An enhanced retriever-reader model for scientific question answering . arXiv preprint arXiv:1808.09492, 2018

  148. [156]

    Q. Ning, H. Wu, H. Peng, and D. Roth. Improving Temporal Relation Extraction with a Globally Acquired Statistical Resource . In Proc. of the Annual Meeting of the North American Association of Computational Linguistics (NAACL), pages 841--851, New Orleans, Louisiana, 6 2018 a ...

  149. [157]

    Q. Ning, B. Zhou, Z. Feng, H. Peng, and D. Roth. CogCompTime: A Tool for Understanding Time in Natural Language . In EMNLP (Demo Track), Brussels, Belgium, 11 2018 b . Association for Computational Linguistics. URL http://cogcomp.org/papers/NZFPR18.pdf

  150. [158]

    G. Novak. Representations of Knowledge in a Program for Solving Physics Problems . In IJCAI-77, 1977

  151. [159]

    Ostermann, M

    S. Ostermann, M. Roth, A. Modi, S. Thater, and M. Pinkal. SemEval-2018 Task 11: Machine Comprehension using Commonsense Knowledge . In Proceedings of The 12th International Workshop on Semantic Evaluation, pages 747--757, 2018

  152. [160]

    Palmer, D

    M. Palmer, D. Gildea, and P. Kingsbury. The proposition bank: An annotated corpus of semantic roles . Computational linguistics, 31 0 (1): 0 71--106, 2005

  153. [161]

    a ckstr\

    A. P. Parikh, O. T \"a ckstr\" o m, D. Das, and J. Uszkoreit. A Decomposable Attention Model for Natural Language Inference . In Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), 2016

  154. [162]

    J. H. Park and W. B. Croft. Using key concepts in a translation model for retrieval . In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 927--930. ACM, 2015

  155. [163]

    J. Pearl. Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 1988. ISBN 1558604790

  156. [164]

    C. S. Peirce. A Theory of Probable Inference . In Studies in Logic by Members of the Johns Hopkins University , pages 126--181. Little, Brown, and Company, 1883

  157. [165]

    Pennington, R

    J. Pennington, R. Socher, and C. Manning. Glove: Global vectors for word representation . In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532--1543, 2014

  158. [166]

    Peters, M

    M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer. Deep Contextualized Word Representations . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volu...

  159. [167]

    L. A. Pizzato and D. Moll \'a . Indexing on semantic roles for question answering . In 2nd workshop on Information Retrieval for Question Answering, pages 74--81, 2008

  160. [168]

    Poliak, J

    A. Poliak, J. Naradowsky, A. Haldar, R. Rudinger, and B. V. Durme. Hypothesis Only Baselines in Natural Language Inference . In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pages 180--191, 2018

  161. [169]

    D. Poole. A methodology for using a default and abductive reasoning system . Int. J. Intell. Syst., 5: 0 521--548, 1990

  162. [170]

    Punyakanok and D

    V. Punyakanok and D. Roth. The use of classifiers in sequential inference. In Proc. of the Conference on Neural Information Processing Systems (NIPS), pages 995--1001. MIT Press, 2001. URL http://cogcomp.org/papers/nips01.pdf

  163. [171]

    Punyakanok, D

    V. Punyakanok, D. Roth, and W. Yih. Mapping Dependencies Trees: An Application to Question Answering . AIM, 1 2004. URL http://cogcomp.org/papers/PunyakanokRoYi04a.pdf

  164. [172]

    Punyakanok, D

    V. Punyakanok, D. Roth, and W. tau Yih. The importance of syntactic parsing and inference in semantic role labeling. Computational Linguistics, 2008

  165. [173]

    M. R. Quillan. Semantic memory . Technical report, BOLT BERANEK AND NEWMAN INC CAMBRIDGE MA, 1966

  166. [174]

    Rajpurkar, J

    P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. SQuAD : 100,000+ Questions for Machine Comprehension of Text . In Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), 2016

  167. [175]

    Rajpurkar, R

    P. Rajpurkar, R. Jia, and P. Liang. Know What You Don't Know: Unanswerable Questions for SQuAD . In Proc. of the Annual Meeting of the Association of Computational Linguistics (ACL), 2018

  168. [176]

    Rashkin, M

    H. Rashkin, M. Sap, E. Allaway, N. A. Smith, and Y. Choi. Event2Mind: Commonsense Inference on Events, Intents, and Reactions . In Proc. of the Annual Meeting of the Association of Computational Linguistics (ACL), pages 463--473, 2018

  169. [177]

    Rasmussen

    J. Rasmussen. The role of hierarchical knowledge representation in decisionmaking and system management . Systems, Man and Cybernetics, IEEE Transactions on, pages 234--243, 1985

  170. [178]

    Ratinov and D

    L. Ratinov and D. Roth. Design challenges and misconceptions in named entity recognition. In Proc. of the Conference on Computational Natural Language Learning (CoNLL), 6 2009. URL http://cogcomp.org/papers/RatinovRo09.pdf

  171. [179]

    Ratinov, D

    L. Ratinov, D. Roth, D. Downey, and M. Anderson. Local and global algorithms for disambiguation to wikipedia. In Proc. of the Annual Meeting of the Association for Computational Linguistics (ACL), 2011. URL http://cogcomp.org/papers/RRDA11.pdf

  172. [180]

    a ckstr \

    S. Reddy, O. T \"a ckstr \"o m, S. Petrov, M. Steedman, and M. Lapata. Universal Semantic Parsing . In Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), pages 89--101, 2017

  173. [181]

    Redman, M

    T. Redman, M. Sammons, and D. Roth. Illinois Named Entity Recognizer: Addendum to R atinov and R oth '09 reporting improved results , 2016. URL http://cogcomp.org/papers/ner-addendum.pdf. Tech Report

  174. [182]

    Richardson and P

    M. Richardson and P. Domingos. M arkov Logic Networks . Machine learning, 62 0 (1--2): 0 107--136, 2006

  175. [183]

    Richardson, C

    M. Richardson, C. J. C. Burges, and E. Renshaw. MCTest: A Challenge Dataset for the Open-Domain Machine Comprehension of Text . In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013 , pages 193--203, 2013. URL http://aclweb.org/a...

  176. [184]

    Rosenblatt

    F. Rosenblatt. The perceptron: a probabilistic model for information storage and organization in the brain. Psychological review, 65 0 (6): 0 386, 1958

  177. [185]

    Roth and W

    D. Roth and W. Yih. A linear programming formulation for global inference in natural language tasks. In H. T. Ng and E. Riloff, editors, Proc. of the Conference on Computational Natural Language Learning (CoNLL), pages 1--8. Association for Computational Linguistics, 2004. URL...

  178. [186]

    Roth and D

    D. Roth and D. Zelenko. Part of speech tagging using a network of linear separators. In ACL-COLING, 1998

  179. [187]

    Roth and M

    M. Roth and M. Lapata. Neural semantic role labeling with dependency path embeddings . Proc. of the Annual Meeting of the Association of Computational Linguistics (ACL), 2016

  180. [188]

    S. Roy, T. Vieira, and D. Roth. Reasoning about quantities in natural language. Transactions of the Association for Computational Linguistics (TACL), 3, 2015. URL http://cogcomp.org/papers/RoyViRo15.pdf

  181. [189]

    D. E. Rumelhart, G. E. Hinton, and R. J. Williams. Learning representations by back-propagating errors . Cognitive modeling, 5, 1988

  182. [190]

    R. C. Schank. Conceptual dependency: A theory of natural language understanding . Cognitive psychology, 3 0 (4): 0 552--631, 1972

  183. [191]

    R. C. Schank and R. P. Abelson. Scripts, plans, and knowledge . In Proc. of the International Joint Conference on Artificial Intelligence (IJCAI), pages 151--157, 1975

  184. [192]

    Selman and H

    B. Selman and H. J. Levesque. Abductive and Default Reasoning: A Computational Core . In Proceedings of the National Conference on Artificial Intelligence (AAAI), 1990

  185. [193]

    M. Seo, A. Kembhavi, A. Farhadi, and H. Hajishirzi. Bidirectional attention flow for machine comprehension . ICLR, 2016

  186. [194]

    Shen and M

    D. Shen and M. Lapata. Using Semantic Roles to Improve Question Answering. In EMNLP-CoNLL, pages 12--21, 2007

  187. [195]

    Socher, D

    R. Socher, D. Chen, C. D. Manning, and A. Y. Ng. Reasoning With Neural Tensor Networks for Knowledge Base Completion . In The Conference on Advances in Neural Information Processing Systems (NIPS), 2013

  188. [196]

    Srikumar and D

    V. Srikumar and D. Roth. A Joint Model for Extended Semantic Role Labeling . In Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), Edinburgh, Scotland, 2011. URL http://cogcomp.org/papers/SrikumarRo11.pdf

  189. [197]

    Srikumar and D

    V. Srikumar and D. Roth. Modeling semantic relations expressed by prepositions. 1: 0 231--242, 2013. URL http://cogcomp.org/papers/SrikumarRo13.pdf

  190. [198]

    Steedman and J

    M. Steedman and J. Baldridge. Combinatory categorial grammar . Non-Transformational Syntax: Formal and explicit models of grammar, pages 181--224, 2011

  191. [199]

    Stern, R

    A. Stern, R. Stern, I. Dagan, and A. Felner. Efficient search for transformation-based inference . In Proc. of the Annual Meeting of the Association of Computational Linguistics (ACL), pages 283--291, 2012

  192. [200]

    M. Steup. Epistemology . In E. N. Zalta, editor, The Stanford Encyclopedia of Philosophy. http://plato.stanford.edu/archives/spr2014/entries/epistemology/, spring 2014 edition, 2014

  193. [201]

    K. Sun, D. Yu, D. Yu, and C. Cardie. Improving machine reading comprehension with general reading strategies . In Proc. of the Annual Meeting of the North American Association of Computational Linguistics (NAACL), 2019

  194. [202]

    W. t. Yih, X. He, and C. Meek. Semantic Parsing for Single-Relation Question Answering. In 52ndACL, pages 643--648. Citeseer, 2014

  195. [203]

    Taddeo and L

    M. Taddeo and L. Floridi. Solving the symbol grounding problem: a critical review of fifteen years of research . Journal of Experimental & Theoretical Artificial Intelligence, 17 0 (4): 0 419--445, 2005

  196. [204]

    P. P. Talukdar, M. Jacob, M. S. Mehmood, K. Crammer, Z. G. Ives, F. Pereira, and S. Guha. Learning to create data-integrating queries . Proceedings of the VLDB Endowment, 1 0 (1): 0 785--796, 2008

  197. [205]

    P. P. Talukdar, Z. G. Ives, and F. Pereira. Automatically incorporating new sources in keyword search-based data integration . In Proceedings of the 2010 ACM SIGMOD International Conference on Management of data, pages 387--398. ACM, 2010

  198. [206]

    Tandon, B

    N. Tandon, B. Dalvi, J. Grus, W. tau Yih, A. Bosselut, and P. Clark. Reasoning about Actions and State Changes by Injecting Commonsense Knowledge . In Proc. of the Conference on Empirical Methods for Natural Language Processing (EMNLP), pages 57--66, 2018

  199. [207]

    Toutanova and D

    K. Toutanova and D. Chen. Observed versus latent features for knowledge base and text inference . In CVSC workshop, 2015

  200. [208]

    Trivedi, H

    H. Trivedi, H. Kwon, T. Khot, A. Sabharwal, and N. Balasubramanian. Entailment-based Question Answering over Multiple Sentences . In Proc. of the Annual Meeting of the North American Association of Computational Linguistics (NAACL), 2019

  201. [209]

    A. M. Turing. Computing machinery and intelligence . Mind, 59 0 (236): 0 433, 1950

  202. [210]

    P. D. Turney. Distributional semantics beyond words: Supervised learning of analogy and paraphrase . TACL, 1: 0 353--366, 2013

  203. [211]

    P. D. Turney and P. Pantel. From frequency to meaning: Vector space models of semantics . Journal of artificial intelligence research, 37: 0 141--188, 2010

  204. [212]

    Tymoshenko, D

    K. Tymoshenko, D. Bonadiman, and A. Moschitti. Convolutional Neural Networks vs. Convolution Kernels: Feature Engineering for Answer Sentence Reranking . In HLT-NAACL, 2016

  205. [213]

    Unger, L

    C. Unger, L. B \"u hmann, J. Lehmann, A.-C. N. Ngomo, D. Gerber, and P. Cimiano. Template-based question answering over RDF data . In Proceedings of the 21st international conference on World Wide Web, pages 639--648. ACM, 2012

  206. [214]

    Vempala, E

    A. Vempala, E. Blanco, and A. Palmer. Determining Event Durations: Models and Error Analysis . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), volume 2, ...

  207. [215]

    B. Wang, K. Liu, and J. Zhao. Inner Attention based Recurrent Neural Networks for Answer Selection . In Proc. of the Annual Meeting of the Association of Computational Linguistics (ACL), 2016

  208. [216]

    C. Wang, N. Xue, S. Pradhan, and S. Pradhan. A Transition-based Algorithm for AMR Parsing. In HLT-NAACL, pages 366--375, 2015

  209. [217]

    H. Wang, D. Yu, K. Sun, J. Chen, D. Yu, D. Roth, and D. McAllester. Evidence Sentence Extraction for Machine Reading Comprehension . arXiv preprint arXiv:1902.08852, 2019

  210. [218]

    W. Wang, M. Yan, and C. Wu. Multi-granularity hierarchical attention fusion networks for reading comprehension and question answering. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1705--1714, 2018

  211. [219]

    D. J. Watts and S. H. Strogatz. Collective dynamics of ‘small-world’networks . nature, 393 0 (6684): 0 440, 1998

  212. [220]

    Wieting, M

    J. Wieting, M. Bansal, K. Gimpel, K. Livescu, and D. Roth. From Paraphrase Database to Compositional Paraphrase Model and Back . TACL, 3: 0 345--358, 2015

  213. [221]

    Williams

    J. Williams. Extracting fine-grained durations for verbs from Twitter . In Proceedings of ACL 2012 Student Research Workshop, pages 49--54. Association for Computational Linguistics, 2012

  214. [222]

    Winograd

    T. Winograd. Understanding natural language . Cognitive psychology, 3 0 (1): 0 1--191, 1972

  215. [223]

    W. A. Woods. Progress in natural language understanding: an application to lunar geology . In Proceedings of the June 4-8, 1973, national computer conference and exposition, pages 441--450. ACM, 1973

  216. [224]

    S. Yang, L. Zou, Z. Wang, J. Yan, and J.-R. Wen. Efficiently Answering Technical Questions-A Knowledge Graph Approach. In Proceedings of the National Conference on Artificial Intelligence (AAAI), pages 3111--3118, 2017

  217. [225]

    Y. Yang, W. Yih, and C. Meek. WikiQA: A Challenge Dataset for Open-Domain Question Answering . In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015 , pages 2013--2018, 2015. URL http://aclweb.org/anthology/D/D15/D15-1237.pdf

  218. [226]

    Y. Yang, L. Birnbaum, J.-P. Wang, and D. Downey. Extracting Commonsense Properties from Embeddings with Limited Human Guidance . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), volume 2, pages 644--649, 2018

  219. [227]

    Yao and B

    X. Yao and B. V. Durme. Information extraction over structured data: Question answering with F reebase . In 52ndACL, 2014

  220. [228]

    W. Yin, S. Ebert, and H. Sch \"u tze. Attention-based convolutional neural network for machine comprehension . In NAACL HCQA Workshop, 2016

  221. [229]

    L. A. Zadeh. The concept of a linguistic variable and its application to approximate reasoning—I . Information sciences, 8 0 (3): 0 199--249, 1975

  222. [230]

    L. A. Zadeh. PRUF—a meaning representation language for natural languages . International Journal of man-machine studies, 10 0 (4): 0 395--460, 1978

  223. [231]

    Zellers, Y

    R. Zellers, Y. Bisk, R. Schwartz, and Y. Choi. SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 93--104, 2018

  224. [232]

    L. S. Zettlemoyer and M. Collins. Learning to map sentences to logical form: Structured classification with probabilistic categorial grammars . UAI, 2005

  225. [233]

    Zhang, R

    S. Zhang, R. Rudinger, K. Duh, and B. V. Durme. Ordinal Common-sense Inference . Transactions of the Association of Computational Linguistics, 5 0 (1): 0 379--395, 2017

  226. [234]

    B. Zhou, D. Khashabi, Q. Ning, and D. Roth. ``going on a vacation'' takes longer than ``going for a walk'': A study of temporal commonsense understanding. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), 2019

  227. [235]

    L. Zou, R. Huang, H. Wang, J. X. Yu, W. He, and D. Zhao. Natural language question answering over RDF : a graph data driven approach . In SIGMOD, pages 313--324, 2014

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.