REVIEW 3 major objections 3 minor 200 references
Teaching Smaller Language Models To Generalise To Unseen Compositional Questions (Full Thesis)
T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A 440M-parameter BART model, trained to reason over retrieved contexts and LLM-generated rationales, can generalise to unseen compositional questions rather than answer from memorised training samples, and can match or exceed much larger…
desk verdict A thorough, candid thesis on training small models for unseen compositional QA; the anti-memorisation claim needs a pinch of salt, but the retrieval-plus-rationale results are solid and worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the context-acquisition and context-scoring pipeline. The Iterator extends multi-hop dense retrieval to an arbitrary number of hops (four here): at each hop it retrieves candidate paragraphs, reranks paragraphs and sentences, and uses an Evidence Set Scorer to decide whether the accumulated sentences are sufficient to answer the question, finally returning a title-prefixed paragraph-fragment context. The Rationale Ranking (RR) model is a smaller Transformer trained, with shared normalisation, to give one score to a question-context pair that reflects both relevance and truthfulness, allowing it to select and filter components from two knowledge sources (LLM-generated rationales and retrieved Wikipedia paragraphs). The Reasoning Model is a multitask-trained BART model; the RATD datasets give it practice weighing partially evidential facts in long noisy contexts of the same form it will see at test time.
What would settle it
One could audit the 'Unmemorisable' subsets by retrieving, for each sample, the nearest training neighbours below the threshold and checking whether the model's correct prediction coincides with a neighbour's gold answer; finding even one such pair would show the filter is incomplete and the generalisation gain is partially memorisation.
Extended reading notes
Core claim
The central claim is that a smaller Language Model is capable of performance beyond simple memorisation in deriving correct answers to challenging compositional questions. To support this, the thesis introduces a method that scores each evaluation question together with its answer against every training sample using sentence-embedding cosine similarity, keeps only samples below a similarity threshold with no answer-term overlap, and then measures the effect of adding new training data on that 'unmemorisable' subset; improvements there (for DROP and ROPES) cannot be attributed to memorisation. The thesis further claims that the Iterator, an $n$-hop dense retrieval, reranking and evidence-set scoring system, can supply useful contexts for arbitrary unseen questions from Wikipedia, and that training on RATD datasets built from those noisy contexts teaches the 440M BART model to reason toward plausible answers from partial evidence. Adding a Rationale Ranking model that scores contexts for truthfulness as well as relevance lets the small reasoner exploit combined LLM-rationale and retrieved-paragraph contexts, with the combined-context model matching or exceeding much larger models when both receive the same knowledge.
Load-bearing premise
The anti-memorisation conclusion rests on the assumption that the similarity threshold of 60 plus the no-answer-term-overlap filter really leaves only evaluation samples that cannot be answered from memory; if a memorisable paraphrase or a discontinuous token overlap falls below that threshold, the reported gains on 'unmemorisable' subsets could be inflated by memorisation.
Editorial extensions
If this is right
- A single workstation with one consumer GPU could run a question-answering system that handles unseen multi-hop and commonsense questions at levels previously associated with models orders of magnitude larger.
- Retrieval-augmented training that includes partial, irrelevant, or missing evidence teaches the reasoner to use noisy combined contexts, so it can exploit a naive concatenation of rationale and retrieved paragraphs without special scoring.
- Scores that combine relevance with truthfulness can filter false LLM rationales and improve mean performance, meaning small models can act as practical truthfulness filters in constrained settings.
- The two knowledge sources are complementary: LLM rationales help commonsense reasoning, while multi-hop retrieval helps questions that need facts from two or more documents, so combining them raises performance beyond either source alone.
- A smaller model given the same knowledge that a LLM generates for itself can beat the LLM's own direct answer, suggesting the bottleneck for small models is knowledge access rather than reasoning capacity.
Reading between the lines
- Editorial inference: the embedding-based memorisability audit could be reused as a general contamination check for QA benchmarks, since scoring both question and answer catches paraphrase and discontinuous overlap that n-gram filters miss.
- Editorial inference: the Rationale Ranking approach points toward a cheap, general route to truthfulness filtering — a small model trained on positive/negative pairs from diverse datasets may substitute for much larger factuality detectors in resource-limited pipelines.
- Editorial inference: the recipe 'give a small reasoner noisy multi-source context and train it on similar noise' could transfer to other tasks such as fact verification, summarisation, or open-domain reading comprehension, reducing the need for specialised retrieval architectures.
- Editorial inference: since the best global RR combination defaulted to naive concatenation for most samples, the untried step the thesis leaves implicit is a per-question selector that predicts which combination strategy will work, rather than applying one strategy globally.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This PhD thesis studies whether a 440M-parameter BART model, trained with multitask question-answering datasets, can answer compositional questions that are unseen during training when provided with retrieved or generated context. Chapter 4 proposes a semantic-similarity method to partition evaluation samples into memorisable and unmemorisable subsets, and an intervention experiment (UQA vs UQA+TDND) that claims the model generalises beyond memorisation on DROP and ROPES. Chapter 5 introduces the Iterator multi-hop dense retrieval system and retrieval-augmented training datasets (RATD), reporting significant improvements over baselines on several unseen datasets. Chapter 6 adds LLM-generated rationales as a second knowledge source and proposes Rationale Ranking (RR) to score and combine contexts; the thesis claims significant gains from combining sources and from the RR method, and shows that the small model can outperform direct LLM prompting given the same context.
Significance. If the findings hold, the practical contribution is substantial: it would demonstrate that a locally deployable 440M model can approach or exceed much larger models on unseen multi-hop and commonsense QA when given the same contextual information. The thesis is strong in empirical breadth: multiple unseen datasets, paired bootstrap and Nemenyi significance tests, ablations for training-regime components, and public code/data links. The external evaluation of the RR model on TruthfulQA (Table 6.3) is a valuable sanity check that the model learns truthfulness, not just relevance. However, two methodological issues (the calibration of the memorisation filter in Chapter 4 and the selection of the RR threshold on the evaluation data in Chapter 6) currently limit the strength of the two headline contributions.
major comments (3)
- [§4.2.3, Table 4.2, Table 4.5] The anti-memorisation claim (Contribution 1, Section 1.3) depends on the 'Unmemorisable' subsets being genuinely unanswerable from memory. The subset is defined by a hand-calibrated similarity threshold T=60 and by removing only samples whose answer has no word overlap with the single most similar training sample. The thesis explicitly acknowledges in Section 4.2.3 that it cannot eliminate evaluation samples whose answer overlaps with any training sample. As a result, a DROP or ROPES evaluation sample whose answer appears in a second-most-similar training sample remains in the subset, and the model could have memorised that sample. In addition, the stsb-roberta-large embeddings used in Eq. (4.1) may miss paraphrased or discontinuous overlaps, so a memorisable pair could score below T=60. If even a modest fraction of the 652 DROP or 197 ROPES 'Unmemorisable' samples are memorisable, the 9.0% and 25.7% gains in Table 4.5 would be inflated, and the conclusion in Section 4.4 ('the improvement is not attributable to memorisation') would be unsupported. To make the claim load-bearing, the filter should be applied against all training samples (e.g. via a nearest-neighbour scan), and the sensitivity of the gains to T should be reported.
- [§6.3.2, Table 6.6] The 'Generally best RR combo' is selected per Reasoning Model as the combination method with the highest unweighted macro-average over the unseen evaluation datasets (Section 6.3.2, text immediately before Table 6.5). These same datasets are then used to compute the +RR -RATD versus -RR -RATD difference in Table 6.6 (45.5 vs 42.7) and to run the Nemenyi significance tests. This is selection on the evaluation data: the threshold t has been chosen to maximise the very quantity being tested, so the reported improvement is an optimistically biased estimate of performance on arbitrary unseen questions. While the thesis correctly states that this method is not usable for a truly unseen question, it nevertheless uses this result as evidence for Contribution 5 ('the RR method significantly outperforms...'). The authors should either (a) choose t on a held-out validation fold and report the resulting test performance, or (b) restrict the RR-value claim to threshold-free evidence such as the Naïve Concatenation result, which also shows a significant combined-context benefit (47.2 vs 42.7).
- [§5.4, Table 5.9] The Chapter 5 conclusion states: 'training on RATD datasets improves performance on all unseen evaluation datasets with retrieved contexts'. This is internally contradicted by Table 5.9, where MusiqueR results significantly degrade from 24.3 F1 (Base) to 22.2 F1 (Base+RATD), while ARC-DA improves. The post-hoc ablative model (excluding Musique RATD, or using only unique-label Musique samples) recovers and improves performance, which is informative, but it does not make the original claim true. The conclusion should be amended to state that RATD improves performance on all retrieved-context datasets except Musique, where the training-distribution bias causes degradation.
minor comments (3)
- [§2.3] 'accessable' should be 'accessible' in the sentence 'All versions of our evaluation (and training) datasets are accessable at github.com/timhartill/unseen_questions.'
- [§6.1] The method is called 'Rational Ranking' in the first bullet point of the introduction but 'Rationale Ranking' everywhere else; the terminology should be consistent.
- [§5.3.2.1, Table 5.3] The Base model's SQAR score of 48.4 (below random) is accompanied by a footnote that prepending 'Yes or no -' improves it to 54.9. This prompt adjustment should be presented as the primary result for transparency, since reporting the below-random score is likely to be misinterpreted.
Circularity Check
Both headline gains are partly self-referential: the RR combination threshold is selected on the target evaluation sets, and the Chapter 4 'Unmemorisable' subset is defined by a hand-calibrated threshold on those same sets; external benchmarks provide independent but partial support.
-
fitted input called prediction
[Section 6.3.2 (Context Combination Methods and Experimental Nomenclature); Tables 6.4 and 6.6]
"We identify combination methods satisfying this criteria as those with the highest unweighted macro-average score over our unseen evaluation datasets (henceforth 'Mean' or 'Mean score') on each Reasoning Model... For the methods that utilize RR model scores we select the highest performing on this measure and refer to it as 'Generally best RR combo' below."
The 'Generally best RR combo' method and its RR threshold are chosen by maximizing the mean score on the exact unseen evaluation datasets whose results are then reported as the demonstration that RR 'significantly improves' performance (Section 6.3.3.2 and Table 6.6: 45.5 vs 42.7). This is test-set selection, not an unseen prediction: the reported gain is the maximum over the evaluated thresholds, and the paper's caveat that the method is not usable on an arbitrary question of unknown type is applied only to 'Best RR combo per dataset', not to 'Generally best RR combo'.
-
fitted input called prediction
[Section 4.2.3 (Similarity Computation Method); Table 4.2; Section 4.3.1]
"We identified a suitable value of T through an iterative process of evaluating the ten most similar sample pairs for each evaluation dataset at a possible value for T and increasing this value at each iteration until we found a value at which no memorisable sample pairs were identified but remaining sample counts are reasonable (Table 4.2). This value was identified as T = 60... For brevity we call this further subset 'Unmemorisable' as shorthand for 'unlikely to be memorisable from our training datasets, including TDND'."
The threshold T=60 is calibrated by the authors' own manual inspection of the most similar eval-train pairs on the very evaluation datasets (DROP, ROPES, etc.) whose 'Unmemorisable' subsets are then used to conclude that the TDND improvements are 'not attributable to memorisation' (Section 4.3.1, Table 4.5). The filter's output is the evidence for the claim, so the conclusion is equivalent to trusting the filter's calibration. The paper itself acknowledges it 'cannot eliminate evaluation samples that have answer overlap with any training sample', and the answer-overlap filter only checks the single most similar training sample, so the 'Unmemorisable' label is a shorthand judgment, not an independently verified property.
full rationale
The thesis has substantial independent content: Chapter 5 reports state-of-the-art finetuned results on IIRCG/IIRCR and competitive DROP results against external systems; Chapter 6 evaluates the Rationale Ranker on TruthfulQA MC1 against external LLM baselines; and the Iterator is compared with Baleen on Hover. These external benchmarks prevent a high circularity score. However, two load-bearing empirical claims are partially self-referential. First, the RR effectiveness claim is quantified with a 'Generally best RR combo' whose threshold is selected by maximizing the mean score over the target unseen evaluation datasets themselves; the paper's own caveat about unusability in a truly unseen setting is attached only to the per-dataset best combo, not to the generally best combo, even though both involve test-set selection. Second, the anti-memorisation claim in Chapter 4 relies on a subset ('Unmemorisable') defined by a threshold T=60 that was hand-calibrated on the same evaluation datasets, and the conclusion then treats performance on that subset as proof that gains are not attributable to memorisation. The paper is transparent about some limitations (e.g., inability to remove all answer-overlap samples), but the central conclusions still depend on the fitted subset and the fitted threshold. The RATD training/evaluation pipeline sharing is not circular by itself: it is an explicitly stated transfer hypothesis about context form. Overall, the central RR improvement and the anti-memorisation conclusion are partially reduced to their own fitted inputs, giving a score of 6 rather than 0-2.
Assumptions & free parameters
free parameters (4)
- Memorisability threshold T =
60
- Reranker sentence score paragraph weight w =
0.5
- Evidence Set scoring combination weight =
0.5p + 0.5se
- RR score threshold t =
grid 0.0005 to 0.9
assumptions (4)
- domain assumption Unseen evaluation datasets are disjoint from training datasets, so performance on them indicates generalization (Sections 1.1, 2.3).
- ad hoc to paper The semantic-similarity threshold T=60 identifies memorisable evaluation samples (Section 4.2.3).
- ad hoc to paper Iterator-generated noisy contexts for RATD training are similar enough to Iterator-generated evaluation contexts that training on the former transfers (Section 5.2.4.3).
- domain assumption The local Wikipedia corpus and the locally run LLMs provide sufficient evidence for the evaluation questions (Chapters 5 and 6).
Cite this review
Pith. "Pith review of Teaching Smaller Language Models To Generalise To Unseen Compositional Questions (Full Thesis)." pith.science (2026). https://pith.science/paper/4MBE7QST
@misc{pith2026241116985,
author = {Pith},
title = {Pith review of: Teaching Smaller Language Models To Generalise To Unseen Compositional Questions (Full Thesis)},
year = {2026},
howpublished = {\url{https://pith.science/paper/4MBE7QST}},
note = {Machine review of arXiv:2411.16985}
}
read the original abstract
Pretrained large Language Models (LLMs) are able to answer questions that are unlikely to have been encountered during training. However a diversity of potential applications exist in the broad domain of reasoning systems and considerations such as latency, cost, available compute resource and internet connectivity are relevant in determining an appropriate approach. We consider the setting where some local compute capacity is available at inference time but internet connectivity is not. Similar to a general-purpose LLM, we assume that our much smaller Reasoning Models may be asked arbitrary questions from unknown distributions, so we focus on evaluation in an unseen setting. We train our models to answer diverse questions by instilling an ability to reason over a retrieved context. We acquire context from two knowledge sources; a Wikipedia corpus queried using a multi-hop dense retrieval system with novel extensions, and from rationales generated from a larger Language Model optimised to run in a lower resource environment. Our main contributions: We propose novel methods to show that our model is capable of answering contextualised questions without memorisation. We establish a comprehensive set of baseline results on unseen evaluation datasets. We show that the addition of novel retrieval-augmented training datasets (RATD) to the training regime of the Reasoning Model significantly improves results. We demonstrate further significant improvement through the application of methods for combining knowledge from two sources. The first method (RR) involves training a novel Rationale Ranking model to score both generated rationales and retrieved contexts with respect to relevance and truthfulness. We use the scores to derive combined contexts. We also show that utilising the RATD datasets enables our model to become proficient at utilising combined noisy contexts.
Figures
Reference graph
Works this paper leans on
-
[1]
Anand, Z
Y. Anand, Z. Nussbaum, B. Duderstadt, B. Schmidt, and A. Mulyar. GPT4All : Training an assistant-style chatbot with large scale data distillation from GPT-3.5-Turbo . https://github.com/nomic-ai/gpt4all, 2023
2023
-
[2]
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, E. Chu, J. H. Clark, L. E. Shafey, Y. Huang, K. Meier-Hellstern, G. Mishra, E. Moreira, M. Omernick, K. Robinson, S. Ruder, Y. Tay, K. Xiao, Y. Xu, Y. Zhang, G. H. Abrego, J. Ahn, J. Austin, P. Barham, J. Botha, J. Bradbury, S. Brahma, K. Brooks, M...
arXiv 2023
-
[3]
Bahdanau, K
D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate. In International Conference On Learning Representations, 2015
2015
-
[4]
Bartolo, A
M. Bartolo, A. Roberts, J. Welbl, S. Riedel, and P. Stenetorp. Beat the AI : Investigating adversarial human annotation for reading comprehension. Transactions of the Association for Computational Linguistics, 8: 0 662--678, Nov. 2020
2020
-
[5]
Bengio, R
Y. Bengio, R. Ducharme, and P. Vincent. A neural probabilistic language model. In Advances In Neural Information Processing Systems, volume 13, 2000
2000
-
[6]
S. Bhakthavatsalam, D. Khashabi, T. Khot, B. D. Mishra, K. Richardson, A. Sabharwal, C. Schoenick, O. Tafjord, and P. Clark. Think you have solved direct-answer question answering? try ARC-DA , the direct-answer AI2 reasoning challenge. arXiv preprint arXiv:2102.03315, 2021
arXiv 2021
-
[7]
S. Bird, E. Klein, and E. Loper. Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit. O'Reilly Media, Inc., June 2009
2009
-
[8]
Y. Bisk, R. Zellers, R. Le bras, J. Gao, and Y. Choi. PIQA : Reasoning about physical commonsense in natural language. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 34(05), pages 7432--7439. Association for the Advancement of Artificial Intelligence, 2020
2020
Show all 200 references
-
[9]
Bordes, S
A. Bordes, S. Chopra, and J. Weston. Question answering with subgraph embeddings. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 615--620, Doha, Qatar, Oct. 2014 a . Association for Computational Linguistics
2014
-
[10]
Bordes, J
A. Bordes, J. Weston, and N. Usunier. Open question answering with weakly supervised embedding models. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2014 , pages 165--180, Berlin, Heidelberg, Sept. 2014 b . Springer-Verlag
2014
-
[11]
Bosselut, H
A. Bosselut, H. Rashkin, M. Sap, C. Malaviya, A. Celikyilmaz, and Y. Choi. COMET : Commonsense transformers for automatic knowledge graph construction. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4762--4779, Florence, Italy...
2019
-
[12]
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. ...
1901
-
[13]
Carlini, D
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang. Quantifying memorization across neural language models. In International Conference on Learning Representations, 2023
2023
-
[14]
Chatterjee
S. Chatterjee. Learning and memorization. In J. Dy and A. Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 755--763. PMLR, 2018
2018
-
[15]
D. Chen, A. Fisch, J. Weston, and A. Bordes. Reading W ikipedia to answer Open-Domain questions. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1870--1879, Vancouver, Canada, 2017. Association for Compu...
2017
-
[16]
X. Chen, K. Lakhotia, B. O g uz, A. Gupta, P. Lewis, S. Peshterliev, Y. Mehdad, S. Gupta, and W.-T. Yih. Salient phrase aware dense retrieval: Can a dense retriever imitate a sparse one? arXiv preprint arXiv:2110.06918, Oct. 2021
-
[17]
X. Chen, M. Lin, N. Sch \"a rli, and D. Zhou. Teaching large language models to Self-Debug . arXiv preprint arXiv:2304.05128, Apr. 2023
2023 arXiv
-
[18]
Chern, S
I.-C. Chern, S. Chern, S. Chen, W. Yuan, K. Feng, C. Zhou, J. He, G. Neubig, and P. Liu. FacTool : Factuality detection in generative AI -- a tool augmented framework for multi-task and multi-domain scenarios. arXiv preprint arXiv:307.13528, July 2023
2023
-
[19]
Chiang, Z
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90\ https://lmsys.org/blog/2023-03-30-vicuna/, March 2023
2023
-
[20]
Choshen, G
L. Choshen, G. Hacohen, D. Weinshall, and O. Abend. The grammar-learning trajectories of neural language models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8281--8297, Stroudsburg, PA, USA, May 2022...
2022
-
[21]
Chowdhery, S
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Au...
2022 arXiv
-
[22]
Clark and M
C. Clark and M. Gardner. Simple and effective Multi-Paragraph reading comprehension. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 845--855, Melbourne, Australia, July 2018. Association for Computation...
2018
-
[23]
Clark, K
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova. BoolQ : Exploring the surprising difficulty of natural Yes/No questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lan...
2019
-
[24]
Clark, M.-T
K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning. ELECTRA : Pre-training text encoders as discriminators rather than generators. In International Conference on Learning Representations, Mar. 2020 a
2020
-
[25]
Clark, O
P. Clark, O. Etzioni, T. Khot, A. Sabharwal, O. Tafjord, P. Turney, and D. Khashabi. Combining retrieval, statistics, and inference to answer elementary science questions. In AAAI Conference on Artificial Intelligence , volume 30. Association for the Advancement of Artificial ...
2016
-
[26]
Clark, I
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord. Think you have solved question answering? try ARC , the AI2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018
2018 arXiv
-
[27]
Clark, O
P. Clark, O. Etzioni, D. Khashabi, T. Khot, B. D. Mishra, K. Richardson, A. Sabharwal, C. Schoenick, O. Tafjord, N. Tandon, S. Bhakthavatsalam, D. Groeneveld, M. Guerquin, and M. Schmitz. From 'f' to 'a' on the n.y. regents science exams: An overview of the aristo project. arX...
1909 arXiv
-
[28]
Clark, O
P. Clark, O. Tafjord, and K. Richardson. Transformers as soft reasoners over language. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence ( IJCAI-20 ) , pages 3882--3890. International Joint Conferences on Artificial Intelligence Organ...
2020
-
[29]
Dankers, E
V. Dankers, E. Bruni, and D. Hupkes. The paradox of the compositionality of natural language: A neural machine translation case study. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4154--4175, Dublin, ...
2022
-
[30]
Dasigi, K
P. Dasigi, K. Lo, I. Beltagy, A. Cohan, N. A. Smith, and M. Gardner. A dataset of information-seeking questions and answers anchored in research papers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2021
-
[31]
Dem s ar
J. Dem s ar. Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7: 0 1--30, 2006
2006
-
[32]
Dettmers, M
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer. LLM.int8() : 8-bit matrix multiplication for transformers at scale. In 36th Conference on Neural Information Processing Systems, pages 30318--30332, Aug. 2022
2022
-
[33]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie...
2019
-
[34]
DeYoung, S
J. DeYoung, S. Jain, N. F. Rajani, E. Lehman, C. Xiong, R. Socher, and B. C. Wallace. ERASER : A benchmark to evaluate rationalized NLP models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4443--4458. Association for Computa...
2020
-
[35]
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. P. Bosma, Z. Zhou, T. Wang, E. Wang, K. Webster, M. Pellat, K. Robinson, K. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. Le, Y. Wu, Z. Chen, and C. Cui. G...
2022
-
[36]
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner. DROP : A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Hum...
2019
-
[37]
Efron and R
B. Efron and R. J. Tibshirani. An Introduction to the Bootstrap. Monographs on Statistics and Applied Probability, 57. Chapman and Hall, New York, NY, 1993
1993
-
[38]
Elangovan, J
A. Elangovan, J. He, and K. Verspoor. Memorization vs. generalization: Quantifying data leakage in NLP performance evaluation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, pages 1325--1335. Association for Comp...
2021
-
[39]
Y. Fang, S. Sun, Z. Gan, R. Pillai, S. Wang, and J. Liu. Hierarchical graph network for multi-hop question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 8823--8838, Online, Nov. 2020. Association for Comp...
2020
-
[40]
V. Feldman. Does learning require memorization? a short tale about a long tail. arXiv preprint arXiv 1906.05271, June 2019
1906 arXiv
-
[41]
Feldman and C
V. Feldman and C. Zhang. What neural networks memorize and why: Discovering the long tail via influence estimation. In Advances in Neural Information Processing Systems 33, pages 2881--2891, 2020
2020
-
[42]
Ferguson, M
J. Ferguson, M. Gardner, H. Hajishirzi, T. Khot, and P. Dasigi. IIRC : A dataset of incomplete information reading comprehension questions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 1137--1147, Stroudsburg, PA, ...
2020
-
[43]
Ferguson, H
J. Ferguson, H. Hajishirzi, P. Dasigi, and T. Khot. Retrieval data augmentation informed by downstream question answering performance. In Proceedings of the Fifth Fact Extraction and VERification Workshop ( FEVER ) , pages 1--5, 2022
2022
-
[44]
Fu, S.-K
J. Fu, S.-K. Ng, Z. Jiang, and P. Liu. GPTScore : Evaluate as you desire. arXiv preprint arXiv:2302.04166, Feb. 2023
2023 arXiv
-
[45]
Gardner, Y
M. Gardner, Y. Artzi, V. Basmov, J. Berant, B. Bogin, S. Chen, P. Dasigi, D. Dua, Y. Elazar, A. Gottumukkala, N. Gupta, H. Hajishirzi, G. Ilharco, D. Khashabi, K. Lin, J. Liu, N. F. Liu, P. Mulcaire, Q. Ning, S. Singh, N. A. Smith, S. Subramanian, R. Tsarfaty, E. Wallace, A. Z...
2020
-
[46]
M. Geva, Y. Goldberg, and J. Berant. Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Jo...
2019
-
[47]
M. Geva, A. Gupta, and J. Berant. Injecting numerical reasoning skills into language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 946--958, Online, 2020. Association for Computational Linguistics
2020
-
[48]
M. Geva, D. Khashabi, E. Segal, T. Khot, D. Roth, and J. Berant. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics, 9: 0 346--361, 2021
2021
-
[49]
Gottumukkala, D
A. Gottumukkala, D. Dua, S. Singh, and M. Gardner. Dynamic sampling strategies for multi-task reading comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 920--924, Stroudsburg, PA, USA, July 2020. Association for Com...
2020
-
[50]
B. F. Green, A. K. Wolf, C. Chomsky, and K. Laughery. Baseball: an automatic question-answerer. In Papers presented at the May 9-11, 1961, western joint IRE-AIEE-ACM computer conference , IRE-AIEE-ACM '61 (Western), pages 219--224, New York, NY, USA, May 1961. Association for ...
1961
-
[51]
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M. Chang. Retrieval augmented language model Pre-Training . In H. D. Iii and A. Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 3929--39...
2020
-
[52]
Hacohen, L
G. Hacohen, L. Choshen, and D. Weinshall. Let's agree to agree: Neural networks share classification order on real datasets. In International Conference on Machine Learning, pages 3950--3960. proceedings.mlr.press, 2020
2020
-
[53]
S. M. Harabagiu, D. I. Moldovan, M. Pasca, R. Mihalcea, M. Surdeanu, R. C. Bunescu, R. Girju, V. Rus, and P. Morarescu. FALCON : Boosting knowledge for answer engines. In TREC , volume 9, pages 479--488. trec.nist.gov, 2000
2000
-
[54]
Hartill, N
T. Hartill, N. TAN, M. Witbrock, and P. J. Riddle. Teaching smaller language models to generalise to unseen compositional questions. Transactions on Machine Learning Research, Aug. 2023
2023
-
[55]
S. Herbold. Autorank: A python package for automated ranking of classifiers. Journal of Open Source Software, 5 0 (48): 0 2173, Apr. 2020
2020
-
[56]
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom. Teaching machines to read and comprehend. In Advances In Neural Information Processing Systems 28, 2015
2015
-
[57]
Hirschman, M
L. Hirschman, M. Light, E. Breck, and J. D. Burger. Deep read: a reading comprehension system. In Proceedings of the 37th annual meeting of the Association for Computational Linguistics on Computational Linguistics, ACL '99, pages 325--332, USA, June 1999. Association for Comp...
1999
-
[58]
Hochreiter and J
S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural Computation, 9 0 (8): 0 1735--1780, Nov. 1997
1997
-
[59]
Holtzman, J
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi. The curious case of neural text degeneration. In International Conference on Learning Representations, Sept. 2019
2019
-
[60]
Hsieh, C.-L
C.-Y. Hsieh, C.-L. Li, C.-K. Yeh, H. Nakhost, Y. Fujii, A. Ratner, R. Krishna, C.-Y. Lee, and T. Pfister. Distilling Step-by-Step ! outperforming larger language models with less training data and smaller model sizes. In Findings of the Association for Computational Linguistic...
2023
-
[61]
Huang, Y
Y. Huang, Y. Li, Y. Xu, L. Zhang, R. Gan, J. Zhang, and L. Wang. MVP-Tuning : Multi-View knowledge retrieval with prompt tuning for commonsense reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages ...
2023
-
[62]
Hupkes, V
D. Hupkes, V. Dankers, M. Mul, and E. Bruni. Compositionality decomposed: How do neural networks generalise? Journal of Artificial Intelligence Research, 67: 0 757--795, 2020
2020
-
[63]
Inoue, P
N. Inoue, P. Stenetorp, and K. Inui. R4C : A benchmark for evaluating RC systems to get the right answer for the right reason. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6740--6750. Association for Computational Linguistics., 2020
2020
-
[64]
Iyyer, J
M. Iyyer, J. Boyd-Graber, L. Claudino, R. Socher, and H. Daum \'e , III. A neural network for factoid question answering over paragraphs. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 633--644, Doha, Qatar, Oct. 201...
2014
-
[65]
Izacard and E
G. Izacard and E. Grave. Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874--880, Online, 2021. Associati...
2021
-
[66]
Izacard, M
G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave. Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research, Aug. 2022
2022
-
[67]
Jhamtani and P
H. Jhamtani and P. Clark. Learning to explain: Datasets and models for identifying valid reasoning chains in multihop Question-Answering . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, pages 137--150, Online, 2020. Association for C...
2020
-
[68]
Jiang, S
Y. Jiang, S. Bordia, Z. Zhong, C. Dognin, M. Singh, and M. Bansal. HoVer : A dataset for Many-Hop fact extraction and claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2020 , pages 3441--3460. Association for Computational Linguistics, 2020
2020
-
[69]
Q. Jin, B. Dhingra, Z. Liu, W. Cohen, and X. Lu. PubMedQA : A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing...
2019
-
[70]
Johnson, M
J. Johnson, M. Douze, and H. Jegou. Billion-scale similarity search with GPUs . IEEE transactions on big data, 7 0 (3): 0 535--547, 2019
2019
-
[71]
Joshi, E
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer. TriviaQA : A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601--1611, Stro...
2017
-
[72]
Jurafsky and J
D. Jurafsky and J. H. Martin. Speech and language processing: An introduction to natural language processing, computational linguistics, and speech recognition (3rd edition draft). https://web.stanford.edu/ jurafsky/slp3/ed3book_jan72023.pdf, 2023. Accessed: 2023-10-17
2023
-
[73]
Kambhatla, T
G. Kambhatla, T. Nguyen, and E. Choi. Quantifying Train-Evaluation overlap with nearest neighbors. In A. Rogers, J. Boyd-Graber, and N. Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023 , pages 2905--2920, Toronto, Canada, July 2023. Associa...
2023
-
[74]
Kandpal, H
N. Kandpal, H. Deng, A. Roberts, E. Wallace, and C. Raffel. Large language models struggle to learn Long-Tail knowledge. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning...
2023
-
[75]
Karpukhin, B
V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-T. Yih. Dense passage retrieval for Open-Domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 6769--6781, Online, Nov. 2...
2020
-
[76]
Khashabi, S
D. Khashabi, S. Chaturvedi, M. Roth, S. Upadhyay, and D. Roth. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: H...
2018
-
[77]
Khashabi, T
D. Khashabi, T. Khot, and A. Sabharwal. More bang for your buck: Natural perturbation for robust question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 163--170, Online, Nov. 2020 a . Association for Comp...
2020
-
[78]
Khashabi, S
D. Khashabi, S. Min, T. Khot, A. Sabharwal, O. Tafjord, P. Clark, and H. Hajishirzi. UNIFIEDQA : Crossing format boundaries with a single QA system. In Findings of the Association for Computational Linguistics: EMNLP 2020 , pages 1896--1907, Online, 2020 b . Association for Co...
2020
-
[79]
Khashabi, Y
D. Khashabi, Y. Kordi, and H. Hajishirzi. UnifiedQA-v2 : Stronger generalization via broader cross-format training. arXiv preprint arXiv: 2202.12359, Feb. 2022
2022 arXiv
-
[80]
Khattab, C
O. Khattab, C. Potts, and M. Zaharia. Baleen: Robust multi-hop reasoning at scale via condensed retrieval. In Advances in Neural Information Processing Systems, 34, pages 27670--27682, 2021
2021
-
[81]
T. Khot, P. Clark, M. Guerquin, P. Jansen, and A. Sabharwal. QASC : A dataset for question answering via sentence composition. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 34(05), pages 8082--8090. Association for the Advancement of Artificial Inte...
2020
-
[82]
T. N. Kipf and M. Welling. Semi-Supervised classification with graph convolutional networks. In International Conference on Learning Representations, 2017
2017
-
[83]
Ko c isk \'y , J
T. Ko c isk \'y , J. Schwarz, P. Blunsom, C. Dyer, K. M. Hermann, G. Melis, and E. Grefenstette. The narrativeqa reading comprehension challenge. Transactions of the Association for Computational Linguistics, 6: 0 317--328, 2018
2018
-
[84]
o pf, Y. Kilcher, D. von R \
A. K \"o pf, Y. Kilcher, D. von R \"u tte, S. Anagnostidis, Z.-R. Tam, K. Stevens, A. Barhoum, N. M. Duc, O. Stanley, R. Nagyfi, E. S. Shahul, S. Suri, D. Glushkov, A. Dantuluri, A. Maguire, C. Schuhmann, H. Nguyen, and A. Mattick. OpenAssistant conversations -- democratizing ...
2023 arXiv
-
[85]
Krishna, A
K. Krishna, A. Roy, and M. Iyyer. Hurdles to progress in long-form question answering. In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou, editors, Proceedings of the 2021 Conference of the North Ame...
2021
-
[86]
Kwiatkowski, J
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov. Natural questions: A benchmark for question answering resea...
2019
-
[87]
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy. RACE : Large-scale ReAding comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA, 2017. Association for Computational Linguistics
2017
-
[88]
Le Scao, A
T. Le Scao, A. Fan, C. Akiki, E. Pavlick, S. Ili \'c , D. Hesslow, R. Castagn \'e , A. S. Luccioni, F. Yvon, M. Gall \'e , J. Tow, A. M. Rush, S. Biderman, A. Webson, P. S. Ammanamanchi, T. Wang, B. Sagot, N. Muennighoff, A. V. del Moral, O. Ruwase, R. Bawden, S. Bekman, A. Mc...
2022 arXiv
-
[89]
Lee, M.-W
K. Lee, M.-W. Chang, and K. Toutanova. Latent retrieval for weakly supervised open domain question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6086--6096. Association for Computational Linguistics, 2019
2019
-
[90]
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini. Deduplicating training data makes language models better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8424--8445,...
2022
-
[91]
Lewis, Y
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer. BART : Denoising Sequence-to-Sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association ...
2020
-
[92]
u ttler, M. Lewis, W.-T. Yih, T. Rockt \
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. K \"u ttler, M. Lewis, W.-T. Yih, T. Rockt \"a schel, S. Riedel, and D. Kiela. Retrieval-Augmented generation for Knowledge-Intensive NLP tasks. In Advances in Neural Information Processing Systems, volume 3...
2020
-
[93]
Lewis, P
P. Lewis, P. Stenetorp, and S. Riedel. Question and answer Test-Train overlap in Open-Domain question answering datasets. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1000--1008, Online, 2021...
2021
-
[94]
D. Li, A. S. Rawat, M. Zaheer, X. Wang, M. Lukasik, A. Veit, F. Yu, and S. Kumar. Large language models with controllable working memory. arXiv preprint arXiv:2211.05110, Nov. 2022
2022 arXiv
-
[95]
L. H. Li, J. Hessel, Y. Yu, X. Ren, K.-W. Chang, and Y. Choi. Symbolic Chain-of-Thought distillation: Small models can also ``think'' step-by-step. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2665--2...
2023
-
[96]
Liang, T
Z. Liang, T. Khot, S. Bethard, M. Surdeanu, and A. Sabharwal. Better retrieval may not lead to better question answering. arXiv preprint arXiv:2205.03685, May 2022
2022 arXiv
-
[97]
K. Lin, O. Tafjord, P. Clark, and M. Gardner. Reasoning over paragraph effects in situations. In Proceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 58--62. Association for Computational Linguistics, 2019
2019
-
[98]
S. Lin, J. Hilton, and O. Evans. TruthfulQA : Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214--3252, Stroudsburg, PA, USA, May 2022. Association for Comp...
2022
-
[99]
J. Liu, A. Liu, X. Lu, S. Welleck, P. West, R. Le Bras, Y. Choi, and H. Hajishirzi. Generated knowledge prompting for commonsense reasoning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3154--3169, St...
2022
-
[100]
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov. RoBERTa : A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692, July 2019
1907 arXiv
-
[101]
Longpre, K
S. Longpre, K. Perisetla, A. Chen, N. Ramesh, C. DuBois, and S. Singh. Entity-Based knowledge conflicts in question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7052--7063, Online and Punta Cana, Dominican Republic...
2021
-
[102]
Lourie, R
N. Lourie, R. Le Bras, C. Bhagavatula, and Y. Choi. UNICORN on RAINBOW : A universal commonsense reasoning model on a new multitask benchmark. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 35 of 15, pages 13480--13488, May 2021
2021
-
[103]
Madaan, N
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Welleck, B. P. Majumder, S. Gupta, A. Yazdanbakhsh, and P. Clark. Self-refine: Iterative refinement with self-feedback. arXiv preprint arXiv:2303.17651, Mar. 2023
2023 arXiv
-
[104]
L. C. Magister, J. Mallinson, J. Adamek, E. Malmi, and A. Severyn. Teaching small language models to reason. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1773--1781, Toronto, Canada, July 2023. Assoc...
2023
-
[105]
Y. A. Malkov and D. A. Yashunin. Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42 0 (4): 0 824--836, 2018
2018
-
[106]
Manakul, A
P. Manakul, A. Liusie, and M. J. F. Gales. SelfCheckGPT : Zero-resource black-box hallucination detection for generative large language models. arXiv preprint arXiv:2303.08896, Mar. 2023
2023 arXiv
-
[107]
Manning and H
C. Manning and H. Schutze. Foundations of Statistical Natural Language Processing. MIT Press, May 1999
1999
-
[108]
A. A. Markov. Essai d'une recherche statistique sur le texte du roman ``eugene onegin'' illustrant la liaison des epreuve en chain (`example of a statistical investigation of the text of ``eugene onegin'' illustrating the dependence between samples in chain'). Izvistia Imperat...
1913
-
[109]
Mihaylov, P
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381--2391, Brussels, Belgium, 2018. Asso...
2018
-
[110]
Mikolov, K
T. Mikolov, K. Chen, G. Corrado, and J. Dean. Efficient estimation of word representations in vector space. In ICLR Workshop , Jan. 2013 a
2013
-
[111]
Mikolov, I
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26, 2013 b
2013
-
[112]
S. Min, E. Wallace, S. Singh, M. Gardner, H. Hajishirzi, and L. Zettlemoyer. Compositional questions do not necessitate multi-hop reasoning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4249--4257, Florence, Italy, July 2019...
2019
-
[113]
S. Min, K. Krishna, X. Lyu, M. Lewis, W.-T. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi. FActScore : Fine-grained atomic evaluation of factual precision in long form text generation. arXiv preprint arXiv:2305.14251, May 2023
2023 arXiv
-
[114]
H. Moravec. Mind Children: The Future of Robot and Human Intelligence. Harvard University Press, 1988
1988
-
[115]
H. T. Ng, L. H. Teo, and J. L. P. Kwan. A machine learning approach to answering questions for reading comprehension tests. In Proceedings of the 2000 Joint SIGDAT conference on Empirical methods in natural language processing and very large corpora: held in conjunction with t...
2000
-
[116]
Y. Onoe, M. J. Q. Zhang, E. Choi, and G. Durrett. CREAK : A dataset for commonsense reasoning over entity knowledge. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), Nov. 2021
2021
-
[117]
GPT-4 technical report
OpenAI . GPT-4 technical report. arXiv preprint arXiv:2303.08774, Mar. 2023
2023 arXiv
-
[118]
Ouyang, J
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe. Training language models to follow instructions ...
2022
-
[119]
X. Pan, W. Yao, H. Zhang, D. Yu, D. Yu, and J. Chen. Knowledge-in-Context : Towards knowledgeable Semi-Parametric language models. In The Eleventh International Conference on Learning Representations, 2023
2023
-
[120]
B. Partee. Compositionality. Varieties of Formal Semantics: Proceedings of the fourth Amsterdam colloquium, 3: 0 281--311, 1984
1984
-
[121]
X. Pi, Q. Liu, B. Chen, M. Ziyadi, Z. Lin, Y. Gao, Q. Fu, J.-G. Lou, and W. Chen. Reasoning like program executors. arXiv preprint arXiv:2201.11473, 2022
2022 arXiv
-
[122]
Piktus, F
A. Piktus, F. Petroni, V. Karpukhin, D. Okhonko, S. Broscheit, G. Izacard, P. Lewis, B. O g uz, E. Grave, W.-T. Yih, and S. Riedel. The web is your oyster - knowledge-intensive NLP against a very large web corpus. arXiv preprint arXiv:2112.09924, Dec. 2021
2021 arXiv
-
[123]
P. Qi, H. Lee, T. Sido, and C. Manning. Answering Open-Domain questions of varying reasoning steps from text. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3599--3614, Online and Punta Cana, Dominican Republic, Nov. 2021. Asso...
2021
-
[124]
Radford, K
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever. Improving language understanding by generative pre-training. http://openai-assets.s3.amazonaws.com/research-covers/language-unsupervised/language_understanding_paper.pdf, 2018
2018
-
[125]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. http://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf, 2019
2019
-
[126]
Raffel, N
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified Text-to-Text transformer. Journal of Machine Learning Research, 21: 0 1--67, 2020
2020
-
[127]
Rajpurkar, J
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang. SQuAD : 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383--2392. Association for Computational Lingustics, 2016
2016
-
[128]
Rajpurkar, R
P. Rajpurkar, R. Jia, and P. Liang. Know what you don't know: Unanswerable questions for SQuAD . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 784--789. Association for Computational Linguistics, 2018
2018
-
[129]
O. Ram, G. Shachaf, O. Levy, J. Berant, and A. Globerson. Learning to retrieve passages without supervision. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2687--2700, Sea...
2022
-
[130]
Razeghi, R
Y. Razeghi, R. L. Logan, IV, M. Gardner, and S. Singh. Impact of pretraining term frequencies on Few-Shot numerical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2022 , pages 840--854, Abu Dhabi, United Arab Emirates, Dec. 2022. Association for...
2022
-
[131]
Reimers and I
N. Reimers and I. Gurevych. Sentence- BERT : Sentence embeddings using S iamese BERT -Networks . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing ( EMNLP-IJCNLP )...
2019
-
[132]
Richardson, C
M. Richardson, C. J. C. Burges, and E. Renshaw. Mctest: A challenge dataset for the open-domain machine comprehension of text. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 193--203. Association for Computational Linguistics, 2013
2013
-
[133]
Riloff and M
E. Riloff and M. Thelen. A rule-based question answering system for reading comprehension tests. In ANLP/NAACL 2000 Workshop on Reading comprehension tests as evaluation for computer-based language understanding systems , Morristown, NJ, USA, 2000. Association for Computationa...
2000
-
[134]
Roberts, C
A. Roberts, C. Raffel, and N. Shazeer. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 5418--5426. Association for Computational Linguistics, 2020
2020
-
[135]
Rogers, O
A. Rogers, O. Kovaleva, M. Downey, and A. Rumshisky. Getting closer to AI complete question answering: A set of prerequisite real tasks. In AAAI Conference on Artificial Intelligence ( AAAI-20 ) , volume 34, pages 8722--8731. Association for the Advancement of Artificial Intel...
2020
-
[136]
Rogers, M
A. Rogers, M. Gardner, and I. Augenstein. QA dataset explosion: A taxonomy of NLP resources for question answering and reading comprehension. ACM Computing Surveys, 55 0 (10): 0 1--45, Feb. 2023
2023
-
[137]
Sakaguchi, R
K. Sakaguchi, R. Le Bras, C. Bhagavatula, and Y. Choi. WinoGrande : An adversarial winograd schema challenge at scale. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 34, pages 8732--8740. Association for the Advancement of Artificial Intelligence, 2020
2020
-
[138]
V. Sanh, A. Webson, C. Raffel, S. H. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, T. Le Scao, A. Raja, M. Dey, M. Saiful Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. Nayak, D. Datta, J. Chang, M. T.-J. Jiang, H. Wang, M. Manica, S. ...
2021
-
[139]
M. Sap, R. Le Bras, E. Allaway, C. Bhagavatula, N. Lourie, H. Rashkin, B. Roof, N. A. Smith, and Y. Choi. ATOMIC : An atlas of machine commonsense for if-then reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence , 33(01), pages 3027--3035, 2019 a
2019
-
[140]
M. Sap, H. Rashkin, D. Chen, R. Le Bras, and Y. Choi. Social IQa : Commonsense reasoning about social interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processi...
2019
-
[141]
Schwarzschild, E
A. Schwarzschild, E. Borgnia, A. Gupta, F. Huang, U. Vishkin, M. Goldblum, and T. Goldstein. Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks. In Advances in Neural Information Processing Systems, volume 34, pages 6695--6706, 2021
2021
-
[142]
P. Sen, A. F. Aji, and A. Saffari. Mintaka: A complex, natural, and multilingual dataset for End-to-End question answering. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1604--1619, Gyeongju, Republic of Korea, Oct. 2022. International...
2022
-
[143]
Shridhar, A
K. Shridhar, A. Stolfo, and M. Sachan. Distilling reasoning capabilities into smaller language models. In Findings of the Association for Computational Linguistics: ACL 2023 , pages 7059--7073, Toronto, Canada, July 2023. Association for Computational Linguistics
2023
-
[144]
Shwartz, P
V. Shwartz, P. West, R. Le Bras, C. Bhagavatula, and Y. Choi. Unsupervised commonsense question answering with self-talk. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 4615--4629, Stroudsburg, PA, USA, Nov. 2020. As...
2020
-
[145]
C. Si, W. Shi, C. Zhao, L. Zettlemoyer, and J. Boyd-Graber. Mixture of prompt experts for generalizable and interpretable question answering. arXiv preprint arXiv 2305.14628, May 2023
2023 arXiv
-
[146]
R. F. Simmons, S. Klein, and K. McConlogue. Indexing and dependency logic for answering english questions. American Documentation, 15 0 (3): 0 196--204, July 1964
1964
-
[147]
Sinha, S
K. Sinha, S. Sodhani, J. Dong, J. Pineau, and W. L. Hamilton. CLUTRR : A diagnostic benchmark for inductive reasoning from text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Lang...
2019
-
[148]
Sp \"a rck Jones
K. Sp \"a rck Jones. A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28 0 (1): 0 11--21, 1972
1972
-
[149]
Speer, J
R. Speer, J. Chin, and C. Havasi. ConceptNet 5.5: An open multilingual graph of general knowledge. In Proceedings of the AAAI Conference on Artificial Intelligence , 31(1), pages 4444--4451, 2017
2017
-
[150]
Srivastava, A
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, A. Kluska, A. Lewkowycz, A. Agarwal, A. Power, A. Ray, A. Warstadt, A. W. Kocurek, A. Safaya, A. Tazarv, A. Xiang, A. Parrish, A. Nie, A. Hussain, A. Ask...
2022 arXiv
-
[151]
Stability AI releases StableVicuna, the AI World’s First Open Source RLHF LLM Chatbot
Stability-AI . Stability AI releases StableVicuna, the AI World’s First Open Source RLHF LLM Chatbot . https://stability.ai/blog/stablevicuna-open-source-rlhf-chatbot/, Apr. 2023. Accessed: 2023-7-5
2023
-
[152]
H. Sun, B. Dhingra, M. Zaheer, K. Mazaitis, R. Salakhutdinov, and W. Cohen. Open domain question answering using early fusion of knowledge bases and text. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4231--4242, Brussels, Bel...
2018
-
[153]
Sutskever, O
I. Sutskever, O. Vinyals, and Q. V. Le. Sequence to sequence learning with neural networks. In Advances In Neural Information Processing Systems 27, volume 27, 2014
2014
-
[154]
Talmor and J
A. Talmor and J. Berant. The web as a Knowledge-Base for answering complex questions. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 641--651, New ...
2018
-
[155]
Talmor, J
A. Talmor, J. Herzig, N. Lourie, and J. Berant. CommonsenseQA : A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vo...
2019
-
[156]
Talmor, O
A. Talmor, O. Yoran, R. Le Bras, C. Bhagavatula, Y. Goldberg, Y. Choi, and J. Berant. CommonsenseQA 2.0: Exposing the limits of AI through gamification. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), Nov. 2021
2021
-
[157]
Taori, I
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto. Stanford alpaca: An instruction-following LLaMA model. https://github.com/tatsu-lab/stanford_alpaca, 2023
2023
-
[158]
Y. Tay, J. Wei, H. W. Chung, V. Q. Tran, D. R. So, S. Shakeri, X. Garcia, H. S. Zheng, J. Rao, A. Chowdhery, D. Zhou, D. Metzler, S. Petrov, N. Houlsby, Q. V. Le, and M. Dehghani. Transcending scaling laws with 0.1\ arXiv preprint arXiv:2210.11399, Oct. 2022
-
[159]
Taylor, M
R. Taylor, M. Kardas, G. Cucurull, T. Scialom, A. Hartshorn, E. Saravia, A. Poulton, V. Kerkez, and R. Stojnic. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085, Nov. 2022
2022 arXiv
-
[160]
W. L. Taylor. ``cloze procedure'': A new tool for measuring readability. Journalism Quarterly, 30 0 (4): 0 415--433, Sept. 1953
1953
-
[161]
Thakur, N
N. Thakur, N. Reimers, A. R \"u ckl \'e , A. Srivastava, and I. Gurevych. BEIR : A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), Oct. 2021
2021
-
[162]
Thorne, A
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal. FEVER : A large-scale dataset for fact extraction and VERification . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, ...
2018
-
[163]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi \`e re, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample. LLaMA : Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, Feb. 2023
2023 arXiv
-
[164]
Trischler, T
A. Trischler, T. Wang, X. Yuan, J. Harris, A. Sordoni, P. Bachman, and K. Suleman. NewsQA : A machine comprehension dataset. In Proceedings of the 2nd Workshop on Representation Learning for NLP , pages 191--200. Association for Computational Linguistics, 2017
2017
-
[165]
Trivedi, N
H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal. MuSiQue : Multihop questions via single-hop question composition. Transactions of the Association for Computational Linguistics, 10: 0 539--554, 2022 a
2022
-
[166]
Trivedi, N
H. Trivedi, N. Balasubramanian, T. Khot, and A. Sabharwal. Teaching broad reasoning skills for Multi-Step QA by generating hard contexts. arXiv preprint arXiv:2205.12496, May 2022 b
2022 arXiv
-
[167]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, pages 5998--6008, 2017
2017
-
[168]
E. M. Voorhees. The TREC question answering track. Natural Language Engineering, 7 0 (4): 0 361--378, Dec. 2001
2001
-
[169]
Wallace, Y
E. Wallace, Y. Wang, S. Li, S. Singh, and M. Gardner. Do NLP models know numbers? probing numeracy in embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing...
2019
-
[170]
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman. SuperGLUE : A stickier benchmark for General-Purpose language understanding systems. In Advances in Neural Information Processing Systems, 32, May 2019 a
2019
-
[171]
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman. GLUE : A Multi-Task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations, 2019 b
2019
-
[172]
S. Wang, M. Yu, X. Guo, Z. Wang, T. Klinger, W. Zhang, S. Chang, G. Tesauro, B. Zhou, and J. Jiang. R ^ 3 : Reinforced Ranker-Reader for open-domain question answering. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 32. Association for the Advancemen...
2018
-
[173]
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou. Self-Consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, Mar. 2022 a
2022 arXiv
-
[174]
Z. Wang, X. Pan, D. Yu, D. Yu, J. Chen, and H. Ji. Zemi: Learning Zero-Shot Semi-Parametric language models from multiple tasks. arXiv preprint arXiv:2210.00185, Oct. 2022 b
2022 arXiv
-
[175]
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le. Finetuned language models are zero-shot learners. In International Conference on Learning Representations, 2021
2021
-
[176]
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou. Chain of thought prompting elicits reasoning in large language models. In Thirty-sixth Conference on Neural Information Processing Systems ( NeurIPS 2022) , Jan. 2022
2022
-
[177]
Wiegreffe and A
S. Wiegreffe and A. Marasovi \'c . Teach me to explain: A review of datasets for explainable NLP . arXiv:2102.12060 [cs.CL], 2021
2021 arXiv
-
[178]
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush. Transformers: State-of-the-art natural l...
2020
-
[179]
Wolfson, M
T. Wolfson, M. Geva, A. Gupta, M. Gardner, Y. Goldberg, D. Deutch, and J. Berant. Break it down: A question understanding benchmark. Transactions of the Association for Computational Linguistics, 8: 0 183--198, 2020
2020
-
[180]
C.-S. Wu, A. Madotto, W. Liu, P. Fung, and C. Xiong. QAConv : Question answering on informative conversations. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5389--5411, Stroudsburg, PA, USA, 2022. Asso...
2022
-
[181]
D. Wu, J. Zhang, and X. Huang. Chain of thought prompting elicits knowledge augmentation. In Findings of the Association for Computational Linguistics: ACL 2023 , pages 6519--6534. Association for Computational Linguistics, July 2023
2023
-
[182]
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin. Unsupervised feature learning via non-parametric instance discrimination. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 3733--3742. IEEE, June 2018
2018
-
[183]
Z. Xie, S. Thiem, J. Martin, E. Wainwright, S. Marmorstein, and P. Jansen. W orld T ree v2: A corpus of science-domain structured explanations and inference patterns supporting multi-hop inference. In Proceedings of the 12th Language Resources and Evaluation Conference, pages ...
2020
-
[184]
Xiong, J
W. Xiong, J. Wu, H. Wang, V. Kulkarni, M. Yu, S. Chang, X. Guo, and W. Y. Wang. TWEETQA : A social media focused question answering dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5020--5031, Florence, Italy, July 2019...
2019
-
[185]
Xiong, X
W. Xiong, X. Li, S. Iyer, J. Du, P. Lewis, W. Y. Wang, Y. Mehdad, S. Yih, S. Riedel, D. Kiela, and B. Oguz. Answering complex Open-Domain questions with Multi-Hop dense retrieval. In International Conference on Learning Representations, 2021
2021
-
[186]
Y. Xu, C. Zhu, S. Wang, S. Sun, H. Cheng, X. Liu, J. Gao, P. He, M. Zeng, and X. Huang. Human parity on CommonsenseQA : Augmenting Self-Attention with external attention. arXiv preprint arXiv: 2112.03254, Dec. 2021
2021 arXiv
-
[187]
Y. Xu, C. Zhu, S. Wang, S. Sun, H. Cheng, X. Liu, J. Gao, P. He, M. Zeng, and X. Huang. Human parity on commonsenseqa: Augmenting self-attention with external attention. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22 , pa...
2022
-
[188]
Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. Cohen, R. Salakhutdinov, and C. D. Manning. HotpotQA : A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369--2380. Associat...
2018
-
[189]
Yoran, A
O. Yoran, A. Talmor, and J. Berant. Turning tables: Generating examples from semi-structured tables for endowing language models with reasoning skills. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 601...
2022
-
[190]
W. Yu, Z. Jiang, Y. Dong, and J. Feng. ReClor : A reading comprehension dataset requiring logical reasoning. In International Conference on Learning Representations, Feb. 2020
2020
-
[191]
W. Yu, C. Zhu, Z. Zhang, S. Wang, Z. Zhang, Y. Fang, and M. Jiang. Retrieval augmentation for commonsense reasoning: A unified approach. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4364--4377. Association for Computational L...
2022
-
[192]
W. Yu, D. Iter, S. Wang, Y. Xu, M. Ju, S. Sanyal, C. Zhu, M. Zeng, and M. Jiang. Generate rather than retrieve: Large language models are strong context generators. In International Conference on Learning Representations, 2023
2023
-
[193]
W. Yuan, G. Neubig, and P. Liu. Bartscore: Evaluating generated text as text generation. Advances in Neural Information Processing Systems, 34: 0 27263--27277, 2021
2021
-
[194]
Zhang, S
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals. Understanding deep learning requires rethinking generalization. In International Conference on Learning Representations, 2017
2017
-
[195]
Zhang, S
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals. Understanding deep learning (still) requires rethinking generalization. Communications of the ACM, 64 0 (3): 0 107--115, Mar. 2021
2021
-
[196]
Zhang, X
S. Zhang, X. Liu, J. Liu, J. Gao, K. Duh, and B. Van Durme. ReCoRD : Bridging the gap between human and machine commonsense reading comprehension. arXiv preprint arXiv:1810.12885, Oct. 2018
2018 arXiv
-
[197]
Zhang, S
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, T. Mihaylov, M. Ott, S. Shleifer, K. Shuster, D. Simig, P. S. Koura, A. Sridhar, T. Wang, and L. Zettlemoyer. OPT : Open pre-trained transformer language models. arXiv preprint ar...
2022 arXiv
-
[198]
W. Zhao, M. Geva, B. Y. Lin, M. Yasunaga, A. Madaan, and T. Yu. Complex reasoning in natural languag. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 6: Tutorial Abstracts), pages 11--20, Toronto, Canada, July 2023. Associatio...
2023
-
[199]
F. Zhu, W. Lei, Y. Huang, C. Wang, S. Zhang, J. Lv, F. Feng, and T.-S. Chua. TAT-QA : A question answering benchmark on a hybrid of tabular and textual content in finance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th I...
2021
-
[200]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.