REVIEW 4 major objections 8 minor 149 references
The Evolution of Natural Language Processing: How Prompt Optimization and Language Models are Shaping the Future
T0 review · 4 major / 8 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This review claims that 45 prompt-optimization strategies can be grouped into 11 distinct working-paradigm classes, and it catalogs the tasks, models, and datasets used to evaluate them.
desk verdict A broad compilation of prompt optimization methods with a useful benchmark critique, but the central 11-class taxonomy is internally inconsistent and needs major correction before it can be used. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is the working-paradigm taxonomy of Figure 6, which assigns each of the 45 methods to one of 11 classes based on the optimization principle used to refine prompts. The classes distinguish whether a method optimizes continuous embeddings or discrete text, and whether it relies on gradients, single-layer or multi-layer insertion, reinforcement learning, enumeration, evolutionary search, in-context learning, an LLM as optimizer, human collaboration, or Bayesian optimization. This taxonomy carries the argument by converting a scattered set of methods into a structured map that can be used for comparison and method selection.
What would settle it
Counting whether any method appears in more than one class of the paper's Figure 6, or comparing the stated 3-soft/42-hard split against the paper's own list of continuous prompt methods, would settle whether the 11 classes partition the 45 strategies.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the existing 45 prompt-optimization strategies can be understood through a taxonomy of 11 working paradigms, ranging from soft prompts optimized at single or multiple layers to hard prompts refined by gradients, reinforcement learning, evolutionary search, in-context learning, human-LLM collaboration, and Bayesian optimization. The paper further claims that mapping each method to its paradigm, its evaluation tasks, and its pretrained model reveals where the field is concentrated and where evaluation practices are inconsistent. It presents this taxonomy and the accompanying performance tables as a foundation for future comparative studies of prompt optimization and LLM-based predictive pipelines.
Load-bearing premise
The load-bearing premise is that each of the 45 methods belongs to exactly one of the 11 classes, so the taxonomy is a clean partition; the paper's own double assignments, such as BPO appearing under both human-LLM collaboration and Bayesian optimization, would undermine that premise.
Editorial extensions
If this is right
- Researchers can locate a new prompt-optimization method within an existing paradigm and compare it against direct relatives rather than against unrelated strategies.
- The review makes it possible to see which NLP tasks, benchmark datasets, and pretrained models dominate evaluation, exposing gaps such as limited multilingual and domain-specific testing.
- The documented inconsistencies in dataset splits, evaluation metrics, and sample sizes imply that reported accuracy numbers cannot yet be read as reliable cross-method rankings.
- The taxonomy highlights emerging directions, especially hybrid approaches that combine human expertise with automated search, as candidates for further investment.
- Because most methods are evaluated on only a few models, the field still lacks evidence on how well prompt-optimization strategies generalize across different model architectures and scales.
Reading between the lines
- The paper's contribution is stronger as a catalog than as a rigorous partition: several of its own placements, such as assigning BPO to both human-LLM collaboration and Bayesian optimization, and AutoPrompt to both gradient-based and hard-prompt categories, suggest the 11 classes overlap and should be treated as a menu of paradigms rather than a mutually exclusive grouping.
- A natural extension of the review would be to convert the taxonomy into a decision procedure that selects a class by prompt type (discrete versus continuous) and model access (white-box versus black-box), which the paper describes but does not formalize.
- The inconsistency the paper documents could be turned into a testable recommendation: evaluate a common suite of methods on one shared set of datasets, metrics, and splits to see whether the reported performance gaps persist or shrink.
- The paper's observation that decoder-based and GPT-family models dominate evaluation implies that prompt-optimization results may be architecture-specific, and testing the same 45 strategies on encoder-only and encoder-decoder models would likely change the comparative picture.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a systematic review of prompt optimization strategies for large language models in NLP. The authors describe a two-stage literature selection process (379 candidate articles reduced to 45), propose an 11-class taxonomy of prompt optimization methods organized by 'working paradigm' (gradient-based, single-layer, multi-layer, interpretable, reinforcement-learning, enumeration, evolutionary, in-context learning, LLM-based, human-LLM collaboration, Bayesian optimization), and survey the NLP tasks, pretrained language models, and benchmark datasets on which these methods have been evaluated, culminating in ten performance tables and a call for standardized benchmarking. The paper's central claim is that it provides 'unique and comprehensive insights' and 'a robust foundation for future comparative studies.'
Significance. The survey addresses a timely and genuinely useful need: a structured map of prompt optimization methods together with the tasks, models, and datasets used to evaluate them. Its strengths are the breadth of the compilation (45 methods, ten detailed tables), an explicit and reproducible-looking selection methodology that follows prior review protocols, and a fair critical discussion of non-standardized evaluation practices, particularly the inconsistencies in train/test splits and metrics across studies. The paper also makes falsifiable quantitative claims—method counts, soft/hard prompt counts, and model-usage percentages—which is a credit to its transparency, though several of these claims fail when checked against the paper's own content. If the internal inconsistencies identified below are corrected, the survey would be a serviceable reference; as it stands, the central organizing claims are not reliable.
major comments (4)
- [Section 1; Section 4; Figure 6] The soft/hard prompt count stated in Section 1 ('3 soft prompts and 42 hard prompt based optimization strategies') is contradicted by the paper's own classifications. Section 4.2–4.3 and Figure 6 identify at least ten continuous-prompt methods—Prompt-Tuning, P-Tuning, Prefix-Tuning, P-Tuning v2, Soft Prompt, BBT, BBTv2, FedBPT, DEPT, and LoPT—and Section 4.11 adds InstructZero's soft-prompt refinement stage. Because the soft/hard dichotomy is the paper's own stated organizing frame for the 45 methods, the enumeration that anchors the entire survey is inaccurate and must be recomputed and reconciled with Figure 6.
- [Section 7; Figure 16] The distribution of methods by number of pretrained language models used does not sum to 100%: the text reports 21% (one PLM) + 34% (two) + 34% (four) + 7% (three) + 8% (five) + 18% (six to seventeen) = 122%. These percentages cannot all be correct for a partition of the 45 methods, so the quantitative claims about model-usage diversity in Section 7 are not supported and need to be recomputed from the underlying data.
- [Section 5; Section 6] The task taxonomy is not internally consistent. Section 5 opens by promising 'nine key NLP tasks' and includes Information Retrieval (5.7), but Section 6—which opens by also saying '9 different classes of NLP Tasks'—contains ten subsections, replaces Information Retrieval with Semantic Parsing (6.5) and Knowledge and Contextual Understanding (6.9), and never returns to IR. The claimed comprehensive task overview must be reconciled so that the tasks enumerated in Section 5 are exactly those analyzed in Section 6.
- [Section 4.10; Section 4.11; Figure 6] The central taxonomy is not a well-defined partition. BPO is introduced in Section 4.10 as the 'sole approach' of Human LLM Collaboration, yet by the criteria given in Section 4.11 (methods that employ Bayesian optimization over candidate prompts with LLM-provided feedback) BPO belongs equally to the Bayesian Optimization class; InstructZero likewise spans soft-prompt tuning and Bayesian optimization; and AutoPrompt is simultaneously gradient-based (Section 4.1) and a hard-prompt method. The paper never states whether the 11 classes are mutually exclusive or orthogonal facets, nor which axis is primary for assignment, so a reader cannot assign a new method to exactly one class and the claim of '11 distinct classes' is undefined.
minor comments (8)
- [Section 4.7] The paragraph on PROMPTBREEDER contains a leftover drafting artifact—'Here's a comprehensive rephrasing of your statement, followed by five candidate sentences:'—which must be deleted.
- [Section 2.2] The screening arithmetic is off: 232 articles pass title/abstract screening, but 185 discarded plus 45 kept equals 230, leaving two articles unaccounted for.
- [Section 7; Figure 16] Section 7 refers to 'Table 16' where the referenced object is Figure 16; the decoder-based model list contains the fragment '-tuning'; and Figure 16 labels 'OPRP' where the text consistently uses OPRO.
- [Section 4; Tables 1-10] Method names are used inconsistently throughout (MOP vs. MoP, PROMPTBREEDER vs. Prompt BREEDER, STablePRompt in Table 9, Doscrete-v in Figure 6), and several methods appearing in the tables are never cited at first mention or at all (e.g., Soft Prompt, Active Examples, AEO, ABO before [60], LongPrompt, and PROPANE in Figure 16).
- [Table 1] The disambiguation-qa rows for MoP and AEO are merged into a single entry with one accuracy value and one model column, so it is unclear which method the reported accuracy of 68 refers to.
- [Section 5.1; Section 2.1; Section 6.1] Numerous typos and grammatical slips remain, including 'multilable' (5.1), 'doamin' (Section 2 sample query), 'aaccuracy' (Section 6.1), 'Diaglouge' (Table 5), and 'To best of our knowledge' in the Abstract and Section 1.
- [Figure 3] The search engine list in Figure 3 includes both 'Elsevier' and 'ScienceDirect' as separate items although ScienceDirect is Elsevier's platform, and the text names 'the archive' without specifying arXiv.
- [Section 4; Figure 6] The order of classes in Figure 6 (Single Layer first) differs from the order in Section 4 (Gradient based first), and the figure's method-to-column assignments for FluentPrompt and the reinforcement-learning group are not clearly consistent with Sections 4.1 and 4.5; the figure and the text should be carefully aligned and checked against the numbered categories.
Circularity Check
No circularity: the taxonomy is an aggregation of published methods, self-citations are only domain examples, and no fitted input is relabeled as a prediction.
full rationale
This is a survey paper, not a derivation-based study. Its central contribution is a taxonomy of 45 published prompt-optimization strategies grouped into 11 classes, plus an inventory of tasks, models, and datasets reported by those methods. There is no equation, fitted parameter, or statistical model whose output is recycled as an input, so the classic circularity patterns (self-definitional fits, fitted-input-called-prediction, ansatz-smuggled-via-citation, uniqueness-imported-from-authors) do not apply. The paper cites the authors' own prior work [7,8] in the Introduction merely as examples of NLP application areas for LLMs, and these citations play no load-bearing role in the classification or in any comparative conclusion. The classification itself is based on external, published method descriptions, not on the authors' previous results. The reader-flagged concerns about overlapping categories and inconsistent counts (e.g., BPO appearing in two classes, AutoPrompt being called gradient-based while also being a discrete hard-prompt method, and the '3 soft / 42 hard' count contradicting later listings) are internal-consistency or correctness risks rather than circularity: the taxonomy could be flawed without any argument reducing to its own inputs. Under the requirement that circularity be demonstrated by a specific reduction, no such reduction is present here.
Assumptions & free parameters
assumptions (2)
- domain assumption The 45 selected articles constitute a representative and comprehensive sample of prompt optimization research.
- ad hoc to paper The 11 proposed classes are mutually exclusive and jointly exhaustive.
Cite this review
Pith. "Pith review of The Evolution of Natural Language Processing: How Prompt Optimization and Language Models are Shaping the Future." pith.science (2026). https://pith.science/paper/AGIVYZTU
@misc{pith2026250617700,
author = {Pith},
title = {Pith review of: The Evolution of Natural Language Processing: How Prompt Optimization and Language Models are Shaping the Future},
year = {2026},
howpublished = {\url{https://pith.science/paper/AGIVYZTU}},
note = {Machine review of arXiv:2506.17700}
}
read the original abstract
Large Language Models (LLMs) have revolutionized the field of Natural Language Processing (NLP) by automating traditional labor-intensive tasks and consequently accelerated the development of computer-aided applications. As researchers continue to advance this field with the introduction of novel language models and more efficient training/finetuning methodologies, the idea of prompt engineering and subsequent optimization strategies with LLMs has emerged as a particularly impactful trend to yield a substantial performance boost across diverse NLP tasks. To best of our knowledge numerous review articles have explored prompt engineering, however, a critical gap exists in comprehensive analyses of prompt optimization strategies. To bridge this gap this paper provides unique and comprehensive insights about the potential of diverse prompt optimization strategies. It analyzes their underlying working paradigms and based on these principles, categorizes them into 11 distinct classes. Moreover, the paper provides details about various NLP tasks where these prompt optimization strategies have been employed, along with details of different LLMs and benchmark datasets used for evaluation. This comprehensive compilation lays a robust foundation for future comparative studies and enables rigorous assessment of prompt optimization and LLM-based predictive pipelines under consistent experimental settings: a critical need in the current landscape. Ultimately, this research will centralize diverse strategic knowledge to facilitate the adaptation of existing prompt optimization strategies for development of innovative predictors across unexplored tasks.
Reference graph
Works this paper leans on
-
[1]
IEEE Transactions on Software Engineering (2024)
Fakhoury, S., Naik, A., Sakkas, G., Chakraborty, S., Lahiri, S.K.: Llm-based test- driven interactive code generation: User study and empirical evaluation. IEEE Transactions on Software Engineering (2024)
2024
-
[2]
ACM Transactions on Software Engineering and Methodology (2024)
Huang, D., Zhang, J.M., Bu, Q., Xie, X., Chen, J., Cui, H.: Bias testing and miti- gation in llm-based code generation. ACM Transactions on Software Engineering and Methodology (2024)
2024
-
[3]
arXiv preprint arXiv:2404.13813 (2024)
Enis, M., Hopkins, M.: From llm to nmt: Advancing low-resource machine translation with claude. arXiv preprint arXiv:2404.13813 (2024)
arXiv 2024
-
[4]
In: Proceedings of the Ninth Conference on Machine Translation, pp
Elshin, D., Karpachev, N., Gruzdev, B., Golovanov, I., Ivanov, G., Antonov, A., Skachkov, N., Latypova, E., Layner, V., Enikeeva, E.,et al.: From general llm to translation: How we dramatically improve translation quality using human evaluation data for llm finetuning. In: Proceedings of the Ninth Conference on Machine Translation, pp. 247–252 (2024)
2024
-
[5]
arXiv preprint arXiv:2505.07888 (2025)
Wu, Y., Deng, X.: Implementing long text style transfer with llms through dual-layered sentence and paragraph structure extraction and mapping. arXiv preprint arXiv:2505.07888 (2025)
arXiv 2025
-
[6]
Plos one17(7), 0270275 (2022)
Asim, M.N., Ibrahim, M.A., Malik, M.I., Dengel, A., Ahmed, S.: Lgca-vhppi: A local-global residue context aware viral-host protein-protein interaction predictor. Plos one17(7), 0270275 (2022)
2022
-
[7]
Complex & Intelligent Systems11(6), 1–22 (2025)
Saleem, S., Asim, M.N., Elst, L.V., Dengel, A.: Generative language mod- els potential for requirement engineering applications: insights into current strengths and limitations. Complex & Intelligent Systems11(6), 1–22 (2025)
2025
-
[8]
PassionNet: An Innovative Framework for Duplicate and Conflicting Requirements Identification
Saleem, S., Asim, M.N., Dengel, A.: Passionnet: An innovative frame- work for duplicate and conflicting requirements identification. arXiv preprint arXiv:2412.01657 (2024)
work page Pith review arXiv 2024
Show all 149 references
-
[9]
IEEE access 78 12, 26839–26874 (2024)
Raiaan, M.A.K., Mukta, M.S.H., Fatema, K., Fahad, N.M., Sakib, S., Mim, M.M.J., Ahmad, J., Ali, M.E., Azam, S.: A review on large language models: Architectures, applications, taxonomies, open issues and challenges. IEEE access 78 12, 26839–26874 (2024)
2024
-
[10]
arXiv preprint arXiv:1810.04805 (2018)
Devlin, J.: Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[11]
https://developer.nvidia.com/blog/ using-deepspeed-and-megatron-to-train-megatron-turing-nlg-530b-the-worlds-largest-and-most-powerful-generative-language-model/
NVIDIA Developer: Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, the World’s Largest and Most Power- ful Generative Language Model. https://developer.nvidia.com/blog/ using-deepspeed-and-megatron-to-train-megatron-turing-nlg-530b-the-worlds-largest-and-most-po...
2021
-
[12]
Artificial Intelligence in Medicine157, 103003 (2024)
Bonfigli, A., Bacco, L., Merone, M., Dell’Orletta, F.: From pre-training to fine- tuning: An in-depth analysis of large language models in the biomedical domain. Artificial Intelligence in Medicine157, 103003 (2024)
2024
-
[13]
https://www.forbes.com/councils/forbestechcouncil/2024/ 04/26/the-untold-story-of-ais-huge-carbon-footprint/ (2024)
Council, F.T.: The Untold Story Of AI’s Huge Carbon Foot- print. https://www.forbes.com/councils/forbestechcouncil/2024/ 04/26/the-untold-story-of-ais-huge-carbon-footprint/ (2024). https://www.forbes.com/councils/forbestechcouncil/2024/04/26/ the-untold-story-of-ais-huge-carb...
2024
-
[14]
arXiv preprint arXiv:2310.14735 (2023)
Chen, B., Zhang, Z., Langren´ e, N., Zhu, S.: Unleashing the potential of prompt engineering in large language models: a comprehensive review. arXiv preprint arXiv:2310.14735 (2023)
2023 arXiv
-
[15]
arXiv preprint arXiv:2104.06599 (2021)
Qin, G., Eisner, J.: Learning how to ask: Querying lms with mixtures of soft prompts. arXiv preprint arXiv:2104.06599 (2021)
2021 arXiv
-
[16]
Advances in Neural Information Processing Systems36, 51008–51025 (2023)
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., Goldstein, T.: Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. Advances in Neural Information Processing Systems36, 51008–51025 (2023)
2023
-
[17]
Journal of Computer Languages70, 101117 (2022)
Dalibor, M., Heithoff, M., Michael, J., Netz, L., Pfeiffer, J., Rumpe, B., Varga, S., Wortmann, A.: Generating customized low-code development platforms for digital twins. Journal of Computer Languages70, 101117 (2022)
2022
-
[18]
Di Ruscio, D., Kolovos, D., Lara, J., Pierantonio, A., Tisi, M., Wimmer, M.: Low-code development and model-driven engineering: Two sides of the same coin? Software and Systems Modeling21(2), 437–446 (2022)
2022
-
[19]
arXiv preprint arXiv:2010.15980 (2020)
Shin, T., Razeghi, Y., Logan IV, R.L., Wallace, E., Singh, S.: Autoprompt: Elic- iting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980 (2020)
2020 arXiv
-
[20]
Shi, W., Han, X., Gonen, H., Holtzman, A., Tsvetkov, Y., Zettlemoyer, L.: Toward human readable prompt tuning: Kubrick’s the shining is a good movie, 79 and a good prompt too? arXiv preprint arXiv:2212.10539 (2022)
2022 arXiv
-
[21]
arXiv preprint arXiv:2104.08691 (2021)
Lester, B., Al-Rfou, R., Constant, N.: The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691 (2021)
2021 arXiv
-
[22]
AI Open5, 208–215 (2024)
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., Tang, J.: Gpt understands, too. AI Open5, 208–215 (2024)
2024
-
[23]
In: International Conference on Machine Learning, pp
Sun, T., Shao, Y., Qian, H., Huang, X., Qiu, X.: Black-box tuning for language- model-as-a-service. In: International Conference on Machine Learning, pp. 20841–20855 (2022). PMLR
2022
-
[24]
arXiv preprint arXiv:2310.01467 (2023)
Sun, J., Xu, Z., Yin, H., Yang, D., Xu, D., Chen, Y., Roth, H.R.: Fedbpt: Efficient federated black-box prompt tuning for large language models. arXiv preprint arXiv:2310.01467 (2023)
2023 arXiv
-
[25]
arXiv preprint arXiv:2309.05173 (2023)
Shi, Z., Lipani, A.: Dept: Decomposed prompt tuning for parameter-efficient fine-tuning. arXiv preprint arXiv:2309.05173 (2023)
2023 arXiv
-
[26]
arXiv preprint arXiv:2406.19486 (2024)
Guo, S., Damani, S., Chang, K.-h.: Lopt: Low-rank prompt tuning for parameter efficient language models. arXiv preprint arXiv:2406.19486 (2024)
2024 arXiv
-
[27]
arXiv preprint arXiv:2101.00190 (2021)
Li, X.L., Liang, P.: Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190 (2021)
2021 arXiv
-
[28]
arXiv preprint arXiv:2110.07602 (2021)
Liu, X., Ji, K., Fu, Y., Tam, W.L., Du, Z., Yang, Z., Tang, J.: P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602 (2021)
2021 arXiv
-
[29]
arXiv preprint arXiv:2205.11200 (2022)
Sun, T., He, Z., Qian, H., Zhou, Y., Huang, X., Qiu, X.: Bbtv2: Towards a gradient-free future with large language models. arXiv preprint arXiv:2205.11200 (2022)
2022 arXiv
-
[30]
In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Technologies, pp
Khashabi, D., Lyu, X., Min, S., Qin, L., Richardson, K., Welleck, S., Hajishirzi, H., Khot, T., Sabharwal, A., Singh, S., Choi, Y.: Prompt waywardness: The curi- ous case of discretized interpretation of continuous prompts. In: Proceedings of the 2022 Conference of the North A...
2022 doi
-
[31]
arXiv preprint arXiv:2205.12548 (2022) 80
Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E.P., Hu, Z.: Rlprompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548 (2022) 80
2022 arXiv
-
[32]
Transactions on Machine Learning Research2023(2023)
Diao, S., Huang, Z., Xu, R., Li, X., Lin, Y., Zhou, X., Zhang, T.: Black- box prompt learning for pre-trained language models. Transactions on Machine Learning Research2023(2023)
2023
-
[33]
arXiv preprint arXiv:2211.11890 (2022)
Zhang, T., Wang, X., Zhou, D., Schuurmans, D., Gonzalez, J.E.: Tempera: Test- time prompting via reinforcement learning. arXiv preprint arXiv:2211.11890 (2022)
2022 arXiv
-
[34]
arXiv preprint arXiv:2407.04118 (2024)
Chen, Y., Wen, Z., Fan, G., Chen, Z., Wu, W., Liu, D., Li, Z., Liu, B., Xiao, Y.: Mapo: Boosting large language model performance with model-adaptive prompt optimization. arXiv preprint arXiv:2407.04118 (2024)
2024 arXiv
-
[35]
arXiv preprint arXiv:2401.08189 (2024)
Kong, W., Hombaiah, S.A., Zhang, M., Mei, Q., Bendersky, M.: Prewrite: Prompt rewriting with reinforcement learning. arXiv preprint arXiv:2401.08189 (2024)
2024 arXiv
-
[36]
arXiv preprint arXiv:2410.07652 (2024)
Kwon, M., Kim, G., Kim, J., Lee, H., Kim, J.: Stableprompt: Automatic prompt tuning using reinforcement learning for large language models. arXiv preprint arXiv:2410.07652 (2024)
2024 arXiv
-
[37]
arXiv preprint arXiv:2309.06553 (2023)
Sun, H., H¨ uy¨ uk, A., Schaar, M.: Query-dependent prompt evaluation and optimization with offline inverse rl. arXiv preprint arXiv:2309.06553 (2023)
2023 arXiv
-
[38]
arXiv preprint arXiv:2310.16427 (2023)
Wang, X., Li, C., Wang, Z., Bai, F., Luo, H., Zhang, J., Jojic, N., Xing, E.P., Hu, Z.: Promptagent: Strategic planning with language models enables expert-level prompt optimization. arXiv preprint arXiv:2310.16427 (2023)
2023 arXiv
-
[39]
arXiv preprint arXiv:2203.07281 (2022)
Prasad, A., Hase, P., Zhou, X., Bansal, M.: Grips: Gradient-free, edit- based instruction search for prompting large language models. arXiv preprint arXiv:2203.07281 (2022)
2022 arXiv
-
[40]
arXiv preprint arXiv:2310.12774 (2023)
Zhou, H., Wan, X., Vuli´ c, I., Korhonen, A.: Survival of the most influential prompts: Efficient black-box prompt search via clustering and pruning. arXiv preprint arXiv:2310.12774 (2023)
2023 arXiv
-
[41]
arXiv preprint arXiv:2309.16797 (2023)
Fernando, C., Banarse, D., Michalewski, H., Osindero, S., Rockt¨ aschel, T.: Promptbreeder: Self-referential self-improvement via prompt evolution. arXiv preprint arXiv:2309.16797 (2023)
2023 arXiv
-
[42]
Guo12, Q., Wang, R., Guo, J., Li23, B., Song, K., Tan, X., Liu, G., Bian, J., Yang, Y.: Connecting large language models with evo-lutionary algorithms yields powerful prompt optimizers
-
[43]
In: 2024 IEEE Congress on Evolutionary Computation (CEC), pp
Liu, S., Chen, C., Qu, X., Tang, K., Ong, Y.-S.: Large language models as evolutionary optimizers. In: 2024 IEEE Congress on Evolutionary Computation (CEC), pp. 1–8 (2024). IEEE 81
2024
-
[44]
arXiv preprint arXiv:2112.08633 (2021)
Rubin, O., Herzig, J., Berant, J.: Learning to retrieve prompts for in-context learning. arXiv preprint arXiv:2112.08633 (2021)
2021 arXiv
-
[45]
arXiv preprint arXiv:2211.04486 (2022)
Zhang, Y., Feng, S., Tan, C.: Active example selection for in-context learning. arXiv preprint arXiv:2211.04486 (2022)
2022 arXiv
-
[46]
arXiv preprint arXiv:2210.03493 (2022)
Zhang, Z., Zhang, A., Li, M., Smola, A.: Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493 (2022)
2022 arXiv
-
[47]
arXiv preprint arXiv:2302.12822 (2023)
Shum, K., Diao, S., Zhang, T.: Automatic prompt augmentation and selec- tion with chain-of-thought from labeled data. arXiv preprint arXiv:2302.12822 (2023)
2023 arXiv
-
[48]
In: International Conference on Machine Learning, pp
Hou, B., O’connor, J., Andreas, J., Chang, S., Zhang, Y.: Promptboost- ing: Black-box text classification with ten forward passes. In: International Conference on Machine Learning, pp. 13309–13324 (2023). PMLR
2023
-
[49]
Wang, R., An, S., Cheng, M., Zhou, T., Hwang, S.J., Hsieh, C.-J.: Mixture-of- experts in prompt optimization (2023)
2023
-
[50]
arXiv preprint arXiv:2405.16122 (2024)
Wu, Z., Lin, X., Dai, Z., Hu, W., Shu, Y., Ng, S.-K., Jaillet, P., Low, B.K.H.: Prompt optimization with ease? efficient ordering-aware automated selection of exemplars. arXiv preprint arXiv:2405.16122 (2024)
2024 arXiv
-
[51]
In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp
Long, D., Zhao, Y., Brown, H., Xie, Y., Zhao, J., Chen, N., Kawaguchi, K., Shieh, M., He, J.: Prompt optimization via adversarial in-context learning. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 7308–...
2024
-
[52]
arXiv preprint arXiv:2305.03495 (2023)
Pryzant, R., Iter, D., Li, J., Lee, Y.T., Zhu, C., Zeng, M.: Automatic prompt optimization with” gradient descent” and beam search. arXiv preprint arXiv:2305.03495 (2023)
2023 arXiv
-
[53]
arXiv preprint arXiv:2309.03409 (2023)
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q.V., Zhou, D., Chen, X.: Large language models as optimizers. arXiv preprint arXiv:2309.03409 (2023)
2023 arXiv
-
[54]
arXiv preprint arXiv:2311.05661 (2023)
Ye, Q., Axmed, M., Pryzant, R., Khani, F.: Prompt engineering a prompt engineer. arXiv preprint arXiv:2311.05661 (2023)
2023 arXiv
-
[55]
arXiv preprint arXiv:2311.09569 (2023)
Lu, Y., Wang, J., Tang, R., Riedel, S., Stenetorp, P.: Strings from the library of babel: Random sampling as a strong baseline for prompt optimisation. arXiv preprint arXiv:2311.09569 (2023)
2023 arXiv
-
[56]
arXiv preprint arXiv:2311.04155 (2023)
Cheng, J., Liu, X., Zheng, K., Ke, P., Wang, H., Dong, Y., Tang, J., Huang, M.: Black-box prompt optimization: Aligning large language models without model training. arXiv preprint arXiv:2311.04155 (2023)
2023 arXiv
-
[57]
Mathematics12(6), 929 (2024)
Sabbatella, A., Ponti, A., Giordani, I., Candelieri, A., Archetti, F.: Prompt 82 optimization in large language models. Mathematics12(6), 929 (2024)
2024
-
[58]
arXiv preprint arXiv:2306.03082 (2023)
Chen, L., Chen, J., Goldstein, T., Huang, H., Zhou, T.: Instructzero: Efficient instruction optimization for black-box large language models. arXiv preprint arXiv:2306.03082 (2023)
2023 arXiv
-
[59]
In: The Eleventh International Conference on Learning Representations (2022)
Zhou, Y., Muresanu, A.I., Han, Z., Paster, K., Pitis, S., Chan, H., Ba, J.: Large language models are human-level prompt engineers. In: The Eleventh International Conference on Learning Representations (2022)
2022
-
[60]
Ma, R., Wang, X., Zhou, X., Li, J., Du, N., Gui, T., Zhang, Q., Huang, X.: Are large language models good prompt optimizers? arXiv preprint arXiv:2402.02101 (2024)
2024 arXiv
-
[61]
arXiv preprint arXiv:1705.00648 (2017)
Wang, W.Y.: ” liar, liar pants on fire”: A new benchmark dataset for fake news detection. arXiv preprint arXiv:1705.00648 (2017)
2017 arXiv
-
[62]
In: Thirteenth International Conference on the Principles of Knowledge Represen- tation and Reasoning (2012)
Levesque, H., Davis, E., Morgenstern, L.: The winograd schema challenge. In: Thirteenth International Conference on the Principles of Knowledge Represen- tation and Reasoning (2012)
2012
-
[63]
Communications of the ACM64(9), 99–106 (2021)
Sakaguchi, K., Bras, R.L., Bhagavatula, C., Choi, Y.: Winogrande: An adver- sarial winograd schema challenge at scale. Communications of the ACM64(9), 99–106 (2021)
2021
-
[64]
Accessed: 2025-02-11 (2022)
Suzgun, M., et al.: BIG-Bench Hard (BBH). Accessed: 2025-02-11 (2022). https: //github.com/suzgunmirac/BIG-Bench-Hard
2022
-
[65]
Complex & Intelligent Systems8(6), 4663–4678 (2022)
Mollas, I., Chrysopoulou, Z., Karlos, S., Tsoumakas, G.: Ethos: a multi-label hate speech detection dataset. Complex & Intelligent Systems8(6), 4663–4678 (2022)
2022
-
[66]
arXiv preprint cs/0409058 (2004)
Pang, B., Lee, L.: A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. arXiv preprint cs/0409058 (2004)
2004 arXiv
-
[67]
In: Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pp
Farha, I.A., Magdy, W.: From arabic sentiment analysis to sarcasm detection: The arsarcasm dataset. In: Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pp. 32–39 (2020)
2020
-
[68]
In: Proceed- ings of the 13th International Workshop on Semantic Evaluation, pp
Kiesel, J., Mestre, M., Shukla, R., Vincent, E., Adineh, P., Corney, D., Stein, B., Potthast, M.: Semeval-2019 task 4: Hyperpartisan news detection. In: Proceed- ings of the 13th International Workshop on Semantic Evaluation, pp. 829–839 (2019)
2019
-
[69]
arXiv preprint arXiv:1805.12471 (2019) 83
Warstadt, A.: Neural network acceptability judgments. arXiv preprint arXiv:1805.12471 (2019) 83
2019 arXiv
-
[70]
arXiv preprint arXiv:1808.09121 (2018)
Pilehvar, M.T., Camacho-Collados, J.: Wic: the word-in-context dataset for evaluating context-sensitive meaning representations. arXiv preprint arXiv:1808.09121 (2018)
2018 arXiv
-
[71]
arXiv preprint arXiv:2205.10782 (2022)
Honovich, O., Shaham, U., Bowman, S.R., Levy, O.: Instruction induction: From few examples to natural language task descriptions. arXiv preprint arXiv:2205.10782 (2022)
2022 arXiv
-
[72]
arXiv preprint arXiv:1710.06071 (2017)
Dernoncourt, F., Lee, J.Y.: Pubmed 200k rct: a dataset for sequential sentence classification in medical abstracts. arXiv preprint arXiv:1710.06071 (2017)
2017 arXiv
-
[73]
Advances in neural information processing systems28(2015)
Zhang, X., Zhao, J., LeCun, Y.: Character-level convolutional networks for text classification. Advances in neural information processing systems28(2015)
2015
-
[74]
In: Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pp
Voorhees, E.M., Tice, D.M.: Building a question answering test collection. In: Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 200–207 (2000)
2000
-
[75]
Semantic web6(2), 167–195 (2015)
Lehmann, J., Isele, R., Jakob, M., Jentzsch, A., Kontokostas, D., Mendes, P.N., Hellmann, S., Morsey, M., Van Kleef, P., Auer, S.,et al.: Dbpedia–a large- scale, multilingual knowledge base extracted from wikipedia. Semantic web6(2), 167–195 (2015)
2015
-
[76]
In: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C.D., Ng, A.Y., Potts, C.: Recursive deep models for semantic compositionality over a sentiment tree- bank. In: Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp. 1631–1642 (2013)
2013
-
[77]
arXiv preprint cs/0506075 (2005)
Pang, B., Lee, L.: Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. arXiv preprint cs/0506075 (2005)
2005 arXiv
-
[78]
In: Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp
Hu, M., Liu, B.: Mining and summarizing customer reviews. In: Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 168–177 (2004)
2004
-
[79]
In: Proceedings of the 7th ACM Conference on Recommender Systems, pp
McAuley, J., Leskovec, J.: Hidden factors and hidden topics: understanding rat- ing dimensions with review text. In: Proceedings of the 7th ACM Conference on Recommender Systems, pp. 165–172 (2013)
2013
-
[80]
arXiv preprint arXiv:1810.12885 (2018)
Zhang, S., Liu, X., Liu, J., Gao, J., Duh, K., Van Durme, B.: Record: Bridg- ing the gap between human and machine commonsense reading comprehension. arXiv preprint arXiv:1810.12885 (2018)
2018 arXiv
-
[81]
arXiv preprint arXiv:1606.05250 (2016)
Rajpurkar, P.: Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250 (2016)
2016 arXiv
-
[82]
Transactions of the Association for Computational Linguistics7, 453–466 (2019)
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, 84 C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K.,et al.: Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics7, 453–466 (2019)
2019
-
[83]
arXiv preprint arXiv:1809.09600 (2018)
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W.W., Salakhutdinov, R., Man- ning, C.D.: Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600 (2018)
2018 arXiv
-
[84]
Dunn, M., Sagun, L., Higgins, M., Guney, V.U., Cirik, V., Cho, K.: Searchqa: A new q&a dataset augmented with context from a search engine. arxiv. arXiv preprint cs.CL/1704.05179 (2017)
2017 arXiv
-
[85]
arXiv preprint arXiv:1611.09830 (2016)
Trischler, A., Wang, T., Yuan, X., Harris, J., Sordoni, A., Bachman, P., Suleman, K.: Newsqa: A machine comprehension dataset. arXiv preprint arXiv:1611.09830 (2016)
2016 arXiv
-
[86]
arXiv preprint arXiv:1905.10044 (2019)
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., Toutanova, K.: Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044 (2019)
2019 arXiv
-
[87]
In: 2011 AAAI Spring Symposium Series (2011)
Roemmele, M., Bejan, C.A., Gordon, A.S.: Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In: 2011 AAAI Spring Symposium Series (2011)
2011
-
[88]
In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp
Khashabi, D., Chaturvedi, S., Roth, M., Upadhyay, S., Roth, D.: Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics:...
2018
-
[89]
arXiv preprint arXiv:1804.07927 (2018)
Saha, A., Aralikatte, R., Khapra, M.M., Sankaranarayanan, K.: Duorc: Towards complex language understanding with paraphrased reading comprehension. arXiv preprint arXiv:1804.07927 (2018)
2018 arXiv
-
[90]
arXiv preprint arXiv:1903.00161 (2019)
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., Gardner, M.: Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. arXiv preprint arXiv:1903.00161 (2019)
2019 arXiv
-
[91]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp
Kembhavi, A., Seo, M., Schwenk, D., Choi, J., Farhadi, A., Hajishirzi, H.: Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4999–5007 (2017)
2017
-
[92]
https://www.bioasq.org/
BioASQ: A Challenge on Large-Scale Biomedical Semantic Indexing and Ques- tion Answering. https://www.bioasq.org/. Accessed: 2025-02-10
2025
-
[93]
arXiv preprint arXiv:1704.04683 (2017)
Lai, G., Xie, Q., Liu, H., Yang, Y., Hovy, E.: Race: Large-scale reading 85 comprehension dataset from examinations. arXiv preprint arXiv:1704.04683 (2017)
2017 arXiv
-
[94]
arXiv preprint arXiv:1809.02789 (2018)
Mihaylov, T., Clark, P., Khot, T., Sabharwal, A.: Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789 (2018)
2018 arXiv
-
[95]
arXiv preprint arXiv:2109.07958 (2021)
Lin, S., Hilton, J., Evans, O.: Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958 (2021)
2021 arXiv
-
[96]
Applied Sciences11(14), 6421 (2021)
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., Szolovits, P.: What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences11(14), 6421 (2021)
2021
-
[97]
arXiv preprint arXiv:1704.05426 (2017)
Williams, A., Nangia, N., Bowman, S.R.: A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426 (2017)
2017 arXiv
-
[98]
arXiv preprint arXiv: 180407461 (2018)
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S.: Glue: A multi- task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv: 180407461 (2018)
2018
-
[99]
In: Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, pp
Giampiccolo, D., Magnini, B., Dagan, I., Dolan, W.B.: The third pascal rec- ognizing textual entailment challenge. In: Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, pp. 1–9 (2007)
2007
-
[100]
Blunsom, P., Camburu, O.-M., Lukasiewicz, T., Rockt¨ aschel, T.: e- snli: Natural language inference with natural language explanations (2018)
2018
-
[101]
In: Proceedings of Sinn und Bedeutung, vol
De Marneffe, M.-C., Simons, M., Tonhauser, J.: The commitmentbank: Investi- gating projection in naturally occurring discourse. In: Proceedings of Sinn und Bedeutung, vol. 23, pp. 107–124 (2019)
2019
-
[102]
In: Proceedings of the AAAI Conference on Artificial Intelligence, vol
Khot, T., Sabharwal, A., Clark, P.: Scitail: A textual entailment dataset from science question answering. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32 (2018)
2018
-
[103]
arXiv preprint arXiv:1508.05326 (2015)
Bowman, S.R., Angeli, G., Potts, C., Manning, C.D.: A large annotated corpus for learning natural language inference. arXiv preprint arXiv:1508.05326 (2015)
2015 arXiv
-
[104]
In: Pro- ceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), pp
Marelli, M., Bentivogli, L., Baroni, M., Bernardi, R., Menini, S., Zamparelli, R.: Semeval-2014 task 1: Evaluation of compositional distributional semantic models on full sentences through semantic relatedness and textual entailment. In: Pro- ceedings of the 8th International ...
2014
-
[105]
arXiv preprint arXiv:1809.05053 (2018)
Conneau, A., Lample, G., Rinott, R., Williams, A., Bowman, S.R., Schwenk, 86 H., Stoyanov, V.: Xnli: Evaluating cross-lingual sentence representations. arXiv preprint arXiv:1809.05053 (2018)
2018 arXiv
-
[106]
arXiv preprint arXiv:1911.12237 (2019)
Gliwa, B., Mochol, I., Biesek, M., Wawer, A.: Samsum corpus: A human- annotated dialogue dataset for abstractive summarization. arXiv preprint arXiv:1911.12237 (2019)
2019 arXiv
-
[107]
arXiv preprint arXiv:1808.08745 (2018)
Narayan, S., Cohen, S.B., Lapata, M.: Don’t give me the details, just the sum- mary! topic-aware convolutional neural networks for extreme summarization. arXiv preprint arXiv:1808.08745 (2018)
2018 arXiv
-
[108]
arXiv preprint arXiv:1602.06023 (2016)
Nallapati, R., Zhou, B., Gulcehre, C., Xiang, B., et al.: Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023 (2016)
2016 arXiv
-
[109]
arXiv preprint arXiv:1706.09254 (2017)
Novikova, J., Duˇ sek, O., Rieser, V.: The e2e dataset: New challenges for end-to- end generation. arXiv preprint arXiv:1706.09254 (2017)
2017 arXiv
-
[110]
In: 10th International Conference on Natural Language Generation, pp
Gardent, C., Shimorina, A., Narayan, S., Perez-Beltrachini, L.: The webnlg challenge: Generating text from rdf data. In: 10th International Conference on Natural Language Generation, pp. 124–133 (2017). ACL Anthology
2017
-
[111]
arXiv preprint arXiv:2007.02871 (2020)
Nan, L., Radev, D., Zhang, R., Rau, A., Sivaprasad, A., Hsieh, C., Tang, X., Vyas, A., Verma, N., Krishna, P., et al.: Dart: Open-domain structured data record to text generation. arXiv preprint arXiv:2007.02871 (2020)
2020 arXiv
-
[112]
arXiv preprint arXiv:2005.00481 (2020)
Alva-Manchego, F., Martin, L., Bordes, A., Scarton, C., Sagot, B., Specia, L.: Asset: A dataset for tuning and evaluation of sentence simplification models with multiple rewriting transformations. arXiv preprint arXiv:2005.00481 (2020)
2020 arXiv
-
[113]
Advances in neural information processing systems30 (2017)
Shen, T., Lei, T., Barzilay, R., Jaakkola, T.: Style transfer from non-parallel text by cross-alignment. Advances in neural information processing systems30 (2017)
2017
-
[114]
In: Proceedings of COLING 2012, pp
Xu, W., Ritter, A., Dolan, W.B., Grishman, R., Cherry, C.: Paraphrasing for style. In: Proceedings of COLING 2012, pp. 2899–2914 (2012)
2012
-
[115]
In: Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1) (2021)
Dumitrescu, S.D., Rebeja, P., Lorincz, B., Gaman, M., Avram, A., Ilie, M., Pruteanu, A., Stan, A., Rosia, L., Iacobescu, C.,et al.: Liro: Benchmark and leaderboard for romanian language tasks. In: Thirty-fifth Conference on Neural Information Processing Systems Datasets and Be...
2021
-
[116]
Qi, Y., Sachan, D.S., Felix, M., Padmanabhan, S.J., Neubig, G.: When and why are pre-trained word embeddings useful for neural machine translation? arXiv preprint arXiv:1804.06323 (2018) 87
2018 arXiv
-
[117]
Transactions of the Association for Computational Linguistics8, 183–198 (2020)
Wolfson, T., Geva, M., Gupta, A., Gardner, M., Goldberg, Y., Deutch, D., Berant, J.: Break it down: A question understanding benchmark. Transactions of the Association for Computational Linguistics8, 183–198 (2020)
2020
-
[118]
arXiv preprint arXiv:2008.09335 (2020)
Li, H., Arora, A., Chen, S., Gupta, A., Gupta, S., Mehdad, Y.: Mtop: A compre- hensive multilingual task-oriented semantic parsing benchmark. arXiv preprint arXiv:2008.09335 (2020)
2020 arXiv
-
[119]
Transactions of the Association for Computational Linguistics8, 556– 571 (2020)
Andreas, J., Bufe, J., Burkett, D., Chen, C., Clausman, J., Crawford, J., Crim, K., DeLoach, J., Dorner, L., Eisner, J.,et al.: Task-oriented dialogue as dataflow synthesis. Transactions of the Association for Computational Linguistics8, 556– 571 (2020)
2020
-
[120]
Bioinformatics33(14), 49–58 (2017)
So˘ gancıo˘ glu, G.,¨Ozt¨ urk, H.,¨Ozg¨ ur, A.: Biosses: a semantic sentence similar- ity estimation system for the biomedical domain. Bioinformatics33(14), 49–58 (2017)
2017
-
[121]
arXiv preprint arXiv:1708.00055 (2017)
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., Specia, L.: Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv preprint arXiv:1708.00055 (2017)
2017 arXiv
-
[122]
In: Third International Workshop on Paraphrasing (IWP2005) (2005)
Dolan, B., Brockett, C.: Automatically constructing a corpus of sentential paraphrases. In: Third International Workshop on Paraphrasing (IWP2005) (2005)
2005
-
[123]
Accessed: 2025-02-11 (2017)
Quora: First Quora Dataset Release: Question Pairs. Accessed: 2025-02-11 (2017). https://quoradata.quora.com/ First-Quora-Dataset-Release-Question-Pairs
2017
-
[124]
arXiv preprint arXiv:1904.01130 (2019)
Zhang, Y., Baldridge, J., He, L.: Paws: Paraphrase adversaries from word scrambling. arXiv preprint arXiv:1904.01130 (2019)
2019 arXiv
-
[125]
In: Proceedings of the 27th International Conference on Computational Linguistics, pp
Liu, X., Chen, Q., Deng, C., Zeng, H., Chen, J., Li, D., Tang, B.: Lcqmc: A large- scale chinese question matching corpus. In: Proceedings of the 27th International Conference on Computational Linguistics, pp. 1952–1962 (2018)
2018
-
[126]
arXiv preprint cs/0306050 (2003)
Sang, E.F., De Meulder, F.: Introduction to the conll-2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050 (2003)
2003 arXiv
-
[127]
In: Proceedings of the Eighth Conference on Computational Natural Language Learning (CoNLL-2004) at HLT-NAACL 2004, pp
Carreras, X., M` arquez, L.: Introduction to the CoNLL-2004 shared task: Seman- tic role labeling. In: Proceedings of the Eighth Conference on Computational Natural Language Learning (CoNLL-2004) at HLT-NAACL 2004, pp. 89–97. Association for Computational Linguistics, Boston, ...
2004
-
[128]
Linguistic Data Consortium, 88 Philadelphia
Weischedel, R., et al.: OntoNotes Release 5.0. Linguistic Data Consortium, 88 Philadelphia. LDC2013T19 (2013)
2013
-
[129]
Journal of biomedical informatics 47, 1–10 (2014)
Do˘ gan, R.I., Leaman, R., Lu, Z.: Ncbi disease corpus: a resource for disease name recognition and concept normalization. Journal of biomedical informatics 47, 1–10 (2014)
2014
-
[130]
In: Joint Conference on EMNLP and CoNLL-shared Task, pp
Pradhan, S., Moschitti, A., Xue, N., Uryupina, O., Zhang, Y.: Conll-2012 shared task: Modeling multilingual unrestricted coreference in ontonotes. In: Joint Conference on EMNLP and CoNLL-shared Task, pp. 1–40 (2012)
2012
-
[131]
In: Proceedings of the Ninth Conference on Computational Natural Language Learning (CoNLL-2005), pp
Carreras, X., M` arquez, L.: Introduction to the conll-2005 shared task: Semantic role labeling. In: Proceedings of the Ninth Conference on Computational Natural Language Learning (CoNLL-2005), pp. 152–164 (2005)
2005
-
[132]
In: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) (2018)
Elsahar, H., Vougiouklis, P., Remaci, A., Gravier, C., Hare, J., Laforest, F., Sim- perl, E.: T-rex: A large scale alignment of natural language with knowledge base triples. In: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 201...
2018
-
[133]
Accessed: 2025-02-11 (2021)
Research, G.: Relation Extraction Corpus. Accessed: 2025-02-11 (2021). https: //github.com/google-research-datasets/relation-extraction-corpus
2021
-
[134]
In: Proceedings of the AAAI Conference on Artificial Intelligence, vol
Speer, R., Chin, J., Havasi, C.: Conceptnet 5.5: An open multilingual graph of general knowledge. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 31 (2017)
2017
-
[135]
arXiv preprint arXiv:1808.09602 (2018)
Luan, Y., He, L., Ostendorf, M., Hajishirzi, H.: Multi-task identification of enti- ties, relations, and coreference for scientific knowledge graph construction. arXiv preprint arXiv:1808.09602 (2018)
2018 arXiv
-
[136]
arXiv preprint arXiv:1706.04115 (2017)
Levy, O., Seo, M., Choi, E., Zettlemoyer, L.: Zero-shot relation extraction via reading comprehension. arXiv preprint arXiv:1706.04115 (2017)
2017 arXiv
-
[137]
Petroni, F., Rockt¨ aschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A.H., Riedel, S.: Language models as knowledge bases? arXiv preprint arXiv:1909.01066 (2019)
2019 arXiv
-
[138]
arXiv preprint arXiv:1811.00937 (2018)
Talmor, A., Herzig, J., Lourie, N., Berant, J.: Commonsenseqa: A ques- tion answering challenge targeting commonsense knowledge. arXiv preprint arXiv:1811.00937 (2018)
2018 arXiv
-
[139]
Transactions of the Association for Computational Linguistics9, 346–361 (2021)
Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., Berant, J.: Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. Transactions of the Association for Computational Linguistics9, 346–361 (2021)
2021
-
[140]
89 Advances in neural information processing systems35, 24824–24837 (2022)
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D.,et al.: Chain-of-thought prompting elicits reasoning in large language models. 89 Advances in neural information processing systems35, 24824–24837 (2022)
2022
-
[141]
arXiv preprint arXiv:2110.14168 (2021)
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al.: Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168 (2021)
2021 arXiv
-
[142]
arXiv preprint arXiv:1608.01413 (2016)
Roy, S., Roth, D.: Solving general arithmetic word problems. arXiv preprint arXiv:1608.01413 (2016)
2016 arXiv
-
[143]
Patel, A., Bhattamishra, S., Goyal, N.: Are nlp models really able to solve simple math word problems? arXiv preprint arXiv:2103.07191 (2021)
2021 arXiv
-
[144]
arXiv preprint arXiv:1705.04146 (2017)
Ling, W., Yogatama, D., Dyer, C., Blunsom, P.: Program induction by ratio- nale generation: Learning to solve and explain algebraic word problems. arXiv preprint arXiv:1705.04146 (2017)
2017 arXiv
-
[145]
In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp
Hosseini, M.J., Hajishirzi, H., Etzioni, O., Kushman, N.: Learning to solve arithmetic word problems with verb categorization. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 523–533 (2014)
2014
-
[146]
Transactions of the Association for Computational Linguistics3, 585–597 (2015)
Koncel-Kedziorski, R., Hajishirzi, H., Sabharwal, A., Etzioni, O., Ang, S.D.: Parsing algebraic word problems into equations. Transactions of the Association for Computational Linguistics3, 585–597 (2015)
2015
-
[147]
arXiv preprint arXiv:2103.03874 (2021)
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., Steinhardt, J.: Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874 (2021)
2021 arXiv
-
[148]
arXiv preprint arXiv:2311.10117 (2023)
Hsieh, C.-J., Si, S., Yu, F.X., Dhillon, I.S.: Automatic engineering of long prompts. arXiv preprint arXiv:2311.10117 (2023)
2023 arXiv
-
[149]
In: Findings of the Association for Computational Linguistics: ACL 2023, pp
Ju, T., Zheng, Y., Wang, H., Zhao, H., Liu, G.: Is continuous prompt a com- bination of discrete prompts? towards a novel view for interpreting continuous prompts. In: Findings of the Association for Computational Linguistics: ACL 2023, pp. 7804–7819 (2023) 90
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.