Pith. sign in

REVIEW 2 major objections 5 minor 277 references

Optimising Language Models for Downstream Tasks: A Post-Training Perspective

T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This thesis argues that post-training techniques—continued pre-training with prompt templates, decomposed prompt tuning, and loss over instructions—substantially improve how language models adapt to downstream tasks.

desk verdict A solid, honest compilation of five peer-reviewed post-training papers; the semi-supervised PCP experiments have a data-overlap problem that should be fixed, and the abstract's AGI line is unsupported. read the letter →

arxiv 2506.20917 v1 pith:5FFK5MUP submitted 2025-06-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagemodelpost-trainingcontinuedpre-trainingprompttuninginstructionparameter-efficientfine-tuningsemi-supervisedlearningmodellingspatialreasoningbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The thesis tries to establish that how a language model is post-trained, after its initial pre-training, determines how well it adapts to specific tasks, and that three simple techniques outperform standard practice. It claims that continuing pre-training on task-related texts formatted with prompt templates and pseudo-labels (PCP) consistently improves prompt-based fine-tuning, sometimes by more than 20% absolute, and beats state-of-the-art semi-supervised methods. It claims that decomposing a soft prompt into a shorter prompt plus low-rank embedding updates (DEPT) matches or beats parameter-efficient fine-tuning baselines while saving over 20% training time and memory, with the savings growing as models get larger. It claims that computing the language-modelling loss over instructions as well as outputs (IM) improves instruction-following and mitigates overfitting, especially when instructions are long relative to outputs or labelled data is scarce. If these claims hold, practitioners gain simple, model-agnostic ways to adapt language models more accurately and cheaply.

What carries the argument

The common mechanism is that post-training should expose the model to task structure, not just task text. PCP carries that structure by embedding a verbalizer-mapped label into a masked template during continued pre-training, so the LM sees the task's input-output format under its own pre-training objective. DEPT carries it by decomposing the soft prompt into a shorter trainable prompt plus a low-rank word-embedding update, preserving capacity while shrinking input length; the use of two different learning rates for the prompt and the low-rank matrices is the empirically load-bearing part of the design. IM carries it by extending the autoregressive loss to instruction tokens, which the thesis argues acts as a regulariser that mitigates overfitting to short outputs or small datasets. The thesis identifies the ratio of instruction length to output length and the number of training examples as the two factors governing when IM's advantage appears.

What would settle it

Re-run the semi-supervised PCP experiments with the unlabelled continued pre-training corpus strictly disjoint from the labelled examples, for instance by holding out a split of the training set, and compare the Macro-F1 gains. If the improvement over prompt-based fine-tuning without PCP drops to near zero, the overlap between labelled and unlabelled data is driving the result; if the gains persist, the overlap is not responsible.

Watch

Extended reading notes

Core claim

On its own terms, the thesis establishes three post-training methods. Prompt-based Continued Pre-training (PCP) replaces the plain task corpus used in task-adaptive pre-training with templated instances in which the masked label position is filled with golden or pseudo labels; the thesis reports that this consistently improves both hard- and soft-prompt fine-tuning on 16 single-sentence and sentence-pair tasks in semi- and fully-supervised settings, outperforming conventional continued pre-training and four state-of-the-art semi-supervised approaches. Decomposed Prompt Tuning (DEPT) factorises a trainable soft prompt of length $l$ into a shorter soft prompt of length $m$ and a pair of low-rank matrices that update frozen word embeddings, trained with two different learning rates; the thesis reports that it outperforms state-of-the-art parameter-efficient fine-tuning on GLUE, SuperGLUE, MRQA, and vision-language tasks while cutting time and memory. Instruction Modelling (IM) computes the next-token loss on instruction tokens as well as completion tokens, excluding template tokens, and the thesis reports that this improves performance across 18 NLP tasks and open-ended generation benchmarks, with AlpacaEval 1.0 win rates rising by over 100% in the most favourable case. A fourth contribution, the StepGame benchmark, is used to show that multi-hop spatial reasoning in text remains a weakness of current language models and to propose a tensor-product memory-augmented architecture for it.

Load-bearing premise

The load-bearing premise is that continued pre-training on the full training set as unlabelled data gives an honest estimate of the semi-supervised gains, even though the labelled examples used for fine-tuning are drawn from that same pool; if this overlap inflates the measured improvements, the reported benefits of PCP over semi-supervised baselines would shrink.

Editorial extensions

If this is right

  • PCP gives practitioners a drop-in replacement for task-adaptive pre-training: continued pre-training on as few as a few hundred templated examples can improve prompt-based fine-tuning, without iterative pseudo-labelling or data augmentation.
  • DEPT makes prompt tuning more deployment-friendly: for the same trainable-parameter budget as vanilla prompt tuning, it reduces input-sequence length and yields faster training and inference, with the speed-up growing with model size.
  • IM provides a simple rule for instruction tuning: computing loss on instructions helps most when instructions are long relative to outputs or when the tuning set is small, so low-resource instruction tuning can be improved without new data.
  • The StepGame benchmark and the tensor-product memory-augmented network indicate that spatial reasoning in text remains a distinct weakness of language models, motivating evaluation of post-training methods beyond standard NLP tasks.
  • The methods are presented as model-agnostic and composable: PCP can initialise semi-supervised backbones, DEPT is orthogonal to LoRA and adapters, and IM can be combined with noise-based regularisation, so the gains stack with existing pipelines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If PCP's gains truly come from presenting task format under the pre-training objective, the same trick should extend to decoder-only and encoder-decoder architectures, which the thesis leaves untested; a concrete next step is PCP with causal or span-corruption objectives.
  • DEPT's efficiency story suggests a broader principle: any PEFT method that adds input tokens pays a quadratic attention cost that can be traded for low-rank input-space updates, so the same decomposition could apply to prefix tuning, in-context example compression, or long-context prompting.
  • IM's instruction-length-to-output-length factor implies that dataset curation for instruction tuning should target this ratio, not just quality or diversity; a practical extension is a curriculum that dynamically reweights instruction loss during training.
  • The semi-supervised overlap concern in Chapter 3 may also affect the reported magnitudes of PCP's advantages; a stricter evaluation with a disjoint unlabelled corpus would clarify whether the headline gains are partially an artefact of the experimental setup.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The thesis proposes and evaluates four post-training techniques for adapting language models to downstream tasks: Prompt-based Continued Pre-training (PCP), which incorporates prompt templates and pseudo/golden labels into continued pre-training; Decomposed Prompt Tuning (DEPT), which replaces a long soft prompt with a shorter prompt plus low-rank updates to word embeddings; Instruction Modelling (IM), which extends the instruction-tuning loss to instruction tokens; and StepGame, a templated multi-hop spatial reasoning benchmark. The thesis claims that these methods improve LM performance in semi-supervised learning, parameter-efficient fine-tuning, low-resource instruction following, and reasoning evaluation, and it backs these claims with experiments on GLUE, SuperGLUE, MRQA, instruction-tuning datasets, and LLM-judged open-ended generation benchmarks, with code publicly released.

Significance. If the claims are correct, PCP offers a simple and stable alternative to iterative self-training, DEPT directly addresses the sequence-length overhead of prompt tuning, and IM gives a low-cost recipe for low-resource instruction tuning; StepGame adds a diagnosable spatial-reasoning probe. Strengths of the manuscript include the breadth of the empirical evaluation, the explicit limitations sections in each chapter, and the open-sourced implementations. The main reservation is that the central semi-supervised claim of Chapter 3 is built on an experimental protocol in which the 'unlabelled' corpus contains the labelled training examples, so the magnitude of the reported gains is not established for a standard semi-supervised setup.

major comments (2)
  1. [§3.4.1, Table 3.4] The semi-supervised protocol draws the labelled examples from the full training set and then uses the full training set as the unlabelled corpus, so every labelled example also appears among the 'unlabelled' data. For PCP, Section 3.3 Step 1 trains a prompt-based classifier on the labelled subset and pseudo-labels the entire unlabelled corpus; for the overlapping examples these pseudo-labels are likely to be correct, and Step 2's MLM objective then provides the model with direct, repeated access to the correct label word for examples that are later used in the prompt-based fine-tuning stage. This is label leakage relative to a standard SSL protocol with a disjoint unlabelled set, and it makes the comparison against the self-training baselines inequitable, since those baselines see the same texts but only PCP converts the overlap into clean supervised signal. No ablation with a disjoint unlabelled set is reported, and the limitations in Section 3.6 do not mention the overlap. Because the inflation is likely largest in the low-label conditions where PCP reports its biggest gains (e.g., IMDB with 20 labelled examples in Table 3.4), the chapter's headline claim that PCP consistently improves prompt-based FT in semi-supervised settings is not established as stated.
  2. [§3.2.1, Table 3.2] The comparison between TAPT and the self-training baselines has the same overlap: TAPT is run on the 'training text corpus including labelled and unlabelled data', and the unlabelled sizes in Figure 3.3 correspond to the full training sets from which the labelled subsets are sampled. Even though TAPT does not use the labels, continued pre-training on the exact text of the labelled examples is not equivalent to training on genuinely unlabelled data, so the claim that TAPT is a strong and stable SSL learner is also weakened. A disjoint-unlabelled version of this comparison should be reported.
minor comments (5)
  1. [§1.3] The name 'Zhengxiang Shi' appears for the author's own papers in the contribution lists, whereas the title page gives 'Zhengyan Shi'; please correct the inconsistency.
  2. [§5.1] In the opening paragraph, 'it does align LMs to act in accordance with the user's intentions' should read 'it does not align' (or 'it fails to align'), as the next sentence describes the need for alignment.
  3. [§3.4.1] The text refers to 'supplementary experiments on the full training sets in Appendix 3.5', but Section 3.5 is in the main text, not an appendix; please renumber or rephrase the reference.
  4. [Table 4.1] DEPT results in Table 4.1 are reported without standard deviations, while some baselines are quoted from prior work; given the small margins over MPT (0.1/0.4 points on mean GLUE/SuperGLUE), please report seed-level variation or explicitly state that the differences are descriptive rather than statistically tested.
  5. [Figure 3.6] The caption on the right-hand panel refers to 'Scaling Laws' but only two model sizes are compared; please soften the wording or add more model sizes to justify the scaling-law interpretation.

Circularity Check

0 steps flagged · score 2.0 of 10

No derivation-level circularity: the thesis methods are evaluated on external benchmarks and are not defined in terms of their own predictions; the Ch. 3 labelled/unlabelled overlap is an experimental-design limitation rather than a circular construction.

full rationale

I walked the claimed derivation chains of the four main contributions. PCP (Ch. 3) is fully specified: Step 1 trains a prompt-based classifier on labelled data, Step 2 pseudo-labels an unlabelled corpus and runs MLM on templated text, and the resulting checkpoint is fine-tuned. The reported gains on GLUE and SSL benchmarks are empirical measurements, not consequences of the definitions. DEPT (Ch. 4) decomposes the soft prompt into a shorter prompt plus low-rank matrices; the parameter-count matching l*d = m*d + (s+d)*r is a deliberate construction, but the efficiency claim follows from the shorter input sequence and the accuracy claims are measured on external GLUE, SuperGLUE, MRQA and VQA benchmarks. IM (Ch. 5) simply changes the loss mask over instruction tokens; the AlpacaEval and NLP-task improvements are externally evaluated, and the overfitting analysis is evidence, not a tautology. StepGame (Ch. 6) is a new benchmark with external bAbI comparisons. The thesis relies heavily on the author's own prior papers, but each method is re-described with full equations and experimental protocols, so the self-citations are not load-bearing in the sense of substituting for evidence. The one substantive concern is in the semi-supervised protocol of Ch. 3: Section 3.4.1 says labelled examples are sampled from the full training set, and Table 3.4 states that the full training set is used as unlabelled data, so the unlabelled corpus contains the labelled examples. This can inflate the reported benefit of PCP, and Section 3.6 does not mention it. However, this is data leakage / experimental validity, not a case where a prediction reduces to its own inputs by construction: the same overlap applies to the ST baselines, and the fully-supervised PCP results in Table 3.3 are independent of the overlap. I therefore find no definitional circularity, and assign a low score reflecting the non-load-bearing self-citation pattern plus the unresolved experimental-design risk.

Assumptions & free parameters 1 free parameters · 2 assumptions · 0 invented entities

The thesis is an empirical compilation, not a derivation of new laws. The main free parameter is the learning-rate split in DEPT, which is essential to its success. The axioms are standard complexity facts and the domain assumption that continued pre-training with prompts helps prompt fine-tuning. No new physical or theoretical entities are introduced.

free parameters (1)
  • Prompt tuning learning rates (alpha_1, alpha_2) = not reported in text; selected by grid search
    In Chapter 4, DEPT requires separate learning rates for the shorter soft prompt and the low-rank matrices. The paper shows that using a single learning rate fails (GLUE performance of 40.8 or 54.7 vs 85.7 with mixed rates). This is a critical hyperparameter choice that the method depends on.
assumptions (2)
  • standard math Transformer self-attention has quadratic complexity in the input sequence length.
    Used in Section 4.1 to motivate DEPT's shorter soft prompt as a way to reduce compute and memory costs.
  • domain assumption Masked language modelling on task-related text with injected labels and templates transfers task knowledge to prompt-based fine-tuning.
    This is the core assumption behind PCP in Chapter 3. It is not proven theoretically, only empirically tested across 16 datasets.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Optimising Language Models for Downstream Tasks: A Post-Training Perspective." pith.science (2026). https://pith.science/paper/5FFK5MUP

@misc{pith2026250620917,
  author       = {Pith},
  title        = {Pith review of: Optimising Language Models for Downstream Tasks: A Post-Training Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5FFK5MUP}},
  note         = {Machine review of arXiv:2506.20917}
}
read the original abstract

Language models (LMs) have demonstrated remarkable capabilities in NLP, yet adapting them efficiently and robustly to specific tasks remains challenging. As their scale and complexity grow, fine-tuning LMs on labelled data often underutilizes available unlabelled data, leads to overfitting on small task-specific sets, and imposes significant computational costs. These limitations hamper their application to the open-ended landscape of real-world language tasks. This thesis proposes a series of methods to better adapt LMs to downstream applications. First, we explore strategies for extracting task-relevant knowledge from unlabelled data, introducing a novel continued pre-training technique that outperforms state-of-the-art semi-supervised approaches. Next, we present a parameter-efficient fine-tuning method that substantially reduces memory and compute costs while maintaining competitive performance. We also introduce improved supervised fine-tuning methods that enable LMs to better follow instructions, especially when labelled data is scarce, enhancing their performance across a range of NLP tasks, including open-ended generation. Finally, we develop new evaluation methods and benchmarks, such as multi-hop spatial reasoning tasks, to assess LM capabilities and adaptation more comprehensively. Through extensive empirical studies across diverse NLP tasks, our results demonstrate that these approaches substantially improve LM robustness, efficiency, and generalization, making them more adaptable to a broad range of applications. These advances mark a significant step towards more robust and efficient LMs, bringing us closer to the goal of artificial general intelligence.

Figures

Figures reproduced from arXiv: 2506.20917 by the authors.

Figure 3.1
Figure 3.1. Mean performance of CLS- and prompt-based FT across 16 NLP tasks when trained by themselves or in combination with either TAPT or our proposed PCP in the semi-supervised setting. Please refer to [PITH_FULL_IMAGE:figures/full_fig_p063_3_1.png] view at source ↗
Figure 3.2
Figure 3.2. The overview of Prompt-based Continued Pre-training (g), in comparison to conventional continued pre-training (e) and instruction tuning (f), along with fine-tuning methods (a,b) and continued pre-training techniques (c,d). The ver￾balizer functions as a mapping from the task label space to individual words. We use masked language modelling for illustrative purposes, where <mask> represents a masked token in the LM … view at source ↗
Figure 3.3
Figure 3.3. The effect of labelled size on TAPT and ST. Average test Macro-F1 score over 5 seeds is reported. From the left to the right, TAPT and ST utilizes 23k, 60k, 100k, 250k, and 500k unlabelled samples respectively. 3.2.2 Empirical Results: Self-Training vs Task Adaptive Pre￾training Overview [PITH_FULL_IMAGE:figures/full_fig_p068_3_3.png] view at source ↗
Figures from the paper (23 more)
Figure 3.4
Figure 3.4. Figure 3.4: The impact of labelled size on the F1 improvement from TAPT over SUPER￾VISED, where unlabelled size is fixed for each dataset. The red vertical line highlights the FULLY-SUPERVISED setting on which prior work [76] focuses. tent advantage over ADAMATCH and FLEXMATCH a…
Figure 3.5
Figure 3.5. Figure 3.5: The performance lower bound of the PCP, where “wrong labels” indicates that all labels in the PCP are incorrect and “random labels” indicates that all labels in the PCP are randomly selected. For each dataset, 16 examples per class are used as labelled data and the f…
Figure 3.6
Figure 3.6. Figure 3.6: (Left) The effect of different unlabelled data sizes using ROBERTA-LARGE. (Right) The effect of Scaling Laws, where ROBERTA-BASE (123M) and ROBERTA-LARGE (354M). All comparison approaches are trained with 16 examples per class for each dataset. #2. What are the requi…
Figure 4.1
Figure 4.1. Figure 4.1: The overview of Fine Tuning (FT), Prompt Tuning (PT), and Prompting Engi￾neering. PT increases the length of the input sequence, leading to much greater computational demands during train and inference phrases. Fine-tuning language models (LMs) [177, 223] on downstre…
Figure 4.2
Figure 4.2. Figure 4.2: The overview of the PETL framework (Top) and our method DEPT (Bottom). DEPT decomposes a trainable soft prompt of the vanilla PT into a shorter soft prompt and a couple of low-rank matrices, where the multiplication of low-rank matrices serves to update frozen word e…
Figure 4.3
Figure 4.3. Figure 4.3: Performance on the GLUE benchmark for different soft prompt lengths m in DEPT, associated with corresponding relative train time and memory cost. The time and memory are averaged over different model sizes using batch size as 16. DEPT consistently uses the same numbe…
Figure 4.4
Figure 4.4. Figure 4.4: Average inference speed on GLUE benchmark using varying soft prompt length m and the rank of low-rank matrices r, keeping the total number of trainable parameters constant. Small texts in blue indicate the speed relative to the vanilla PT (represented by brown) (m=10…
Figure 4.5
Figure 4.5. Figure 4.5: Test results on GLUE benchmark using T5-BASE, showing the importance of training DEPT with different learning rates. trices is trained with a lower rate. In our experiments, the first option obtains an average performance of 40.8 on the GLUE benchmark. The second opt…
Figure 5.1
Figure 5.1. Figure 5.1: Performance differences between INSTRUCTION TUNING (IT) and our pro￾posed method INSTRUCTION MODELLING (IM) trained on 7 instruction tun￾ing datasets. (Left) The mean performance across 18 traditional NLP tasks. (Right) The win rate on the AlpacaEval 1.0 benchmark. P…
Figure 5.2
Figure 5.2. Figure 5.2: (Left) Performance improvement, achieved by our approach INSTRUCTION MODELLING (IM) compared to INSTRUCTION TUNING (IT) on the AlpacaE￾val 1.0, against the ratio between average instruction length and average out￾put length in instruction tuning datasets (training si…
Figure 5.3
Figure 5.3. Figure 5.3: An example of instruction tuning training data. and completions up to that point): P(x) = P(I1,I2,...,Im,C1,C2,...,Cn) = m+n ∏ t=1 P(st |s1,s2,...,st−1) (5.1) The loss function, L, for instruction modelling calculates the negative log￾likelihood for both instruction …
Figure 5.4
Figure 5.4. Figure 5.4: (Left) Training loss distribution for each example between our approach IN￾STRUCTION MODELLING (IM) and INSTRUCTION TUNING (IT) on the LIMA dataset. (Right) Test loss distribution for each example between IM and IT on the Tulu V2 dataset, using a 10% randomly sampled…
Figure 5.5
Figure 5.5. Figure 5.5: (Left) Training loss distribution for each example between our approach INSTRUCTION MODELLING (IM) and INSTRUCTION TUNING (IT) on the Alpagasus Dolly 3k dataset. (Right) Test loss distribution for each example between IM and IT on the Tulu V2 dataset, using a 10% sam…
Figure 5.6
Figure 5.6. Figure 5.6: (Left) Training loss distribution for each example between our approach IN￾STRUCTION MODELLING (IM) and INSTRUCTION TUNING (IT) on the Less MMLU Chat dataset. (Right) Test loss distribution for each example between IM and IT on the Tulu V2 dataset, using a 10% sample…
Figure 5.7
Figure 5.7. Figure 5.7: Mean performance on 18 NLP tasks over epochs using LLAMA-2-7B-BASE. This analysis suggests that IM experiences a lower instruction tuning tax com￾pared to IT. #3. Instruction Tuning Tax on the NLP tasks. Previous works show that training LMs with RLHF causes an Align…
Figure 5.8
Figure 5.8. Figure 5.8: Comparison of INSTRUCTION TUNING (IT) and INSTRUCTION MODELLING (IM) methods using OPT-6.7B (Top Row) and LLAMA-2-13B-BASE (Bottom Row) trained on 7 instruction tuning datasets. (Left) The mean per￾formance across 18 NLP tasks. (Right) The win rate on the AlpacaEval …
Figure 5.9
Figure 5.9. Figure 5.9: (Left) Output length comparison between our approach INSTRUCTION MOD￾ELLING (IM) and INSTRUCTION TUNING (IT) across various data utilisa￾tion levels from the Tulu V2 dataset, as evaluated on the AlpacaEval dataset. (Right) Performance comparison (measured by win rate…
Figure 5.10
Figure 5.10. Figure 5.10: AlpacaEval 1.0 performance trends for IM and IT approaches on the LIMA and Alpagasus Dolly 9k datasets across different epochs. LIMA and Alpagasus Dolly 9k datasets. We evaluate the performance of IM and IT over different numbers of epochs. IM consistently surpasses…
Figure 5.11
Figure 5.11. Figure 5.11: Comparative analysis of output lengths for IM and IT across different epochs on Alpagasus Dolly 3k, Alpagasus Dolly 9k, LIMA, and Less Tydiqa datasets. #6. The impact of Epochs on Output Lengths [PITH_FULL_IMAGE:figures/full_fig_p119_5_11.png]
Figure 6.1
Figure 6.1. Figure 6.1: An example of generating a StepGame sample with k = 4. unbinding methods are used to store and retrieve information from and to the TPR M, which we will call memory. 6.3 The StepGame Dataset To design a benchmark dataset that explicitly tests models’ spatial reasonin…
Figure 6.2
Figure 6.2. Figure 6.2: On the left-hand side we have the original chain. Orange entities are those targeted by the question. Besides, we show the same chain with the addition of noise. In green, we represent irrelevant, disconnected and supporting entities. 6.3.3 Distracting Noise To make …
Figure 6.3
Figure 6.3. Figure 6.3: The TP-MANN architecture. PE stands for positional encoder, the sign in the box below the symbol E represents a feed-forward neural network, the ⊗ sign represents the outer-product operator, the • sign represents the inner product operator, and LN represents a layer …
Figure 6.4
Figure 6.4. Figure 6.4: Analysis of TP-MANN’s number of recurrent layers (T). The x-axis is T with which the model has been trained. Each line represents a different value of k of the StepGame dataset. set with k between 6 and 10 with noise. These ks are larger than the largest k used durin…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

277 extracted references · 34 canonical work pages

  1. [2]

    Nemotron-4 340b technical report.arXiv preprint arXiv:2406.11704, 2024

    Bo Adler, Niket Agarwal, Ashwath Aithal, Dong H Anh, Pallab Bhat- tacharya, Annika Brundyn, Jared Casper, Bryan Catanzaro, Sharon Clay, Jonathan Cohen, et al. Nemotron-4 340b technical report.arXiv preprint arXiv:2406.11704, 2024

  2. [3]

    Publicly available clinical BERT embeddings

    Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott. Publicly available clinical BERT embeddings. InProceedings of the 2nd Clinical Natural Language Processing Workshop, pages 72–78, Minneapolis, Minnesota, USA, 2019. Association for Computational Linguistics

  3. [4]

    Reid, Stephen Gould, and Anton van den Hengel

    Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian D. Reid, Stephen Gould, and Anton van den Hengel. Vision- and-language navigation: Interpreting visually-grounded navigation instruc- tions in real environments. In2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-...

  4. [5]

    Pseudo-labeling and confirmation bias in deep semi-supervised learning

    Eric Arazo, Diego Ortego, Paul Albert, Noel E O’Connor, and Kevin McGuinness. Pseudo-labeling and confirmation bias in deep semi-supervised learning. InArXiv preprint, volume abs/1908.02983, 2019. 116 BIBLIOGRAPHY

  5. [6]

    Tran, Dara Bahri, Jianmo Ni, Jai Prakash Gupta, Kai Hui, Sebastian Ruder, and Donald Metzler

    Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran, Dara Bahri, Jianmo Ni, Jai Prakash Gupta, Kai Hui, Sebastian Ruder, and Donald Metzler. Ext5: Towards extreme multi-task scaling for transfer learning. InThe Tenth Inter- national Conference on Learning Representations, ICLR 2022, V...

  6. [7]

    A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings

    Mikel Artetxe, Gorka Labaka, and Eneko Agirre. A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings. InProceedings of the 56th Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 789–798, Melbourne, Australia, 2018. Association for Computational Linguistics

  7. [8]

    ATTEMPT: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts

    Akari Asai, Mohammadreza Salehi, Matthew Peters, and Hannaneh Ha- jishirzi. ATTEMPT: Parameter-efficient multi-task tuning via attentional mixtures of soft prompts. InProceedings of the 2022 Conference on Empiri- cal Methods in Natural Language Processing, pages 6655–6672, Abu Dhabi, United Arab Emirates, 2022. Association for Computational Linguistics

  8. [9]

    Program synthesis with large language models.ArXiv preprint, abs/2108.07732, 2021

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models.ArXiv preprint, abs/2108.07732, 2021

Show all 277 references
  1. [10]

    A general theoret- ical paradigm to understand learning from human preferences

    Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello. A general theoret- ical paradigm to understand learning from human preferences. InInterna- tional Conference on Artificial Intelligence and Statistics, p...

  2. [11]

    Qwen technical report

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023. BIBLIOGRAPHY 117

  3. [12]

    Training a helpful and harmless assistant with reinforcement learning from human feedback.ArXiv preprint, abs/2204.05862, 2022

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback.ArXiv preprint, abs/2204.05862, 2022

  4. [13]

    Brain power.Proceedings of the National Academy of Sciences, 118(32):e2107022118, 2021

    Vijay Balasubramanian. Brain power.Proceedings of the National Academy of Sciences, 118(32):e2107022118, 2021

  5. [14]

    SciBERT: A pretrained language model for scientific text

    Iz Beltagy, Kyle Lo, and Arman Cohan. SciBERT: A pretrained language model for scientific text. InProceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), ...

  6. [15]

    BitFit: Sim- ple parameter-efficient fine-tuning for transformer-based masked language- models

    Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. BitFit: Sim- ple parameter-efficient fine-tuning for transformer-based masked language- models. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1–9, Du...

  7. [16]

    Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell

    Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? InProceedings of the 2021 ACM Conference on Fairness, Account- ability, and Transparency, FAccT ’21, page 610–623, New York,...

  8. [17]

    Bender and Alexander Koller

    Emily M. Bender and Alexander Koller. Climbing towards NLU: On mean- ing, form, and understanding in the age of data. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5185–5198, Online, 2020. Association for Computational Linguistics

  9. [18]

    Curriculum learning

    Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In Andrea Pohoreckyj Danyluk, Léon Bottou, and 118 BIBLIOGRAPHY Michael L. Littman, editors,Proceedings of the 26th Annual International Conference on Machine Learning, ICML 2009, Montreal...

  10. [19]

    The fifth PASCAL recognizing textual entailment challenge

    Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. The fifth PASCAL recognizing textual entailment challenge. InTAC, 2009

  11. [20]

    Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel

    David Berthelot, Nicholas Carlini, Ekin D. Cubuk, Alex Kurakin, Kihyuk Sohn, Han Zhang, and Colin Raffel. Remixmatch: Semi-supervised learn- ing with distribution matching and augmentation anchoring. In8th Interna- tional Conference on Learning Representations, ICLR 2020, Addi...

  12. [21]

    Goodfellow, Nicolas Papernot, Avital Oliver, and Colin Raffel

    David Berthelot, Nicholas Carlini, Ian J. Goodfellow, Nicolas Papernot, Avital Oliver, and Colin Raffel. Mixmatch: A holistic approach to semi- supervised learning. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelz- imer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett...

  13. [22]

    Adamatch: A unified approach to semi-supervised learning and domain adaptation

    David Berthelot, Rebecca Roelofs, Kihyuk Sohn, Nicholas Carlini, and Alexey Kurakin. Adamatch: A unified approach to semi-supervised learning and domain adaptation. InThe Tenth International Conference on Learn- ing Representations, ICLR 2022, Virtual Event, April 25-29, 2022....

  14. [23]

    Shih, Yejin Choi, and Daniel Marcu

    Yonatan Bisk, Kevin J. Shih, Yejin Choi, and Daniel Marcu. Learning in- terpretable spatial operations in a rich 3d blocks world. In Sheila A. McIl- raith and Kilian Q. Weinberger, editors,Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18), ...

  15. [24]

    PIQA: reasoning about physical commonsense in natural language

    Yonatan Bisk, Rowan Zellers, Ronan LeBras, Jianfeng Gao, and Yejin Choi. PIQA: reasoning about physical commonsense in natural language. InThe Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intellige...

  16. [25]

    Bowman, Gabor Angeli, Christopher Potts, and Christopher D

    Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. A large annotated corpus for learning natural language inference. InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal, 2015. Ass...

  17. [26]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Ka- plan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Je...

  18. [27]

    Semi-supervised semantic role labeling with cross-view training

    Rui Cai and Mirella Lapata. Semi-supervised semantic role labeling with cross-view training. InProceedings of the 2019 Conference on Empir- 120 BIBLIOGRAPHY ical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (E...

  19. [28]

    Findings of the 2009 Workshop on Statistical Machine Translation

    Chris Callison-Burch, Philipp Koehn, Christof Monz, and Josh Schroeder. Findings of the 2009 Workshop on Statistical Machine Translation. InPro- ceedings of the Fourth Workshop on Statistical Machine Translation, pages 1–28, Athens, Greece, 2009. Association for Computational ...

  20. [29]

    Extracting training data from large language models

    Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-V oss, Katherine Lee, Adam Roberts, Tom B Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. InUSENIX Security Symposium, volume 6, 2021

  21. [31]

    SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation

    Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Spe- cia. SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation. InProceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 1–14, V...

  22. [32]

    Importance of semantic representation: dataless classification

    Ming-Wei Chang, Lev Ratinov, Dan Roth, and Vivek Srikumar. Importance of semantic representation: dataless classification. InProceedings of the 23rd national conference on Artificial intelligence-Volume 2, pages 830–835, 2008

  23. [33]

    Semi-supervised BIBLIOGRAPHY 121 learning (chapelle, o

    Olivier Chapelle, Bernhard Scholkopf, and Alexander Zien. Semi-supervised BIBLIOGRAPHY 121 learning (chapelle, o. et al., eds.; 2006)[book reviews].IEEE Transactions on Neural Networks, 20(3):542–542, 2009

  24. [34]

    Code alpaca: An instruction-following llama model for code generation

    Sahil Chaudhary. Code alpaca: An instruction-following llama model for code generation. https://github.com/sahil280114/codealpaca, 2023

  25. [35]

    Debiased self-training for semi-supervised learning

    Baixu Chen, Junguang Jiang, Ximei Wang, Pengfei Wan, Jianmin Wang, and Mingsheng Long. Debiased self-training for semi-supervised learning. In Advances in Neural Information Processing Systems, NIPS’22, 2022

  26. [36]

    Unseen filler generalization in attention-based natural language reasoning models

    Chin-Hui Chen, Yi-Fu Fu, Hsiao-Hua Cheng, and Shou-De Lin. Unseen filler generalization in attention-based natural language reasoning models. In2020 IEEE Second International Conference on Cognitive Machine Intelligence (CogMI), pages 42–51. IEEE, 2020

  27. [37]

    TOUCHDOWN: natural language navigation and spatial reasoning in visual street environments

    Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. TOUCHDOWN: natural language navigation and spatial reasoning in visual street environments. InIEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019. Co...

  28. [38]

    MixText: Linguistically-informed interpolation of hidden space for semi-supervised text classification

    Jiaao Chen, Zichao Yang, and Diyi Yang. MixText: Linguistically-informed interpolation of hidden space for semi-supervised text classification. InPro- ceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2147–2157, Online, 2020. Associati...

  29. [39]

    Alpagasus: Training a better alpaca model with fewer data

    Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Ya- dav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. Alpagasus: Training a better alpaca model with fewer data. InThe Twelfth International Conference on Learning Representations, 2024

  30. [40]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, 122 BIBLIOGRAPHY Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Bro...

  31. [41]

    Microsoft coco captions: Data collection and evaluation server.ArXiv preprint, abs/1504.00325, 2015

    Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server.ArXiv preprint, abs/1504.00325, 2015

  32. [42]

    Adapt- ing language models to compress contexts

    Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen. Adapt- ing language models to compress contexts. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing, pages 3829–3846, ...

  33. [43]

    Gonza- lez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonza- lez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023

  34. [44]

    Palm: Scaling language modeling with path- ways.ArXiv preprint, abs/2204.02311, 2022

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gau- rav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sut- ton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with path- ways.ArXiv preprint, abs/2204.02311, 2022. BIBLIOGRAPHY 123

  35. [45]

    Deep reinforcement learning from human preferences

    Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences. In I. Guyon, U. V on Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vish- wanathan, and R. Garnett, editors,Advances in Neural Information P...

  36. [46]

    Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robin- ...

  37. [47]

    BoolQ: Exploring the surprising difficulty of natural yes/no questions

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computatio...

  38. [48]

    Manning, and Quoc Le

    Kevin Clark, Minh-Thang Luong, Christopher D. Manning, and Quoc Le. Semi-supervised sequence modeling with cross-view training. InProceed- ings of the 2018 Conference on Empirical Methods in Natural Language Pro- cessing, pages 1914–1925, Brussels, Belgium, 2018. Association f...

  39. [49]

    Think you have solved ques- 124 BIBLIOGRAPHY tion answering? try arc, the ai2 reasoning challenge.ArXiv preprint, abs/1803.05457, 2018

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved ques- 124 BIBLIOGRAPHY tion answering? try arc, the ai2 reasoning challenge.ArXiv preprint, abs/1803.05457, 2018

  40. [50]

    Training verifiers to solve math word problems, 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021

  41. [51]

    Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

    Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

  42. [52]

    The PASCAL recog- nising textual entailment challenge

    Ido Dagan, Oren Glickman, and Bernardo Magnini. The PASCAL recog- nising textual entailment challenge. Inthe First International Conference on Machine Learning Challenges: Evaluating Predictive Uncertainty Visual Object Classification, and Recognizing Textual Entailment, 2005

  43. [53]

    Flashattention-2: Faster attention with better parallelism and work partitioning, 2023

    Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning, 2023

  44. [54]

    The commitmentbank: Investigating projection in naturally occurring discourse

    Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. The commitmentbank: Investigating projection in naturally occurring discourse. Inproceedings of Sinn und Bedeutung, volume 23, pages 107–124, 2019

  45. [55]

    Universal transformers

    Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. Universal transformers. In7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9,

  46. [56]

    BERT: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langua...

  47. [57]

    Object-based at- tention for spatio-temporal reasoning: Outperforming neuro-symbolic mod- els with flexible distributed architectures.ArXiv preprint, abs/2012.08508, 2020

    David Ding, Felix Hill, Adam Santoro, and Matt Botvinick. Object-based at- tention for spatio-temporal reasoning: Outperforming neuro-symbolic mod- els with flexible distributed architectures.ArXiv preprint, abs/2012.08508, 2020

  48. [59]

    Dolan and Chris Brockett

    William B. Dolan and Chris Brockett. Automatically constructing a corpus of sentential paraphrases. InProceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005

  49. [60]

    A robust self-learning framework for cross- lingual text classification

    Xin Dong and Gerard de Melo. A robust self-learning framework for cross- lingual text classification. InProceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJC...

  50. [61]

    The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

  51. [62]

    Searchqa: A new q&a dataset augmented with context from a search engine.ArXiv preprint, abs/1704.05179, 2017

    Matthew Dunn, Levent Sagun, Mike Higgins, V Ugur Guney, V olkan Cirik, and Kyunghyun Cho. Searchqa: A new q&a dataset augmented with context from a search engine.ArXiv preprint, abs/1704.05179, 2017

  52. [63]

    MRQA 2019 shared task: Evaluating generalization in reading com- prehension

    Adam Fisch, Alon Talmor, Robin Jia, Minjoon Seo, Eunsol Choi, and Danqi Chen. MRQA 2019 shared task: Evaluating generalization in reading com- prehension. InProceedings of the 2nd Workshop on Machine Reading for Question Answering, pages 1–13, Hong Kong, China, 2019. Associati...

  53. [64]

    Making pre-trained language models better few-shot learners

    Tianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: ...

  54. [65]

    Zero-shot text classification with self-training

    Ariel Gera, Alon Halfon, Eyal Shnarch, Yotam Perlitz, Liat Ein-Dor, and Noam Slonim. Zero-shot text classification with self-training. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Process- ing, pages 1107–1119, Abu Dhabi, United Arab Emirates, ...

  55. [67]

    The third PASCAL recognizing textual entailment challenge

    Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. The third PASCAL recognizing textual entailment challenge. InProceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, pages 1–9, Prague, 2007. Association for Computational Linguistics

  56. [68]

    Pars: Pseudo-label aware robust sample selection for learning with noisy labels.ArXiv preprint, abs/2201.10836, 2022

    Arushi Goel, Yunlong Jiao, and Jordan Massiah. Pars: Pseudo-label aware robust sample selection for learning with noisy labels.ArXiv preprint, abs/2201.10836, 2022

  57. [69]

    Making the V in VQA matter: Elevating the role of image under- standing in visual question answering

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image under- standing in visual question answering. In2017 IEEE Conference on Com- puter Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, ...

  58. [70]

    Semi-supervised learning by en- tropy minimization

    Yves Grandvalet and Yoshua Bengio. Semi-supervised learning by en- tropy minimization. InAdvances in Neural Information Processing Systems BIBLIOGRAPHY 127 17 [Neural Information Processing Systems, NIPS 2004, December 13-18, 2004, Vancouver, British Columbia, Canada], pages 5...

  59. [71]

    Projected language models: A large model pre-segmented into smaller ones

    David Grangier, Angelos Katharopoulos, Pierre Ablin, and Awni Hannun. Projected language models: A large model pre-segmented into smaller ones. InICML 2024 Workshop on Foundation Models in the Wild, 2024

  60. [72]

    PPT: Pre-trained prompt tuning for few-shot learning

    Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang. PPT: Pre-trained prompt tuning for few-shot learning. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8410–8423, Dublin, Ireland, 2022. Association for Co...

  61. [73]

    Textbooks are all you need.ArXiv preprint, abs/2306.11644, 2023

    Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al. Textbooks are all you need.ArXiv preprint, abs/2306.11644, 2023

  62. [74]

    Parameter-efficient transfer learning with diff pruning

    Demi Guo, Alexander Rush, and Yoon Kim. Parameter-efficient transfer learning with diff pruning. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long...

  63. [75]

    Suchin Gururangan, Tam Dang, Dallas Card, and Noah A. Smith. Varia- tional pretraining for semi-supervised text classification. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5880–5894, Florence, Italy, 2019. Association for Co...

  64. [76]

    Suchin Gururangan, Ana Marasovi ´c, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. Don’t stop pretraining: Adapt language models to domains and tasks. InProceedings of the 58th Annual 128 BIBLIOGRAPHY Meeting of the Association for Computational Lingu...

  65. [77]

    W ARP: Word-level Adversarial ReProgramming

    Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. W ARP: Word-level Adversarial ReProgramming. InProceedings of the 59th An- nual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume...

  66. [78]

    ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection

    Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. InProceedings of the 60th Annual Meeting of the Association for Computational Lingui...

  67. [79]

    Preserving pre-trained features helps calibrate fine-tuned language models.ArXiv preprint, abs/2305.19249, 2023

    Guande He, Jianfei Chen, and Jun Zhu. Preserving pre-trained features helps calibrate fine-tuned language models.ArXiv preprint, abs/2305.19249, 2023

  68. [80]

    Extending clip for category-to-image re- trieval in e-commerce

    Mariya Hendriksen, Maurits Bleeker, Svitlana Vakulenko, Nanne van Noord, Ernst Kuiper, and Maarten de Rijke. Extending clip for category-to-image re- trieval in e-commerce. InAdvances in Information Retrieval: 44th European Conference on IR Research, ECIR 2022, Stavanger, Norw...

  69. [81]

    Aligning AI with shared human values

    Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. Aligning AI with shared human values. In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021

  70. [82]

    Measuring massive multitask language BIBLIOGRAPHY 129 understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language BIBLIOGRAPHY 129 understanding. In9th International Conference on Learning Representa- tions, ICLR 2021, Virtual Event, Austria, May 3-7,...

  71. [83]

    Training compute-optimal large language models.ArXiv preprint, abs/2203.15556, 2022

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models.ArXiv preprint, abs/2203.15556, 2022

  72. [84]

    Orpo: Monolithic preference optimization without reference model.arXiv preprint arXiv:2403.07691, 2(4):5, 2024

    Jiwoo Hong, Noah Lee, and James Thorne. Orpo: Monolithic preference optimization without reference model.arXiv preprint arXiv:2403.07691, 2(4):5, 2024

  73. [85]

    Unnatural in- structions: Tuning language models with (almost) no human labor

    Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. Unnatural in- structions: Tuning language models with (almost) no human labor. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors,Proceedings of the 61st Annual Meeting of the Association for Computational L...

  74. [86]

    Human feedback is not gold standard

    Tom Hosking, Phil Blunsom, and Max Bartolo. Human feedback is not gold standard. InThe Twelfth International Conference on Learning Representa- tions, 2024

  75. [87]

    Meta-learning the differ- ence: Preparing large language models for efficient adaptation.Transactions of the Association for Computational Linguistics, 10:1249–1265, 2022

    Zejiang Hou, Julian Salazar, and George Polovets. Meta-learning the differ- ence: Preparing large language models for efficient adaptation.Transactions of the Association for Computational Linguistics, 10:1249–1265, 2022

  76. [88]

    Parameter-efficient transfer learning for NLP

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for NLP. In Kamalika Chaud- huri and Ruslan Salakhutdinov, editors,Proceedings of the 36th Inte...

  77. [89]

    Universal language model fine-tuning for text classification

    Jeremy Howard and Sebastian Ruder. Universal language model fine-tuning for text classification. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 328–339, Melbourne, Australia, 2018. Association for Comput...

  78. [90]

    Can llms learn from a single exam- ple?, 2023

    Jeremy Howard and Jonathan Whitaker. Can llms learn from a single exam- ple?, 2023

  79. [91]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. InThe Tenth International Conference on Learn- ing Representations, ICLR 2022, Virtual Event, April 25-29, 2022. O...

  80. [92]

    Mining and summarizing customer reviews

    Minqing Hu and Bing Liu. Mining and summarizing customer reviews. In ACM SIGKDD international conference on Knowledge discovery and data mining, 2004

  81. [93]

    Instruction fine-tuning: Does prompt loss mat- ter?, 2024

    Mathew Huerta-Enochian. Instruction fine-tuning: Does prompt loss mat- ter?, 2024

  82. [94]

    Llama guard: Llm-based input-output safeguard for human-ai conver- sations.arXiv preprint arXiv:2312.06674, 2023

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conver- sations.arXiv preprint arXiv:2312.06674, 2023

  83. [95]

    Hyperdecoders: Instance-specific de- coders for multi-task NLP

    Hamish Ivison and Matthew Peters. Hyperdecoders: Instance-specific de- coders for multi-task NLP. InFindings of the Association for Computational Linguistics: EMNLP 2022, pages 1715–1730, Abu Dhabi, United Arab Emi- rates, 2022. Association for Computational Linguistics. BIBLI...

  84. [96]

    Camels in a changing climate: Enhancing lm adaptation with tulu 2, 2023

    Hamish Ivison, Yizhong Wang, Valentina Pyatkin, Nathan Lambert, Matthew Peters, Pradeep Dasigi, Joel Jang, David Wadden, Noah A Smith, Iz Beltagy, et al. Camels in a changing climate: Enhancing lm adaptation with tulu 2, 2023

  85. [97]

    Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein

    Neel Jain, Ping yeh Chiang, Yuxin Wen, John Kirchenbauer, Hong-Min Chu, Gowthami Somepalli, Brian R. Bartoldson, Bhavya Kailkhura, Avi Schwarzschild, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein. NEFTune: Noisy embeddings improve instruction finetuning. In ...

  86. [98]

    Representation learning for grounded spatial reasoning.Transactions of the Association for Computational Linguistics, 6:49–61, 2018

    Michael Janner, Karthik Narasimhan, and Regina Barzilay. Representation learning for grounded spatial reasoning.Transactions of the Association for Computational Linguistics, 6:49–61, 2018

  87. [99]

    Limit: Less is more for instruction tuning across evaluation paradigms.ArXiv preprint, abs/2311.13133, 2023

    Aditi Jha, Sam Havens, Jeremey Dohmann, Alex Trott, and Jacob Portes. Limit: Less is more for instruction tuning across evaluation paradigms.ArXiv preprint, abs/2311.13133, 2023

  88. [100]

    Scaling laws for neural language models.ArXiv preprint, abs/2001.08361, 2020

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Ben- jamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.ArXiv preprint, abs/2001.08361, 2020

  89. [101]

    Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks

    Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson. Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internat...

  90. [102]

    Looking beyond the surface: A challenge set for reading compre- hension over multiple sentences

    Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. Looking beyond the surface: A challenge set for reading compre- hension over multiple sentences. InProceedings of the 2018 Conference of 132 BIBLIOGRAPHY the North American Chapter of the Associat...

  91. [103]

    UNIFIEDQA: Crossing for- mat boundaries with a single QA system

    Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. UNIFIEDQA: Crossing for- mat boundaries with a single QA system. InFindings of the Association for Computational Linguistics: EMNLP 2020, pages 1896–1907, Online, 2...

  92. [104]

    Scitail: A textual entail- ment dataset from science question answering.Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), Apr

    Tushar Khot, Ashish Sabharwal, and Peter Clark. Scitail: A textual entail- ment dataset from science question answering.Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), Apr. 2018

  93. [105]

    Decomposed prompting: A modular ap- proach for solving complex tasks

    Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. Decomposed prompting: A modular ap- proach for solving complex tasks. InThe Eleventh International Conference on Learning Representations, 2022

  94. [107]

    Kipf and Max Welling

    Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In5th International Conference on Learning Repre- sentations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017

  95. [108]

    Big-loop recurrence within the hippocampal system supports integration of information across episodes.Neuron, 99(6):1342–1354, 2018

    Raphael Koster, Martin J Chadwick, Yi Chen, David Berron, Andrea Banino, Emrah Düzel, Demis Hassabis, and Dharshan Kumaran. Big-loop recurrence within the hippocampal system supports integration of information across episodes.Neuron, 99(6):1342–1354, 2018. BIBLIOGRAPHY 133

  96. [109]

    Situated dialogue and spatial organization: What, where

    Geert-Jan M Kruijff, Hendrik Zender, Patric Jensfelt, and Henrik I Chris- tensen. Situated dialogue and spatial organization: What, where. . . and why? International Journal of Advanced Robotic Systems, 2007

  97. [110]

    Generalization through the recurrent interaction of episodic memories: a model of the hippocampal sys- tem.Psychological review, 2012

    Dharshan Kumaran and James L McClelland. Generalization through the recurrent interaction of episodic memories: a model of the hippocampal sys- tem.Psychological review, 2012

  98. [111]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob De- vlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming- Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and S...

  99. [112]

    Temporal ensembling for semi-supervised learning

    Samuli Laine and Timo Aila. Temporal ensembling for semi-supervised learning. In5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceed- ings. OpenReview.net, 2017

  100. [113]

    T\" ulu 3: Pushing frontiers in open language model post-training.arXiv preprint arXiv:2411.15124, 2024

    Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, et al. T\" ulu 3: Pushing frontiers in open language model post-training.arXiv preprint arXiv:2411.15124, 2024

  101. [114]

    A review of spatial reasoning and interaction for real-world robotics.Advanced Robotics, 2017

    Christian Landsiedel, Verena Rieser, Matthew Walter, and Dirk Wollherr. A review of spatial reasoning and interaction for real-world robotics.Advanced Robotics, 2017

  102. [115]

    Self-attentive associative memory

    Hung Le, Truyen Tran, and Svetha Venkatesh. Self-attentive associative memory. InProceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 ofPro- ceedings of Machine Learning Research, pages 5682–5691. PMLR, 202...

  103. [116]

    Bloom: A 176b-parameter open-access mul- tilingual language model.arXiv preprint, 2023

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili ´c, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. Bloom: A 176b-parameter open-access mul- tilingual language model.arXiv preprint, 2023

  104. [117]

    Teven Le Scao and Alexander Rush. How many data points is a prompt worth? InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2627–2636, Online, 2021. Association for Computa- t...

  105. [118]

    Pseudo-label: The simple and efficient semi- supervised learning method for deep neural networks

    Dong-Hyun Lee et al. Pseudo-label: The simple and efficient semi- supervised learning method for deep neural networks. InWorkshop on chal- lenges in representation learning, ICML, page 896, 2013

  106. [119]

    Biobert: a pre-trained biomedical lan- guage representation model for biomedical text mining.ArXiv preprint, abs/1901.08746, 2019

    Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. Biobert: a pre-trained biomedical lan- guage representation model for biomedical text mining.ArXiv preprint, abs/1901.08746, 2019

  107. [120]

    Scalable agent alignment via reward modeling: a research di- rection.ArXiv preprint, abs/1811.07871, 2018

    Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg. Scalable agent alignment via reward modeling: a research di- rection.ArXiv preprint, abs/1811.07871, 2018

  108. [121]

    The power of scale for parameter-efficient prompt tuning

    Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic, 2021. Association for ...

  109. [122]

    The winograd schema challenge

    Hector Levesque, Ernest Davis, and Leora Morgenstern. The winograd schema challenge. InThirteenth International Conference on the Principles of Knowledge Representation and Reasoning, 2012. BIBLIOGRAPHY 135

  110. [123]

    Levesque, Ernest Davis, and Leora Morgenstern

    Hector J. Levesque, Ernest Davis, and Leora Morgenstern. The winograd schema challenge. In13th International Conference on the Principles of Knowledge Representation and Reasoning, KR 2012, Proceedings of the In- ternational Conference on Knowledge Representation and Reasoning...

  111. [124]

    Semi-supervised text clas- sification with balanced deep representation distributions

    Changchun Li, Ximing Li, and Jihong Ouyang. Semi-supervised text clas- sification with balanced deep representation distributions. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural L...

  112. [125]

    Hogg, and Anthony G

    Fangjun Li, David C. Hogg, and Anthony G. Cohn. Advancing spatial rea- soning in large language models: An in-depth evaluation and enhancement using the stepgame benchmark.Proceedings of the AAAI Conference on Ar- tificial Intelligence, 38(17):18500–18507, Mar. 2024

  113. [126]

    Junnan Li, Richard Socher, and Steven C. H. Hoi. Dividemix: Learning with noisy labels as semi-supervised learning. In8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30,

  114. [127]

    Learn- ing to transfer prompts for text generation

    Junyi Li, Tianyi Tang, Jian-Yun Nie, Ji-Rong Wen, and Xin Zhao. Learn- ing to transfer prompts for text generation. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies, pages 3506–35...

  115. [128]

    Task-adaptive pre-training and self-training are complementary for natural language un- 136 BIBLIOGRAPHY derstanding

    Shiyang Li, Semih Yavuz, Wenhu Chen, and Xifeng Yan. Task-adaptive pre-training and self-training are complementary for natural language un- 136 BIBLIOGRAPHY derstanding. InFindings of the Association for Computational Linguistics: EMNLP 2021, pages 1006–1015, Punta Cana, Domi...

  116. [129]

    Self-alignment with instruction backtranslation

    Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettle- moyer, Jason E Weston, and Mike Lewis. Self-alignment with instruction backtranslation. InThe Twelfth International Conference on Learning Rep- resentations, 2024

  117. [130]

    Prefix-tuning: Optimizing continuous prompts for generation

    Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Pape...

  118. [131]

    Hashimoto

    Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Alpacaeval: An automatic evaluator of instruction-following models.https://github. com/tatsu-lab/alpaca_eval, 5 2023

  119. [132]

    Scaling down to scale up: A guide to parameter-efficient fine-tuning.ArXiv preprint, abs/2303.15647, 2023

    Vladislav Lialin, Vijeta Deshpande, and Anna Rumshisky. Scaling down to scale up: A guide to parameter-efficient fine-tuning.ArXiv preprint, abs/2303.15647, 2023

  120. [133]

    The unlock- ing spell on base llms: Rethinking alignment via in-context learning.ArXiv preprint, abs/2312.01552, 2023

    Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. The unlock- ing spell on base llms: Rethinking alignment via in-context learning.ArXiv preprint, abs/2312.01552, 2023

  121. [134]

    TruthfulQA: Measuring how models mimic human falsehoods

    Stephanie Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring how models mimic human falsehoods. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland, 2022. Association for Com...

  122. [135]

    Fangyu Liu, Qianchu Liu, Shruthi Bannur, Fernando Pérez-García, Naoto Usuyama, Sheng Zhang, Tristan Naumann, Aditya Nori, Hoifung Poon, Javier Alvarez-Valle, Ozan Oktay, and Stephanie L. Hyland. Compositional zero-shot domain transfer with text-to-text models.Transactions of t...

  123. [136]

    Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning

    Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A Raffel. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors,Advances ...

  124. [137]

    SummEQuAL: Summariza- tion evaluation via question answering using large language models

    Junyuan Liu, Zhengyan Shi, and Aldo Lipani. SummEQuAL: Summariza- tion evaluation via question answering using large language models. In Bhavana Dalvi Mishra, Greg Durrett, Peter Jansen, Ben Lipkin, Danilo Neves Ribeiro, Lionel Wong, Xi Ye, and Wenting Zhao, editors,Proceed- i...

  125. [138]

    Statistical rejection sampling improves preference optimization

    Tianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J Liu, and Jialu Liu. Statistical rejection sampling improves preference optimization. InThe Twelfth International Conference on Learning Repre- sentations, 2024

  126. [139]

    What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning.ArXiv preprint, abs/2312.15685, 2023

    Wei Liu, Weihao Zeng, Keqing He, Yong Jiang, and Junxian He. What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning.ArXiv preprint, abs/2312.15685, 2023

  127. [140]

    Gpt understands, too.ArXiv preprint, abs/2103.10385, 2021

    Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. Gpt understands, too.ArXiv preprint, abs/2103.10385, 2021

  128. [141]

    Roberta: A robustly optimized bert pretraining approach.ArXiv preprint, abs/1907.11692, 2019

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi 138 BIBLIOGRAPHY Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach.ArXiv preprint, abs/1907.11692, 2019

  129. [142]

    Zero-shot entity linking by reading entity descriptions

    Lajanugen Logeswaran, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, Jacob Devlin, and Honglak Lee. Zero-shot entity linking by reading entity descriptions. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3449–3460, Florence, I...

  130. [143]

    The ai scientist: Towards fully automated open-ended scientific discovery.arXiv preprint arXiv:2408.06292, 2024

    Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery.arXiv preprint arXiv:2408.06292, 2024

  131. [144]

    Biogpt: generative pre-trained transformer for biomedi- cal text generation and mining.Briefings in bioinformatics, 23(6):bbac409, 2022

    Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. Biogpt: generative pre-trained transformer for biomedi- cal text generation and mining.Briefings in bioinformatics, 23(6):bbac409, 2022

  132. [145]

    Maas, Raymond E

    Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y . Ng, and Christopher Potts. Learning word vectors for sentiment analysis. InProceedings of the 49th Annual Meeting of the Association for Compu- tational Linguistics: Human Language Technologies, pages 142–15...

  133. [146]

    Com- pacter: Efficient low-rank hypercomplex adapter layers

    Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. Com- pacter: Efficient low-rank hypercomplex adapter layers. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jen- nifer Wortman Vaughan, editors,Advances in Neural Information Processin...

  134. [147]

    UniPELT: A unified framework BIBLIOGRAPHY 139 for parameter-efficient language model tuning

    Yuning Mao, Lambert Mathias, Rui Hou, Amjad Almahairi, Hao Ma, Ji- awei Han, Scott Yih, and Madian Khabsa. UniPELT: A unified framework BIBLIOGRAPHY 139 for parameter-efficient language model tuning. InProceedings of the 60th Annual Meeting of the Association for Computational...

  135. [148]

    On the importance of effectively adapting pretrained language models for active learning

    Katerina Margatina, Loic Barrault, and Nikolaos Aletras. On the importance of effectively adapting pretrained language models for active learning. In Proceedings of the 60th Annual Meeting of the Association for Computa- tional Linguistics (Volume 2: Short Papers), pages 825–8...

  136. [149]

    McAuley and Jure Leskovec

    Julian J. McAuley and Jure Leskovec. Hidden factors and hidden topics: un- derstanding rating dimensions with review text. In Qiang Yang, Irwin King, Qing Li, Pearl Pu, and George Karypis, editors,Seventh ACM Conference on Recommender Systems, RecSys ’13, Hong Kong, China, Oct...

  137. [150]

    Effective self- training for parsing

    David McClosky, Eugene Charniak, and Mark Johnson. Effective self- training for parsing. InProceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 152–159, New York City, USA, 2006. Association for Computational Linguistics

  138. [151]

    Can a suit of armor conduct electricity? a new dataset for open book question an- swering

    Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question an- swering. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, Brussels, Belgi...

  139. [152]

    Efficient esti- mation of word representations in vector space.Proceedings of Workshop at ICLR, 2013

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient esti- mation of word representations in vector space.Proceedings of Workshop at ICLR, 2013

  140. [153]

    MetaICL: Learning to learn in context

    Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. MetaICL: Learning to learn in context. InProceedings of the 2022 Con- 140 BIBLIOGRAPHY ference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies, pages...

  141. [154]

    Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Han- naneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demon- strations: What makes in-context learning work? InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, p...

  142. [155]

    Cross-task generalization via natural language crowdsourcing instructions

    Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. Cross-task generalization via natural language crowdsourcing instructions. InProceedings of the 60th Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), pages 3470–34...

  143. [156]

    Vir- tual adversarial training: a regularization method for supervised and semi- supervised learning.ArXiv preprint, abs/1704.03976, 2017

    Takeru Miyato, Shin-ichi Maeda, Masanori Koyama, and Shin Ishii. Vir- tual adversarial training: a regularization method for supervised and semi- supervised learning.ArXiv preprint, abs/1704.03976, 2017

  144. [157]

    Learning to compose soft prompts for compositional zero-shot learning

    Nihal V Nayak, Peilin Yu, and Stephen Bach. Learning to compose soft prompts for compositional zero-shot learning. InThe Eleventh International Conference on Learning Representations, 2022

  145. [158]

    Gpt-4 technical report.ArXiv preprint, abs/2303.08774, 2023

    OpenAI. Gpt-4 technical report.ArXiv preprint, abs/2303.08774, 2023

  146. [159]

    fairseq: A fast, extensible toolkit for sequence modeling

    Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. fairseq: A fast, extensible toolkit for sequence modeling. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Lingu...

  147. [160]

    Training language models to follow instructions with hu- man feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, an...

  148. [161]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, a...

  149. [162]

    Smaug: Fixing failure modes of preference op- timisation with dpo-positive.arXiv preprint arXiv:2402.13228, 2024

    Arka Pal, Deep Karkhanis, Samuel Dooley, Manley Roberts, Siddartha Naidu, and Colin White. Smaug: Fixing failure modes of preference op- timisation with dpo-positive.arXiv preprint arXiv:2402.13228, 2024

  150. [163]

    Recurrent relational networks

    Rasmus Berg Palm, Ulrich Paquet, and Ole Winther. Recurrent relational networks. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kristen Grauman, Nicolò Cesa-Bianchi, and Roman Garnett, editors,Advances in Neural Information Processing Systems 31: Annual Conference on Neura...

  151. [164]

    A sentimental education: Sentiment analysis us- ing subjectivity summarization based on minimum cuts

    Bo Pang and Lillian Lee. A sentimental education: Sentiment analysis us- ing subjectivity summarization based on minimum cuts. InProceedings of the 42nd Annual Meeting of the Association for Computational Linguistics (ACL-04), pages 271–278, Barcelona, Spain, 2004. 142 BIBLIOGRAPHY

  152. [165]

    Seeing stars: Exploiting class relationships for sen- timent categorization with respect to rating scales

    Bo Pang and Lillian Lee. Seeing stars: Exploiting class relationships for sen- timent categorization with respect to rating scales. InProceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL’05), pages 115–124, Ann Arbor, Michigan, 2005. Ass...

  153. [166]

    Iterative reasoning preference optimization

    Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho, He He, Sainbayar Sukhbaatar, and Jason Weston. Iterative reasoning preference optimization. arXiv preprint arXiv:2404.19733, 2024

  154. [167]

    The LAMBADA dataset: Word prediction requiring a broad discourse context

    Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The LAMBADA dataset: Word prediction requiring a broad discourse context. InProceedings of the 54th Annual Meeting of th...

  155. [168]

    Bleu: a method for automatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. InProceedings of the 40th Annual Meeting of the Association for Computational Linguis- tics, pages 311–318, Philadelphia, Pennsylvania, USA, 2002. Assoc...

  156. [169]

    Miriam R. L. Petruck and Michael J. Ellsworth. Representing spatial rela- tions in FrameNet. InProceedings of the First International Workshop on Spatial Language Understanding, pages 41–45, New Orleans, 2018. Associ- ation for Computational Linguistics

  157. [170]

    WiC: the word-in- context dataset for evaluating context-sensitive meaning representations

    Mohammad Taher Pilehvar and Jose Camacho-Collados. WiC: the word-in- context dataset for evaluating context-sensitive meaning representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language T...

  158. [171]

    SemEval-2015 task 8: SpaceE- val

    James Pustejovsky, Parisa Kordjamshidi, Marie-Francine Moens, Aaron Levine, Seth Dworman, and Zachary Yocum. SemEval-2015 task 8: SpaceE- val. InProceedings of the 9th International Workshop on Semantic Evalua- tion (SemEval 2015), pages 884–894, Denver, Colorado, 2015. Associ...

  159. [172]

    Agent q: Advanced reasoning and learning for autonomous ai agents.arXiv preprint arXiv:2408.07199, 2024

    Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov. Agent q: Advanced reasoning and learning for autonomous ai agents.arXiv preprint arXiv:2408.07199, 2024

  160. [173]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and ...

  161. [174]

    Improving language understanding by generative pre-training.arXiv preprint, 2018

    Alec Radford and Karthik Narasimhan. Improving language understanding by generative pre-training.arXiv preprint, 2018

  162. [175]

    Language models are unsupervised multitask learners.Ope- nAI blog, 1(8), 2019

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners.Ope- nAI blog, 1(8), 2019

  163. [176]

    Direct preference optimization: Your lan- guage model is secretly a reward model

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Ste- fano Ermon, and Chelsea Finn. Direct preference optimization: Your lan- guage model is secretly a reward model. InThirty-seventh Conference on Neural Information Processing Systems, 2023

  164. [177]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits 144 BIBLIOGRAPHY of transfer learning with a unified text-to-text transformer.J. Mach. Learn. Res., 21:140:1–140:67, 2020

  165. [179]

    SQuAD: 100,000+ questions for machine comprehension of text

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. SQuAD: 100,000+ questions for machine comprehension of text. InPro- ceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas, 2016. Association for Com...

  166. [180]

    Condita: A state machine like architecture for multi- modal task bots.Alexa Prize TaskBot Challenge Proceedings, 2022

    Jerome Ramos, To Eun Kim, Zhengxiang Shi, Xiao Fu, Fanghua Ye, Yue Feng, and Aldo Lipani. Condita: A state machine like architecture for multi- modal task bots.Alexa Prize TaskBot Challenge Proceedings, 2022

  167. [181]

    Siva Reddy, Danqi Chen, and Christopher D. Manning. CoQA: A conver- sational question answering challenge.Transactions of the Association for Computational Linguistics, 7:249–266, 2019

  168. [182]

    Code llama: Open foundation models for code.arXiv preprint arXiv:2308.12950, 2023

    Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Re- mez, et al. Code llama: Open foundation models for code.arXiv preprint arXiv:2308.12950, 2023

  169. [183]

    AdapterDrop: On the efficiency of adapters in transformers

    Andreas Rücklé, Gregor Geigle, Max Glockner, Tilman Beck, Jonas Pfeif- fer, Nils Reimers, and Iryna Gurevych. AdapterDrop: On the efficiency of adapters in transformers. InProceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing, pages 7930–7946...

  170. [185]

    Winogrande: An adversarial winograd schema challenge at scale

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. InThe Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence...

  171. [186]

    Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V . Nayak, ...

  172. [187]

    Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter W. Battaglia, and Tim Lillicrap. A simple neural network module for relational reasoning. In Isabelle Guyon, Ulrike von 146 BIBLIOGRAPHY Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergu...

  173. [188]

    Anna C Schapiro, Nicholas B Turk-Browne, Matthew M Botvinick, and Ken- neth A Norman. Complementary learning systems within the hippocampus: a neural network modelling approach to reconciling episodic memory with statistical learning.Philosophical Transactions of the Royal Soc...

  174. [189]

    Exploiting cloze-questions for few-shot text classification and natural language inference

    Timo Schick and Hinrich Schütze. Exploiting cloze-questions for few-shot text classification and natural language inference. InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269, Online, 2021....

  175. [190]

    It’s not just size that matters: Small language models are also few-shot learners

    Timo Schick and Hinrich Schütze. It’s not just size that matters: Small language models are also few-shot learners. InProceedings of the 2021 Conference of the North American Chapter of the Association for Compu- tational Linguistics: Human Language Technologies, pages 2339–23...

  176. [191]

    Learning to reason with third order tensor products

    Imanol Schlag and Jürgen Schmidhuber. Learning to reason with third order tensor products. In Samy Bengio, Hanna M. Wallach, Hugo Larochelle, Kris- ten Grauman, Nicolò Cesa-Bianchi, and Roman Garnett, editors,Advances in Neural Information Processing Systems 31: Annual Confere...

  177. [192]

    Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and BIBLIOGRAPHY 147 Oleg Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017

  178. [193]

    Edinburgh neural ma- chine translation systems for WMT 16

    Rico Sennrich, Barry Haddow, and Alexandra Birch. Edinburgh neural ma- chine translation systems for WMT 16. InProceedings of the First Confer- ence on Machine Translation: Volume 2, Shared Task Papers, pages 371– 376, Berlin, Germany, 2016. Association for Computational Linguistics

  179. [194]

    Learning to execute actions or ask clarification questions

    Zhengxiang Shi, Yue Feng, and Aldo Lipani. Learning to execute actions or ask clarification questions. In Marine Carpuat, Marie-Catherine de Marn- effe, and Ivan Vladimir Meza Ruiz, editors,Findings of the Association for Computational Linguistics: NAACL 2022, pages 2060–2070,...

  180. [195]

    Don’t stop pretraining? make prompt- based fine-tuning powerful learner

    Zhengxiang Shi and Aldo Lipani. Don’t stop pretraining? make prompt- based fine-tuning powerful learner. InThirty-seventh Conference on Neural Information Processing Systems, 2023

  181. [196]

    Rethink the effectiveness of text data aug- mentation: An empirical analysis.arXiv preprint arXiv:2306.07664, 2023

    Zhengxiang Shi and Aldo Lipani. Rethink the effectiveness of text data aug- mentation: An empirical analysis.arXiv preprint arXiv:2306.07664, 2023

  182. [197]

    DePT: Decomposed prompt tuning for parameter-efficient fine-tuning

    Zhengxiang Shi and Aldo Lipani. DePT: Decomposed prompt tuning for parameter-efficient fine-tuning. InThe Twelfth International Conference on Learning Representations, 2024

  183. [198]

    Attention-based ingredient phrase parser.arXiv preprint arXiv:2210.02535, 2022

    Zhengxiang Shi, Pin Ni, Meihui Wang, To Eun Kim, and Aldo Lipani. Attention-based ingredient phrase parser.arXiv preprint arXiv:2210.02535, 2022

  184. [199]

    When and what to ask through world states and text instructions: Iglu nlp challenge solution.arXiv preprint arXiv:2305.05754, 2023

    Zhengxiang Shi, Jerome Ramos, To Eun Kim, Xi Wang, Hossein A Rah- mani, and Aldo Lipani. When and what to ask through world states and text instructions: Iglu nlp challenge solution.arXiv preprint arXiv:2305.05754, 2023. 148 BIBLIOGRAPHY

  185. [200]

    Lexical entrainment for conversation systems

    Zhengxiang Shi, Procheta Sen, and Aldo Lipani. Lexical entrainment for conversation systems. InFindings of the Association for Computational Lin- guistics: EMNLP 2023. Association for Computational Linguistics, 2023

  186. [201]

    Rethinking semi-supervised learning with language models

    Zhengxiang Shi, Francesco Tonolini, Nikolaos Aletras, Emine Yilmaz, Gabriella Kazai, and Yunlong Jiao. Rethinking semi-supervised learning with language models. InFindings of the Association for Computational Linguistics: ACL 2023, pages 5614–5634, Toronto, Canada, 2023. Assoc...

  187. [202]

    Self contrastive learning for session-based recommendation

    Zhengxiang Shi, Xi Wang, and Aldo Lipani. Self contrastive learning for session-based recommendation. In Nazli Goharian, Nicola Tonellotto, Yulan He, Aldo Lipani, Graham McDonald, Craig Macdonald, and Iadh Ou- nis, editors,Advances in Information Retrieval, pages 3–20, Cham, 2...

  188. [203]

    Stepgame: A new bench- mark for robust multi-hop spatial reasoning in texts

    Zhengxiang Shi, Qiang Zhang, and Aldo Lipani. Stepgame: A new bench- mark for robust multi-hop spatial reasoning in texts. InThirty-Sixth AAAI Conference on Artificial Intelligence, AAAI 2022, Thirty-Fourth Conference on Innovative Applications of Artificial Intelligence, IAAI...

  189. [204]

    Understanding likelihood over-optimisation in direct alignment algo- rithms.arXiv preprint arXiv:2410.11677, 2024

    Zhengyan Shi, Sander Land, Acyr Locatelli, Matthieu Geist, and Max Bar- tolo. Understanding likelihood over-optimisation in direct alignment algo- rithms.arXiv preprint arXiv:2410.11677, 2024

  190. [205]

    Yang, Bin Wu, Laurence Aitchison, Emine Yilmaz, and Aldo Lipani

    Zhengyan Shi, Adam X. Yang, Bin Wu, Laurence Aitchison, Emine Yilmaz, and Aldo Lipani. Instruction tuning with loss over instructions. InThe Thirty- eighth Annual Conference on Neural Information Processing Systems, 2024

  191. [206]

    Logan IV , Eric Wallace, and Sameer Singh

    Taylor Shin, Yasaman Razeghi, Robert L. Logan IV , Eric Wallace, and Sameer Singh. AutoPrompt: Eliciting Knowledge from Language Models BIBLIOGRAPHY 149 with Automatically Generated Prompts. InProceedings of the 2020 Con- ference on Empirical Methods in Natural Language Proces...

  192. [207]

    Aya dataset: An open-access collection for multilingual instruction tuning.ArXiv preprint, abs/2402.06619, 2024

    Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataci- unas, Laura OMahony, et al. Aya dataset: An open-access collection for multilingual instruction tuning.ArXiv preprint, abs/2402.06619, 2024

  193. [208]

    Tensor product variable binding and the representation of symbolic structures in connectionist systems.Artificial intelligence, 46(1- 2):159–216, 1990

    Paul Smolensky. Tensor product variable binding and the representation of symbolic structures in connectionist systems.Artificial intelligence, 46(1- 2):159–216, 1990

  194. [210]

    Manning, Andrew Ng, and Christopher Potts

    Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. InProceedings of the 2013 Conference on Empirical Methods in Natural Language Process...

  195. [211]

    Fix- match: Simplifying semi-supervised learning with consistency and confi- dence

    Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li. Fix- match: Simplifying semi-supervised learning with consistency and confi- dence. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsel...

  196. [212]

    Learn- ing to summarize with human feedback

    Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F Christiano. Learn- ing to summarize with human feedback. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors,Advances in Neural Info...

  197. [213]

    On transferability of prompt tuning for natural language processing

    Yusheng Su, Xiaozhi Wang, Yujia Qin, Chi-Min Chan, Yankai Lin, Huadong Wang, Kaiyue Wen, Zhiyuan Liu, Peng Li, Juanzi Li, Lei Hou, Maosong Sun, and Jie Zhou. On transferability of prompt tuning for natural language processing. InProceedings of the 2022 Conference of the North ...

  198. [214]

    How to fine-tune bert for text classification? InArXiv preprint, volume abs/1905.05583, 2019

    Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang. How to fine-tune bert for text classification? InArXiv preprint, volume abs/1905.05583, 2019

  199. [215]

    Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks

    Yi-Lin Sung, Jaemin Cho, and Mohit Bansal. Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks. InArXiv preprint, volume abs/2112.06825, 2021

  200. [216]

    LST: Ladder side-tuning for parameter and memory efficient transfer learning

    Yi-Lin Sung, Jaemin Cho, and Mohit Bansal. LST: Ladder side-tuning for parameter and memory efficient transfer learning. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors,Advances in Neu- ral Information Processing Systems, 2022

  201. [217]

    Training neural networks with fixed sparse masks

    Yi-Lin Sung, Varun Nair, and Colin Raffel. Training neural networks with fixed sparse masks. In Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan, editors,Advances in Neural Information Processing Systems 34: Annual Conference ...

  202. [218]

    Challenging BIG-bench tasks and whether chain-of- thought can solve them

    Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, and Jason Wei. Challenging BIG-bench tasks and whether chain-of- thought can solve them. In Anna Rogers, Jordan Boyd-Graber, and Naoak...

  203. [219]

    Source-target inference models for spatial in- struction understanding

    Hao Tan and Mohit Bansal. Source-target inference models for spatial in- struction understanding. In Sheila A. McIlraith and Kilian Q. Weinberger, editors,Proceedings of the Thirty-Second AAAI Conference on Artificial In- telligence, (AAAI-18), the 30th innovative Applications...

  204. [220]

    Hashimoto

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model, 2023

  205. [221]

    Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results

    Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V . N. Vishwanathan, and Roman Garne...

  206. [223]

    Llama 2: Open foundation and fine-tuned chat models

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. ArXiv preprint, abs/2307.09288, 2023

  207. [224]

    NewsQA: A machine com- prehension dataset

    Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sor- doni, Philip Bachman, and Kaheer Suleman. NewsQA: A machine com- prehension dataset. InProceedings of the 2nd Workshop on Representation Learning for NLP, pages 191–200, Vancouver, Canada, 2017. Association...

  208. [225]

    Betty van Aken, Benjamin Winter, Alexander Löser, and Felix A. Gers. How does BERT answer questions?: A layer-wise analysis of transformer repre- sentations. In Wenwu Zhu, Dacheng Tao, Xueqi Cheng, Peng Cui, Elke A. Rundensteiner, David Carmel, Qi He, and Jeffrey Xu Yu, editor...

  209. [226]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In I. Guyon, U. V on Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors,Advances in Neural ...

  210. [227]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wal- lach, Rob Fergus, S. V . N. Vishwanathan, and Roman Garnet...

  211. [228]

    Graph attention networks

    Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph attention networks. In6th Inter- national Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. ...

  212. [229]

    Learning to follow navigational directions

    Adam V ogel and Daniel Jurafsky. Learning to follow navigational directions. InProceedings of the 48th Annual Meeting of the Association for Computa- tional Linguistics, pages 806–814, Uppsala, Sweden, 2010. Association for Computational Linguistics

  213. [230]

    Building a question answering test col- lection

    Ellen M V oorhees and Dawn M Tice. Building a question answering test col- lection. Inthe 23rd annual international ACM SIGIR conference on Research and development in information retrieval, 2000

  214. [231]

    SPoT: Better frozen model adaptation through soft prompt transfer

    Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou’, and Daniel Cer. SPoT: Better frozen model adaptation through soft prompt transfer. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5039–5059, Dublin, Ire...

  215. [232]

    Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’A...

  216. [233]

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and 154 BIBLIOGRAPHY Samuel R. Bowman. GLUE: A multi-task benchmark and analysis plat- form for natural language understanding. In7th International Conference on Learning Representations, ICLR 2019, New Orleans...

  217. [234]

    Meihui Wang, James Haworth, Huanfa Chen, Yunzhe Liu, and Zhengxiang Shi. Investigating the potential of crowdsourced street-level imagery in un- derstanding the spatiotemporal dynamics of cities: A case study of walkabil- ity in inner london.Cities, 153:105243, 2024

  218. [235]

    Self-tuning for data-efficient deep learning

    Ximei Wang, Jinghan Gao, Mingsheng Long, and Jianmin Wang. Self-tuning for data-efficient deep learning. In Marina Meila and Tong Zhang, edi- tors,Proceedings of the 38th International Conference on Machine Learn- ing, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 ofPr...

  219. [236]

    Chi, Sha- ran Narang, Aakanksha Chowdhery, and Denny Zhou

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sha- ran Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency im- proves chain of thought reasoning in language models. InThe Eleventh In- ternational Conference on Learning Representations, 2023

  220. [237]

    USB: A unified semi-supervised learning benchmark for classification

    Yidong Wang, Hao Chen, Yue Fan, Wang SUN, Ran Tao, Wenxin Hou, Ren- jie Wang, Linyi Yang, Zhi Zhou, Lan-Zhe Guo, Heli Qi, Zhen Wu, Yu-Feng Li, Satoshi Nakamura, Wei Ye, Marios Savvides, Bhiksha Raj, Takahiro Shi- nozaki, Bernt Schiele, Jindong Wang, Xing Xie, and Yue Zhang. US...

  221. [238]

    Smith, Daniel Khashabi, and Hannaneh Hajishirzi

    Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language models with self-generated instructions. In Anna Rogers, Jordan Boyd- Graber, and Naoaki Okazaki, editors,Proceedings of the 61st A...

  222. [239]

    OpenReview.net, 2019

  223. [240]

    Multitask prompt tuning enables parameter-efficient transfer learning

    Zhen Wang, Rameswar Panda, Leonid Karlinsky, Rogerio Feris, Huan Sun, and Yoon Kim. Multitask prompt tuning enables parameter-efficient transfer learning. InThe Eleventh International Conference on Learning Represen- tations, 2023

  224. [241]

    Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. Neural network acceptability judgments.Transactions of the Association for Computational Linguistics, 7:625–641, 2019

  225. [242]

    Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

    Jason Wei, Maarten Bosma, Vincent Y . Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V . Le. Finetuned lan- guage models are zero-shot learners. InThe Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-...

  226. [243]

    Towards 156 BIBLIOGRAPHY ai-complete question answering: A set of prerequisite toy tasks

    Jason Weston, Antoine Bordes, Sumit Chopra, and Tomás Mikolov. Towards 156 BIBLIOGRAPHY ai-complete question answering: A set of prerequisite toy tasks. In Yoshua Bengio and Yann LeCun, editors,4th International Conference on Learning Representations, ICLR 2016, San Juan, Puer...

  227. [244]

    Annotating expressions of opinions and emotions in language.Language resources and evaluation, 39(2-3), 2005

    Janyce Wiebe, Theresa Wilson, and Claire Cardie. Annotating expressions of opinions and emotions in language.Language resources and evaluation, 39(2-3), 2005

  228. [245]

    Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks

    Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuzni...

  229. [246]

    A broad-coverage challenge corpus for sentence understanding through inference

    Adina Williams, Nikita Nangia, and Samuel Bowman. A broad-coverage challenge corpus for sentence understanding through inference. InProceed- ings of the 2018 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Human Language Technologie...

  230. [247]

    Prompt com- pression and contrastive conditioning for controllability and toxicity reduc- tion in language models

    David Wingate, Mohammad Shoeybi, and Taylor Sorensen. Prompt com- pression and contrastive conditioning for controllability and toxicity reduc- tion in language models. InFindings of the Association for Computational Linguistics: EMNLP 2022, pages 5621–5634, Abu Dhabi, United ...

  231. [248]

    Adaptive compositional continual meta-learning

    Bin Wu, Jinyuan Fang, Xiangxiang Zeng, Shangsong Liang, and Qiang Zhang. Adaptive compositional continual meta-learning. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors,Proceedings of the 40th International Con...

  232. [249]

    Dynamic bayesian contrastive predictive coding model for personalized product search.ACM Trans

    Bin Wu, Zaiqiao Meng, and Shangsong Liang. Dynamic bayesian contrastive predictive coding model for personalized product search.ACM Trans. Web, 17(4), 2023

  233. [250]

    Meta-learning helps personalized product search

    Bin Wu, Zaiqiao Meng, Qiang Zhang, and Shangsong Liang. Meta-learning helps personalized product search. InProceedings of the ACM Web Confer- ence 2022, WWW ’22, page 2277–2287, New York, NY , USA, 2022. Asso- ciation for Computing Machinery

  234. [252]

    Less: Selecting influential data for instruction tuning, 2024

    Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. Less: Selecting influential data for instruction tuning, 2024

  235. [253]

    Hovy, Thang Luong, and Quoc Le

    Qizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong, and Quoc Le. Un- supervised data augmentation for consistency training. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors,Advances in Neural Information Processing Syste...

  236. [254]

    Hovy, and Quoc V

    Qizhe Xie, Minh-Thang Luong, Eduard H. Hovy, and Quoc V . Le. Self- training with noisy student improves imagenet classification. In2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 10684–10695. IEEE, 2020

  237. [255]

    WizardLM: Empowering large pre-trained language models to follow complex instructions

    Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. WizardLM: Empowering large pre-trained language models to follow complex instructions. InThe Twelfth International Conference on Learning Representations, 2024. 158...

  238. [256]

    Rethinking the instruction quality: Lift is what you need, 2023

    Yang Xu, Yongqiang Yao, Yufan Huang, Mengnan Qi, Maoquan Wang, Bin Gu, and Neel Sundaresan. Rethinking the instruction quality: Lift is what you need, 2023

  239. [257]

    Understanding the role of user profile in the personalization of large language models.arXiv preprint arXiv:2406.17803, 2024

    Bin Wu, Zhengyan Shi, Hossein A Rahmani, Varsha Ramineni, and Emine Yilmaz. Understanding the role of user profile in the personalization of large language models.arXiv preprint arXiv:2406.17803, 2024

  240. [258]

    To repeat or not to repeat: Insights from scaling LLM under token-crisis

    Fuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng, and Yang You. To repeat or not to repeat: Insights from scaling LLM under token-crisis. In Thirty-seventh Conference on Neural Information Processing Systems, 2023

  241. [259]

    mT5: A massively multi- lingual pre-trained text-to-text transformer

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mT5: A massively multi- lingual pre-trained text-to-text transformer. InProceedings of the 2021 Con- ference of the North American Chapter of the Association fo...

  242. [260]

    Bayesian reward models for llm alignment.arXiv preprint arXiv:2402.13210, 2024

    Adam X Yang, Maxime Robeyns, Thomas Coste, Zhengyan Shi, Jun Wang, Haitham Bou-Ammar, and Laurence Aitchison. Bayesian reward models for llm alignment.arXiv preprint arXiv:2402.13210, 2024

  243. [261]

    A theory of representation learning gives a deep generalisation of kernel methods

    Adam X Yang, Maxime Robeyns, Edward Milsom, Ben Anson, Nandi Schoots, and Laurence Aitchison. A theory of representation learning gives a deep generalisation of kernel methods. InInternational Conference on Ma- chine Learning, pages 39380–39415. PMLR, 2023

  244. [262]

    Yang, Maxime Robeyns, Xi Wang, and Laurence Aitchison

    Adam X. Yang, Maxime Robeyns, Xi Wang, and Laurence Aitchison. Bayesian low-rank adaptation for large language models. InThe Twelfth International Conference on Learning Representations, 2023. BIBLIOGRAPHY 159

  245. [263]

    Dash: Semi-supervised learning with dynamic thresholding

    Yi Xu, Lei Shang, Jinxing Ye, Qi Qian, Yu-Feng Li, Baigui Sun, Hao Li, and Rong Jin. Dash: Semi-supervised learning with dynamic thresholding. In Marina Meila and Tong Zhang, editors,Proceedings of the 38th Inter- national Conference on Machine Learning, ICML 2021, 18-24 July ...

  246. [264]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Rus- lan Salakhutdinov, and Christopher D. Manning. HotpotQA: A dataset for diverse, explainable multi-hop question answering. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Proc...

  247. [265]

    Unsupervised word sense disambiguation rivaling super- vised methods

    David Yarowsky. Unsupervised word sense disambiguation rivaling super- vised methods. In33rd Annual Meeting of the Association for Computational Linguistics, pages 189–196, Cambridge, Massachusetts, USA, 1995. Associ- ation for Computational Linguistics

  248. [266]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy, 2019. Association for Comp...

  249. [267]

    Flexmatch: Boosting semi- supervised learning with curriculum pseudo labeling

    Bowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu, Jindong Wang, Manabu Okumura, and Takahiro Shinozaki. Flexmatch: Boosting semi- supervised learning with curriculum pseudo labeling. In Marc’Aurelio Ran- zato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wort- man...

  250. [268]

    Differentiable prompt makes pre-trained 160 BIBLIOGRAPHY language models better few-shot learners

    Ningyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng, Zhen Bi, Chuanqi Tan, Fei Huang, and Huajun Chen. Differentiable prompt makes pre-trained 160 BIBLIOGRAPHY language models better few-shot learners. InThe Tenth International Con- ference on Learning Representations, ICLR 2022,...

  251. [269]

    Robust and in- terpretable grounding of spatial references with relation networks

    Tsung-Yen Yang, Andrew Lan, and Karthik Narasimhan. Robust and in- terpretable grounding of spatial references with relation networks. InFind- ings of the Association for Computational Linguistics: EMNLP 2020, pages 1908–1923, Online, 2020. Association for Computational Linguistics

  252. [270]

    Opt: Open pre-trained transformer language models.ArXiv preprint, abs/2205.01068, 2022

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models.ArXiv preprint, abs/2205.01068, 2022

  253. [271]

    Character-level convolu- tional networks for text classification

    Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. Character-level convolu- tional networks for text classification. In Corinna Cortes, Neil D. Lawrence, Daniel D. Lee, Masashi Sugiyama, and Roman Garnett, editors,Advances in Neural Information Processing Systems 28: Annual Confere...

  254. [272]

    PAWS: Paraphrase adver- saries from word scrambling

    Yuan Zhang, Jason Baldridge, and Luheng He. PAWS: Paraphrase adver- saries from word scrambling. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), ...

  255. [273]

    Long is more for alignment: A simple but tough-to-beat baseline for instruction fine-tuning, 2024

    Hao Zhao, Maksym Andriushchenko, Francesco Croce, and Nicolas Flam- marion. Long is more for alignment: A simple but tough-to-beat baseline for instruction fine-tuning, 2024

  256. [274]

    Slic-hf: Sequence likelihood calibration with human feed- back.arXiv preprint arXiv:2305.10425, 2023

    Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, BIBLIOGRAPHY 161 and Peter J Liu. Slic-hf: Sequence likelihood calibration with human feed- back.arXiv preprint arXiv:2305.10425, 2023

  257. [275]

    OpenReview.net, 2022

  258. [276]

    Record: Bridging the gap between hu- man and machine commonsense reading comprehension.ArXiv preprint, abs/1810.12885, 2018

    Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme. Record: Bridging the gap between hu- man and machine commonsense reading comprehension.ArXiv preprint, abs/1810.12885, 2018

  259. [277]

    Riot: Efficient prompt refinement with residual optimization tree

    Chenyi Zhou, Zhengyan Shi, Yuan Yao, Lei Liang, Huajun Chen, and Qiang Zhang. Riot: Efficient prompt refinement with residual optimization tree. In Proceedings of the Association for Computational Linguistics (ACL), 2025

  260. [278]

    LIMA: Less is more for alignment

    Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, LILI YU, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. LIMA: Less is more for alignment. InThirty-seventh Conference on Neural Information Processin...

  261. [279]

    When can proxies improve the sample complexity of preference learning? 2024

    Yuchen Zhu, Daniel Augusto de Souza, Zhengyan Shi, Mengyue Yang, Pasquale Minervini, Alexander D’Amour, and Matt J Kusner. When can proxies improve the sample complexity of preference learning? 2024

  262. [280]

    unique horror movie

    Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences.arXiv preprint arXiv:1909.08593, 2019. Appendix A The power of Continued Pre-training 164 Appendix A....

  263. [282]

    Calibrate before use: Improving few-shot performance of language models

    Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. In Marina Meila and Tong Zhang, editors,Proceedings of the 38th International Con- ference on Machine Learning, ICML 2021, 18-24 July 2021,...

  264. [283]

    Gonzalez, and Ion Stoica

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-bench and chatbot arena. InThirty-seventh Conference on Neural Infor-...

  265. [288]

    </s>", "Q

    We choose a batch size of 32. We find that training LoRA on the MRQA dataset presents challenges, despite conducting a thorough search for optimal learning rates and training steps. The reasons for these difficulties remain uncertain. For prompt tuning and DEPT, as shown in Ta...

  266. [289]

    Is this wrong?

    We follow the HuggingFace Open LLM Leaderboard to 0 few-shot examples. The output generation terminates upon encountering a new line followed by "Q:". The mean F1 score is used as the evaluation metric. PIQA.Evaluation on the dataset at the huggingface dataset 9 involves a mul...

  267. [2020]

    OpenReview.net, 2020

  268. [2021]

    Association for Computing Machinery

  269. [2022]

    Association for Computational Linguistics

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.