Pith. sign in

REVIEW 3 major objections 4 minor 228 references

Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This thesis argues that task bots can adapt to new user behavior, take on new tasks, and avoid hallucinated facts through self-learning loops that minimize human annotation.

desk verdict A thesis that repackages three already-published papers: no new results, but the underlying methods are solid and the limitations are honestly stated at the end. read the letter →

arxiv 2508.19689 v1 pith:WL35XDB3 submitted 2025-08-27 cs.CL

classification cs.CL
keywords task-orienteddialogueself-learningrewardmodelreinforcementlearningschema-guidedpromptinghallucinationmitigationdirectpreferenceoptimizationfactualityalignment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The thesis targets three post-deployment failures of task-oriented bots: unfamiliar user phrasing, new task definitions, and factually wrong responses. Its central claim is that each can be handled by a self-learning loop that replaces human labels with internal signals. SL-AGENT refines a generative bot on unlabeled human-bot logs using a pre-trained turn-level reward model and reinforcement learning, approaching the performance of turn-level human feedback. SGP-TOD shows that a frozen LLM guided by a task schema, consisting of belief instructions and a dialog policy skeleton, can zero-shot outperform few-shot prompting. Self-Alignment for Factuality shows that an LLM's own true/false judgments about its generated claims can serve as preference labels for DPO, reducing hallucinations in question answering and biography generation. If these loops hold, task bots could be maintained with drastically less human data collection and annotation.

What carries the argument

Three mechanisms carry the argument. (1) A turn-level reward model trained with a contrastive objective on positive examples (original and back-translated user utterances) and five negative categories (repetition, inconsistency, partial information, non-fluency, misunderstanding); this model judges response quality in unlabeled logs and drives REINFORCE refinement. (2) A task schema, composed of a task-specific ontology listing slots and values and a policy skeleton of template dialog turns; two prompters, the DST Prompter and Policy Prompter, convert the schema and dialog history into prompts for a frozen LLM, with belief states expressed as SQL-like queries and database state fetched expli

What would settle it

Take unlabeled human-bot logs from a deployed bot, have human annotators score each turn, and compare with the SL-AGENT reward model's scores; if rank correlation is near zero on turns containing unseen slot values or phrasings, then the RL refinement loop cannot systematically distinguish good from bad responses and will amplify reward-model noise rather than improve the bot.

Watch

Extended reading notes

Core claim

The paper establishes, across three frameworks, that internal signals can substitute for external supervision in task-bot maintenance. In SL-AGENT, a reward model is trained on a few labeled dialogs plus synthetic positive and negative examples, then used to score unlabeled human-bot logs; REINFORCE-style refinement of the dialog model improves Inform, Success, and Combined scores on four single-domain tasks, matching or nearing the upper bound set by turn-level human feedback. In SGP-TOD, a frozen LLM prompted with belief instructions and a policy skeleton achieves state-of-the-art zero-shot results on MultiWOZ, RADDLE, and STAR, outperforming few-shot prompting baselines and, on domain-ext

Load-bearing premise

The load-bearing premise is that the pre-trained reward model scores response quality correctly on unlabeled human-bot logs, including novel user phrasings and newly extended tasks; if that signal is wrong, reinforcement learning entrenches the error instead of fixing it.

Editorial extensions

If this is right

  • A deployed task bot can improve on unseen user phrasings with zero new human annotations, reaching levels close to what turn-level human feedback would provide.
  • New task capabilities can be added by editing the task schema: inserting, amending, or removing template turns in the policy skeleton, with no new training data.
  • Factuality of an LLM can be improved by using its own claim-level true/false evaluations as DPO training signals, reducing the need for human preference labeling.
  • The three mechanisms can be composed: machine teaching supplies a few corrected dialogs for a new function, the reward model then lets the bot self-refine from logs, and self-evaluation guards the responses against hallucination.
  • Confidence calibration of self-evaluation, improved by SK-TUNING, makes the factuality signal more reliable and can likely transfer to other self-improvement settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The five negative-response categories in SL-AGENT are a portable recipe: any domain where one can define repetition, inconsistency, partial information, non-fluency, and misunderstanding could reuse the same reward-model construction without hand-labeling.
  • SGP-TOD's schema-guided prompting and SL-AGENT's self-refinement could be combined so that after a schema extension is deployed, the bot automatically adapts to how real users phrase queries about the new slots, potentially removing even the machine-teaching correction step.
  • SELF-EVAL's claim-level factuality scores could also be used at inference time to rank multiple candidate responses, not only as DPO training labels; the paper does not test this directly.
  • The main risk in self-alignment is that when the LLM's internal knowledge is wrong on a topic, its self-evaluation will confidently label the wrong claim as true; detecting knowledge boundaries, which the thesis lists as future work, would be a natural safeguard.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript presents three self-contained chapters addressing three post-deployment challenges for task-oriented dialogue systems: adaptability to unseen user behaviors (SL-Agent), extensibility to new tasks (SGP-TOD), and factuality of generated responses (Self-Alignment for Factuality). SL-Agent trains a turn-level reward model on synthetic positive/negative examples and uses REINFORCE on unlabeled human-bot logs to refine a SOLOIST dialog model. SGP-TOD uses hand-built belief instructions and policy skeletons to prompt a frozen LLM to perform belief-state tracking, action selection, and response generation in a zero-shot manner. Self-Alignment for Factuality samples multiple candidate responses, estimates their factuality via the LLM's own self-evaluation (with optional SK-Tuning), constructs preference pairs, and fine-tunes with DPO. The three components are evaluated on MultiWOZ/RADDLE/STAR and TruthfulQA/BioGEN with automatic and human metrics. The thesis claims minimal or zero human intervention across all three axes.

Significance. If the claims are fully supported, the work would be practically significant: SL-Agent offers a way to exploit unlabeled interaction logs without human reward annotation; SGP-TOD provides a training-free alternative to fine-tuning for new task schemas; Self-Alignment points to a DPO-based route to factuality using the model's own knowledge signal. The experiments are broad and include comparisons with several strong baselines, ablations, and human evaluations. However, the strongest advertised property—zero-human-annotation adaptation to truly unseen behaviors—is weakened by the simulation design, and the factuality pipeline delegates key steps to an external model and uses golden-answer supervision during SK-Tuning. The contribution is therefore a useful set of empirical frameworks rather than a fully autonomous self-learning system.

major comments (3)
  1. [§3.3.2, Eq. (3.2), Table 3.4] The simulation evidence for zero-annotation adaptability is circular as presented. The reward model is trained to discriminate five hand-defined corruption categories in §3.2.3 (repetition, inconsistency, partial information, non-fluency, misunderstanding), and the 'unseen' human-bot logs in §3.3.2 are generated 'by introducing noise through response corruption.' If the same corruption taxonomy was used to create the simulated logs, then Table 3.4 demonstrates that SL-Agent can recognize and correct the exact error types its reward model was built to detect, not that it generalizes to genuinely novel user behaviors. The manuscript should specify the corruption procedure in the simulation and, for a load-bearing test, hold out one or more error categories from reward training. The real-scenario experiment (Table 3.5) is less circular but uses only 30 logs and does not measure reward-model
  2. [§5.2.1–5.2.3, Eq. (5.3)] The factuality framework is self-referential: preference labels are derived from the same model's self-evaluation p(True|q,a), and DPO then trains that model to prefer responses with high self-evaluation scores. Systematic overconfidence or task-specific bias in SELF-EVAL will therefore be reinforced rather than corrected. The manuscript acknowledges overconfidence in §5.2.2 and introduces SK-Tuning as mitigation, but SK-Tuning itself requires ground-truth answers and Deberta-Large-MNLI entailment to construct True/False labels, so the pipeline is not purely self-supervised. Moreover, claim extraction and question generation are delegated to GPT-3.5-turbo, so the 'self' is partly external. The paper should quantify how much of the DPO gain survives when preference labels are replaced by oracle factuality labels, and should compare against a non-self-referential reward model.
  3. [§4.2.4, §4.3.1, Table 4.1] The zero-shot extensibility claim rests on manually engineered task schemas: belief instructions contain all slot names and plausible values, and policy skeletons contain 10–20 hand-written template turns per task. The 'zero-shot' label is standard in the prompting literature, but the central claim of 'minimal human effort' is sensitive to schema-engineering cost. The manuscript should include an explicit accounting of the human effort needed to author a new schema (or a study of how much performance degrades when the schema is imperfect), and it should state that the method does not remove the need for symbolic task design.
minor comments (4)
  1. [§3.3.2, Table 3.4] The text says 'Table 4.2 presents the end-to-end evaluation results'; this should be Table 3.4. Similar cross-reference errors appear in §3.3.3 ('reported in Table 4.2') and §3.2.2.
  2. [Table 3.4, Table 3.5] Significance is reported only as 'p < 0.01 based on Combined' without stating the test, the number of runs, or the variance. Given that the table reports per-domain results, a paired test across domains or a confidence interval would be more informative.
  3. [§5.2.2, Table 5.1] The prompt shown uses a True/False format, but the narrative sometimes refers to 'A'/'B' as the output. Clarify whether the reported p(True) is the probability assigned to the 'True' token or the probability of the letter 'A'.
  4. [Chapter 3–5] The manuscript repeatedly uses inconsistent spacing in model names such as 'S OLOIST', 'SL-S OLOIST', and 'S GP-TOD'. A final formatting pass is needed.

Circularity Check

1 steps flagged · score 4.0 of 10

One self-referential reward loop in the factuality chapter; adaptability and extensibility chapters are externally benchmarked and not demonstrated circular.

  1. self definitional [Section 5.2.1 Overview (Steps 2-3), Section 5.2.2 Eq. 5.1, Section 5.2.3 Eq. 5.3]
    "In this step, we evaluate the factuality of the generated candidate responses ... by leveraging the intrinsic knowledge of LLMs. ... we select the top α responses as the preferred responses y_w and the remaining responses as the dis-preferred ones y_l, resulting in a set of preference pairs D = {(x,y_w,y_l)}. ... Finally, we align the LLM with these preference data via DPO. (Sec. 5.2.1) ... p(True|q,a) = f_M(q,a) (Eq. 5.1)"

    The preference labels used for DPO are produced by SELF-EVAL, which is built on the same LLM M whose responses it judges (p(True|q,a)=f_M(q,a)). DPO (Eq. 5.3) then fits the policy to prefer exactly those responses that M's own evaluator scores higher. If p(True) is miscalibrated for a claim, the loop reinforces the error rather than correcting it; the paper itself concedes overconfidence with SELF-EVAL-P(TRUE) and adds external SK-Tuning with golden answers to patch the signal. The external test benchmarks and SK-Tuning provide independent evidence, so the circularity is partial: the training labels are self-defined, but the final factuality claim is externally evaluated.

full rationale

The thesis contains three largely independent contributions. SGP-TOD (Ch. 4) is a prompting strategy evaluated against external benchmarks (MultiWOZ, RADDLE, STAR, domain-extension); the manual policy skeleton is derived from a few training dialogs, which weakens the 'zero-shot' label but is not a circular derivation. SL-Agent (Ch. 3) trains a reward model on human-annotated error categories and then applies it to unlabeled logs; this is a legitimate two-stage pipeline. The main unresolved risk is the simulation: Section 3.3.2 says the 45 'unseen' dialogs are made imperfect 'by introducing noise through response corruption,' while Section 3.2.3 defines the reward model's negative examples using five specific corruption categories. The manuscript does not explicitly state that the simulation corruptions are the same five categories, so I do not count this as a demonstrated reduction, but the lack of specification leaves the zero-annotation adaptability claim vulnerable to that critique. The clearest circular feature is in Chapter 5: the factuality preference data are labeled by the model's own self-evaluation (Eq. 5.1) and then used as DPO targets (Eq. 5.3), making the raw reward loop self-referential. The paper acknowledges overconfidence and adds SK-Tuning with external golden labels, and the final evaluation is external, so the central claim still has independent content. Overall score 4.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The thesis contributes three frameworks, each resting on domain assumptions about model self-evaluation and hand-crafted knowledge. No novel physical or conceptual entities are introduced.

free parameters (2)
  • top fraction alpha for preference selection
    In Self-Alignment, the top alpha fraction of responses are marked preferred for DPO; value is chosen by hand, not derived, and affects preference data quality.
  • DPO beta
    Regularization strength in DPO (Eq. 5.3) controls deviation from reference policy; chosen as hyperparameter, not fitted.
assumptions (4)
  • domain assumption A pre-trained reward model can assess response quality in unlabeled human-bot logs, including novel user behaviors.
    SL-Agent's RL loop relies on the reward model to provide correct quality signals (Section 3.2.3).
  • domain assumption An LLM's self-evaluation probability p(True|q,a) correlates with factual correctness.
    Self-Alignment uses self-evaluation as the preference label for DPO (Section 5.2.2).
  • domain assumption Hand-crafted task schemas (ontology plus dialog flow) are sufficient to guide a frozen LLM to complete new tasks.
    SGP-TOD relies on the premise that a schema captures all needed task knowledge (Section 4.2).
  • standard math REINFORCE policy gradient and DPO loss are valid optimization objectives.
    Used as training algorithms (Eq. 3.3, Eq. 5.3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality." pith.science (2026). https://pith.science/paper/WL35XDB3

@misc{pith2026250819689,
  author       = {Pith},
  title        = {Pith review of: Building Task Bots with Self-learning for Enhanced Adaptability, Extensibility, and Factuality},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WL35XDB3}},
  note         = {Machine review of arXiv:2508.19689}
}
read the original abstract

Developing adaptable, extensible, and accurate task bots with minimal or zero human intervention is a significant challenge in dialog research. This thesis examines the obstacles and potential solutions for creating such bots, focusing on innovative techniques that enable bots to learn and adapt autonomously in constantly changing environments.

Figures

Figures reproduced from arXiv: 2508.19689 by the authors.

Figure 1.1
Figure 1.1. Architectures of two end-to-end task-oriented dialogue systems: [PITH_FULL_IMAGE:figures/full_fig_p018_1_1.png] view at source ↗
Figure 2.1
Figure 2.1. An example of a dialog turn sequence at turn [PITH_FULL_IMAGE:figures/full_fig_p028_2_1.png] view at source ↗
Figure 2.2
Figure 2.2. Evolution of training paradigms for developing end-to-end task bots: [PITH_FULL_IMAGE:figures/full_fig_p030_2_2.png] view at source ↗
Figures from the paper (30 more)
Figure 2.3
Figure 2.3. Figure 2.3: An illustrative example of a dialog model that employs an auto-regressive [PITH_FULL_IMAGE:figures/full_fig_p032_2_3.png]
Figure 2.4
Figure 2.4. Figure 2.4: Illustration of proactive learning in a self-feeding chatbot. The model [PITH_FULL_IMAGE:figures/full_fig_p037_2_4.png]
Figure 2.5
Figure 2.5. Figure 2.5: Illustration of the machine teaching process [ [PITH_FULL_IMAGE:figures/full_fig_p038_2_5.png]
Figure 2.6
Figure 2.6. Figure 2.6: Dialog policy optimization in an RL loop, where the interaction between a [PITH_FULL_IMAGE:figures/full_fig_p039_2_6.png]
Figure 2.7
Figure 2.7. Figure 2.7: An example of a task schema from the MultiWOZ dataset [ [PITH_FULL_IMAGE:figures/full_fig_p043_2_7.png]
Figure 2.8
Figure 2.8. Figure 2.8: An overview of the ANYTOD system, cited from Zhao et al. [ [PITH_FULL_IMAGE:figures/full_fig_p045_2_8.png]
Figure 2.9
Figure 2.9. Figure 2.9: Illustration of the prompting paradigm in zero-shot, one-shot, and few-shot [PITH_FULL_IMAGE:figures/full_fig_p047_2_9.png]
Figure 2.10
Figure 2.10. Figure 2.10: An example of hallucinations in LLMs: given the same prompt, an LLM [PITH_FULL_IMAGE:figures/full_fig_p050_2_10.png]
Figure 2.11
Figure 2.11. Figure 2.11: A diagram illustrating the three steps of Reinforcement Learning from [PITH_FULL_IMAGE:figures/full_fig_p053_2_11.png]
Figure 2.12
Figure 2.12. Figure 2.12: Comparison of PPO vs. DPO, where DPO optimizes for human prefer [PITH_FULL_IMAGE:figures/full_fig_p055_2_12.png]
Figure 3.1
Figure 3.1. Figure 3.1: Illustration of the proposed SL-AGENT with a human-bot dialog example. (i) The human-bot dialog example, containing an inappropriate response related to unseen user behaviors (upper part). (ii) Demonstration of the refining process in SL￾AGENT with the exhibited dial…
Figure 3.2
Figure 3.2. Figure 3.2: The proposed SL-AGENT operates as follows: (i) Fine-tune the bot using available task-specific dialogs. (ii) Deploy the bot online to gather unlabeled human￾bot dialog logs. (iii) Refine the dialog model using reinforcement learning with the fine-tuned reward model. …
Figure 3.3
Figure 3.3. Figure 3.3: Illustration of synthetic dialog construction. Slot values in the delexicalized [PITH_FULL_IMAGE:figures/full_fig_p065_3_3.png]
Figure 3.4
Figure 3.4. Figure 3.4: The summarized five types of dialog turns featuring inappropriate or in [PITH_FULL_IMAGE:figures/full_fig_p067_3_4.png]
Figure 3.5
Figure 3.5. Figure 3.5: Illustration of the training example, i.e., the processed dialog turn in the [PITH_FULL_IMAGE:figures/full_fig_p074_3_5.png]
Figure 3.6
Figure 3.6. Figure 3.6: Two interactive examples. (a) An interactive example between user and [PITH_FULL_IMAGE:figures/full_fig_p082_3_6.png]
Figure 4.1
Figure 4.1. Figure 4.1: The proposed SGP-TOD is depicted with a dialog example, where the [PITH_FULL_IMAGE:figures/full_fig_p087_4_1.png]
Figure 4.2
Figure 4.2. Figure 4.2: Illustration of belief state prediction utilizing DST Prompter. The predicted [PITH_FULL_IMAGE:figures/full_fig_p090_4_2.png]
Figure 4.3
Figure 4.3. Figure 4.3: Illustration of system action determination and response generation em [PITH_FULL_IMAGE:figures/full_fig_p092_4_3.png]
Figure 4.4
Figure 4.4. Figure 4.4: Zero-shot end-to-end evaluation results on [PITH_FULL_IMAGE:figures/full_fig_p103_4_4.png]
Figure 4.5
Figure 4.5. Figure 4.5: Detailed belief instructions in DST Prompter. [PITH_FULL_IMAGE:figures/full_fig_p112_4_5.png]
Figure 4.6
Figure 4.6. Figure 4.6: A formatting example in Policy Prompter. [PITH_FULL_IMAGE:figures/full_fig_p113_4_6.png]
Figure 4.7
Figure 4.7. Figure 4.7: Policy Prompter of SGP-TOD on STAR. The relevant template turn within the input, the generated user template utterance, and the system action in the output are accentuated. 99 [PITH_FULL_IMAGE:figures/full_fig_p114_4_7.png]
Figure 4.8
Figure 4.8. Figure 4.8: Policy Prompter of SGP-TOD-E2E on [PITH_FULL_IMAGE:figures/full_fig_p115_4_8.png]
Figure 5.1
Figure 5.1. Figure 5.1: Illustration of Self-Alignment for Factuality. Given a prompt to write a biography, before factuality alignment, the LLM generates some facts that are not accurate. Through self-evaluation, the LLM is capable of identifying these inaccurate facts. The feedback from t…
Figure 5.2
Figure 5.2. Figure 5.2: A diagram illustrating the three steps of our [PITH_FULL_IMAGE:figures/full_fig_p120_5_2.png]
Figure 5.3
Figure 5.3. Figure 5.3: The process of constructing training data for SK-T [PITH_FULL_IMAGE:figures/full_fig_p121_5_3.png]
Figure 5.4
Figure 5.4. Figure 5.4: Results of pairwise comparisons on BioGEN across four dimensions: factu￾ality, helpfulness, relevance and naturalness, as evaluated by GPT-4. The left and right sections present the win rates of Self-Alignment for Factuality w/ SELF-EVAL-SKT against FACTTUNE-MC and S…
Figure 5.5
Figure 5.5. Figure 5.5: Calibration curves of utilizing SELF-EVAL-P(TRUE) and SELF-EVAL￾SKT on LLAMA2-7B in the CommonsenseQA task. Following Kadavath et al. [65], we plot confidence vs. frequency that a prediction is correct. The dashed line indicates perfect calibration. tion. We present …
Figure 5.6
Figure 5.6. Figure 5.6: Calibration curves of utilizing SELF-EVAL-P(TRUE) and SELF-EVAL￾SKT (without duplicates) on LLAMA2-7B in the CommonsenseQA task. Following Kadavath et al. [65], we plot confidence vs. frequency that a prediction is correct. The dashed line indicates perfect calibrati…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

228 extracted references · 14 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Zeyuan Allen-Zhu and Yuanzhi Li. 2023. http://arxiv.org/abs/2309.14402 Physics of language models: Part 3.2, knowledge manipulation

  4. [4]

    Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403

  5. [5]

    Anthropic. 2024. https://www.anthropic.com/news/claude-3-5-sonnet Claude 3.5 sonnet . Anthropic Blog

  6. [6]

    Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan. 2021. http://arxiv.org/abs/2112.00861 A ...

  7. [7]

    Amos Azaria and Tom Mitchell. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.68 The internal state of an LLM knows when it ' s lying . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 967--976, Singapore. Association for Computational Linguistics

  8. [8]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, ...

Show all 228 references
  1. [9]

    Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan P...

  2. [10]

    Ralph Allan Bradley and Milton E Terry. 1952. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4):324--345

  3. [11]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...

  4. [12]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  5. [13]

    Pawe Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, I \ n igo Casanueva, Ultes Stefan, Ramadan Osman, and Milica Ga s i\'c. 2018 a . Multiwoz - a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling. In Proceedings of the 2018 Conference on Empir...

  6. [14]

    Pawe Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Inigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Ga s i \'c . 2018 b . Multiwoz--a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling. arXiv preprint arXiv:1810.00278

  7. [15]

    Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeff Wu. 2023. http://arxiv.org/abs/2312.09390 Weak-to-strong generalization: Eliciting strong capabilit...

  8. [16]

    Yu, Qiang Yang, and Xing Xie

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2023. http://arxiv.org/abs/2307.03109 A survey on evaluation of large language models

  9. [17]

    Jiefeng Chen, Jinsung Yoon, Sayna Ebrahimi, Sercan O Arik, Tomas Pfister, and Somesh Jha. 2023 a . http://arxiv.org/abs/2310.11689 Adaptation with self-evaluation to improve selective prediction in llms

  10. [18]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374

  11. [19]

    Xinyun Chen, Renat Aksitov, Uri Alon, Jie Ren, Kefan Xiao, Pengcheng Yin, Sushant Prakash, Charles Sutton, Xuezhi Wang, and Denny Zhou. 2023 b . http://arxiv.org/abs/2311.17311 Universal self-consistency for large language model generation

  12. [20]

    Smith, and Tao Yu

    Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Tao Yu. 2023. Binding language models in symbolic languages. ICLR, abs/2210.02875

  13. [21]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vino...

  14. [22]

    Christiano, Jan Leike, Tom B

    Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. https://proceedings.neurips.cc/paper/2017/hash/d5e2c0adad503c91f91df240d0cd4e49-Abstract.html Deep reinforcement learning from human preferences . In Advances in Neural Information ...

  15. [23]

    Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. 2023. Dola: Decoding by contrasting layers improves factuality in large language models. arXiv preprint arXiv:2309.03883

  16. [24]

    Roi Cohen, Mor Geva, Jonathan Berant, and Amir Globerson. 2023. https://doi.org/10.18653/v1/2023.findings-eacl.139 Crawling the internal knowledge-base of language models . In Findings of the Association for Computational Linguistics: EACL 2023, pages 1856--1869, Dubrovnik, Cr...

  17. [25]

    ContextualAI. 2024. https://contextual.ai/introducing-rag2/ Introducing rag 2.0

  18. [26]

    Yinpei Dai, Hangyu Li, Chengguang Tang, Yongbin Li, Jian Sun, and Xiaodan Zhu. 2020. Learning low-resource end-to-end goal-oriented dialog for fast and reliable system deployment. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages...

  19. [27]

    Google DeepMind. 2024. https://deepmind.google/technologies/gemini/ Gemini 2.0 . Google Blog

  20. [29]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 a . http://arxiv.org/abs/1810.04805 Bert: Pre-training of deep bidirectional transformers for language understanding

  21. [30]

    Jacob Devlin, Ming - Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 b . https://doi.org/10.18653/v1/n19-1423 BERT: pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North American Chapter of the Assoc...

  22. [31]

    Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023. http://arxiv.org/abs/2309.11495 Chain-of-verification reduces hallucination in large language models

  23. [32]

    Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. 2023. http://arxiv.org/abs/2301.00234 A survey on in-context learning

  24. [33]

    Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018. Understanding back-translation at scale. arXiv preprint arXiv:1808.09381

  25. [34]

    Teddy Ferdinan, Jan Kocoń, and Przemysław Kazienko. 2024. http://arxiv.org/abs/2402.09147 Into the unknown: Self-learning large language models

  26. [35]

    Jan-Philipp Fränken, Eric Zelikman, Rafael Rafailov, Kanishk Gandhi, Tobias Gerstenberg, and Noah D. Goodman. 2024. http://arxiv.org/abs/2404.14313 Self-supervised alignment with mutual information: Learning to follow principles without preference labels

  27. [36]

    Zeyu Gan and Yong Liu. 2024. http://arxiv.org/abs/2410.01720 Towards a theoretical understanding of synthetic data in llm post-training: A reverse-bottleneck perspective

  28. [37]

    Jianfeng Gao, Michel Galley, and Lihong Li. 2018 a . https://doi.org/10.18653/v1/P18-5002 Neural approaches to conversational AI . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics: Tutorial Abstracts, pages 2--7, Melbourne, Australia. ...

  29. [38]

    Jianfeng Gao, Michel Galley, and Lihong Li. 2018 b . Neural approaches to conversational ai. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, pages 1371--1374

  30. [39]

    Jianfeng Gao, Michel Galley, and Lihong Li. 2019. http://arxiv.org/abs/1809.08267 Neural approaches to conversational ai

  31. [40]

    Luyu Gao, Zhuyun Dai, Panupong Pasupat, Anthony Chen, Arun Tejasvi Chaganty, Yicheng Fan, Vincent Zhao, Ni Lao, Hongrae Lee, Da-Cheng Juan, and Kelvin Guu. 2023. https://doi.org/10.18653/v1/2023.acl-long.910 RARR : Researching and revising what language models say, using langu...

  32. [41]

    Silin Gao, Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020. Paraphrase augmented task-oriented dialog generation. arXiv preprint arXiv:2004.07462

  33. [42]

    Milica Ga s i \'c , Filip Jur c \' c ek, Blaise Thomson, Kai Yu, and Steve Young. 2011. On-line policy optimisation of spoken dialogue systems via live interaction with human subjects. In 2011 IEEE Workshop on Automatic Speech Recognition & Understanding, pages 312--317. IEEE

  34. [43]

    Milica Gasic, Dongho Kim, Pirros Tsiakoulis, Catherine Breslin, Matthew Henderson, Martin Szummer, Blaise Thomson, and Steve J. Young. 2014 a . Incremental on-line adaptation of pomdp-based dialogue managers to extended domains. In INTERSPEECH

  35. [44]

    Milica Gasic, Dongho Kim, Pirros Tsiakoulis, Catherine Breslin, Matthew Henderson, Martin Szummer, Blaise Thomson, and Steve J. Young. 2014 b . http://www.isca-speech.org/archive/interspeech\_2014/i14\_0140.html Incremental on-line adaptation of pomdp-based dialogue managers t...

  36. [45]

    Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig. 2024. http://arxiv.org/abs/2405.05904 Does fine-tuning llms on new knowledge encourage hallucinations?

  37. [46]

    Anirudh Goyal and Yoshua Bengio. 2022. Inductive biases for deep learning of higher-level cognition. Proceedings of the Royal Society A, 478(2266):20210068

  38. [47]

    Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Y...

  39. [48]

    Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan, Yang Liu, Chao-Wei Huang, et al. 2020. Overview of the ninth dialog system technology challenge: Dstc9. arXiv preprint arXiv:2011.06486

  40. [49]

    Weinberger

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. http://arxiv.org/abs/1706.04599 On calibration of modern neural networks

  41. [50]

    Donghoon Ham, Jeong-Gwan Lee, Youngsoo Jang, and Kee-Eung Kim. 2020. https://www.aclweb.org/anthology/2020.acl-main.54/ End-to-end neural pipeline for goal-oriented dialogue systems using gpt-2 . In Proceedings of the 58th Annual Meeting of the Association for Computational Li...

  42. [51]

    Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazare, and Jason Weston. 2019. Learning from dialogue after deployment: Feed yourself, chatbot! arXiv preprint arXiv:1901.05415

  43. [52]

    Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. https://openreview.net/forum?id=XPZIaotutsD Deberta: Decoding-enhanced bert with disentangled attention . In International Conference on Learning Representations

  44. [53]

    Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao, Dermot Liu, Peng Jiang, Min Yang, Fei Huang, Luo Si, et al. 2022. Galaxy: A generative pre-trained model for task-oriented dialog with semi-supervised learning and explicit policy injection. Proceedings of the AAAI Con...

  45. [54]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR)

  46. [55]

    John R Hershey and Peder A Olsen. 2007. Approximating the kullback leibler divergence between gaussian mixture models. In 2007 IEEE International Conference on Acoustics, Speech and Signal Processing-ICASSP'07, volume 4, pages IV--317. IEEE

  47. [56]

    Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751

  48. [57]

    Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. https://openreview.net/forum?id=rygGQyrFvH The curious case of neural text degeneration . In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . Ope...

  49. [58]

    Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, and Richard Socher. 2020. https://arxiv.org/abs/2005.00796 A simple language model for task-oriented dialogue . arXiv preprint arXiv:2005.00796

  50. [59]

    Smith, and Mari Ostendorf

    Yushi Hu, Chia-Hsuan Lee, Tianbao Xie, Tao Yu, Noah A. Smith, and Mari Ostendorf. 2022. https://aclanthology.org/2022.findings-emnlp.193 In-context learning for few-shot dialogue state tracking . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 2...

  51. [60]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. http://arxiv.org/abs/2311.05232 A survey on hallucination in large language models: Principles, taxonomy, challenges, and o...

  52. [61]

    Zhen Huang, Zengzhi Wang, Shijie Xia, Xuefeng Li, Haoyang Zou, Ruijie Xu, Run-Ze Fan, Lyumanshan Ye, Ethan Chern, Yixin Ye, Yikai Zhang, Yuqing Yang, Ting Wu, Binjie Wang, Shichao Sun, Yang Xiao, Yiyuan Li, Fan Zhou, Steffi Chern, Yiwei Qin, Yan Ma, Jiadi Su, Yixiu Liu, Yuxian...

  53. [62]

    Vojt e ch Hude c ek and Ondrej Dusek. 2023. https://doi.org/10.18653/v1/2023.sigdial-1.21 Are large language models all you need for task-oriented dialogue? In Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 216--228, Pragu...

  54. [63]

    Vojtech Hudecek and Ondrej Dusek. 2023. https://doi.org/10.48550/arXiv.2304.06556 Are llms all you need for task-oriented dialogue? CoRR, abs/2304.06556

  55. [64]

    Smith, Yejin Choi, and Hannaneh Hajishirzi

    Hamish Ivison, Yizhong Wang, Jiacheng Liu, Zeqiu Wu, Valentina Pyatkin, Nathan Lambert, Noah A. Smith, Yejin Choi, and Hannaneh Hajishirzi. 2024. http://arxiv.org/abs/2406.09279 Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback

  56. [65]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2023. https://doi.org/10.1145/3571730 Survey of hallucination in natural language generation . ACM Comput. Surv. , 55(12):248:1--248:38

  57. [66]

    Zhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Victoria Lin, Wen tau Yih, and Srinivasan Iyer. 2024. http://arxiv.org/abs/2402.12847 Instruction-tuned language models are better knowledge learners

  58. [67]

    Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav For...

  59. [68]

    Mihir Kale and Abhinav Rastogi. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.527 Template guided text generation for task-oriented dialogue . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 202...

  60. [69]

    Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980

  61. [70]

    Bongard, Andrew P

    Dhireesha Kudithipudi, Mario Aguilar - Simon, Jonathan Babb, Maxim Bazhenov, Douglas Blackiston, Josh C. Bongard, Andrew P. Brna, Suraj Chakravarthi Raja, Nick Cheney, Jeff Clune, Anurag Reddy Daram, Stefano Fusi, Peter Helfer, Leslie Kay, Nicholas Ketz, Zsolt Kira, Soheil Kol...

  62. [71]

    Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar. 2023. https://openreview.net/pdf?id=VD-AYtP0dve Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation . In The Eleventh International Conference on Learning Representations, ICLR 2...

  63. [72]

    Wai-Chung Kwan, Hong-Ru Wang, Hui-Min Wang, and Kam-Fai Wong. 2023. https://doi.org/10.1007/s11633-022-1347-y A survey on recent advances and challenges in reinforcement learning methods for task-oriented dialogue policy learning . Machine Intelligence Research, 20(3):318–334

  64. [73]

    Jinhyuk Lee, Anthony Chen, Zhuyun Dai, Dheeru Dua, Devendra Singh Sachan, Michael Boratko, Yi Luan, Sébastien M. R. Arnold, Vincent Perot, Siddharth Dalmia, Hexiang Hu, Xudong Lin, Panupong Pasupat, Aida Amini, Jeremy R. Cole, Sebastian Riedel, Iftekhar Naim, Ming-Wei Chang, a...

  65. [74]

    Nayeon Lee, Wei Ping, Peng Xu, Mostofa Patwary, Pascale Fung, Mohammad Shoeybi, and Bryan Catanzaro. 2023. http://arxiv.org/abs/2206.04624 Factuality enhanced language models for open-ended text generation

  66. [75]

    Wenqiang Lei, Xisen Jin, Min-Yen Kan, Zhaochun Ren, Xiangnan He, and Dawei Yin. 2018. Sequicity: Simplifying task-oriented dialogue systems with single sequence-to-sequence architectures. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistic...

  67. [76]

    Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang, Linjie Li, Lijuan Wang, and Jianfeng Gao. 2024 a . https://doi.org/10.1561/0600000110 Multimodal foundation models: From specialists to general-purpose assistants . Found. Trends Comput. Graph. Vis., 16(1-2):1--214

  68. [77]

    Jiacheng Li, Ming Wang, Jin Li, Jinmiao Fu, Xin Shen, Jingbo Shang, and Julian McAuley. 2023 a . http://arxiv.org/abs/2305.13731 Text is all you need: Learning language representations for sequential recommendation

  69. [78]

    Jinchao Li, Baolin Peng, Sungjin Lee, Jianfeng Gao, Ryuichi Takanobu, Qi Zhu, Minlie Huang, Hannes Schulz, Adam Atkinson, and Mahmoud Adada. 2020. Results of the multi-domain task-completion dialog challenge. In Proceedings of the 34th AAAI Conference on Artificial Intelligenc...

  70. [79]

    Junyi Li, Jie Chen, Ruiyang Ren, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2024 b . The dawn after the dark: An empirical study on factuality hallucination in large language models. arXiv preprint arXiv:2401.03205

  71. [80]

    Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023 b . http://arxiv.org/abs/2306.03341 Inference-time intervention: Eliciting truthful answers from a language model

  72. [81]

    Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi. 2023 c . https://doi.org/10.18653/v1/2023.acl-long.150 Symbolic chain-of-thought distillation: Small models can also `` think '' step-by-step . In Proceedings of the 61st Annual Meeting of t...

  73. [82]

    Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023 d . https://doi.org/10.18653/v1/2023.acl-long.687 Contrastive decoding: Open-ended text generation as optimization . In Proceedings of the 61st Ann...

  74. [83]

    Xingxuan Li, Ruochen Zhao, Yew Ken Chia, Bosheng Ding, Shafiq Joty, Soujanya Poria, and Lidong Bing. 2023 e . http://arxiv.org/abs/2305.13269 Chain-of-knowledge: Grounding large language models via dynamic knowledge adapting over heterogeneous sources

  75. [84]

    Zekun Li, Wenhu Chen, Shiyang Li, Hong Wang, Jing Qian, and Xifeng Yan. 2022. https://aclanthology.org/2022.findings-emnlp.318 Controllable dialogue simulation with in-context learning . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 4330--4347...

  76. [85]

    Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2023. http://arxiv.org/abs/2305.20050 Let's verify step by step

  77. [86]

    Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. https://doi.org/10.18653/v1/2022.acl-long.229 T ruthful QA : Measuring how models mimic human falsehoods . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pa...

  78. [87]

    Zhen Lin, Shubhendu Trivedi, and Jimeng Sun. 2023. http://arxiv.org/abs/2305.19187 Generating with confidence: Uncertainty quantification for black-box large language models

  79. [88]

    Zachary Lipton, Xiujun Li, Jianfeng Gao, Lihong Li, Faisal Ahmed, and Li Deng. 2018. Bbq-networks: Efficient exploration in deep reinforcement learning for task-oriented dialogue systems. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32

  80. [89]

    Bing Liu and Ian Lane. 2017. Iterative policy learning in end-to-end trainable task-oriented neural dialog models. In 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 482--489. IEEE

  81. [90]

    Bing Liu, Gokhan Tur, Dilek Hakkani-Tur, Pararth Shah, and Larry Heck. 2018. Dialogue learning with human teaching and feedback in end-to-end trainable task-oriented dialogue systems. arXiv preprint arXiv:1804.06512

  82. [91]

    Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022. https://doi.org/10.18653/v1/2022.deelio-1.10 What makes good in-context examples for GPT -3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extr...

  83. [92]

    Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. 2023 a . http://arxiv.org/abs/2308.05374 Trustworthy llms: a survey and guideline for evaluating large language models' alignment

  84. [93]

    Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, et al. 2023 b . Summary of chatgpt-related research and perspective towards the future of large language models. Meta-Radiology, page 100017

  85. [94]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692

  86. [95]

    Yixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev. 2023 c . https://doi.org/10.18653/v1/2023.acl-long.228 Revisiting the gold standard: Grounding summarization evaluation with ro...

  87. [96]

    Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, and Daphne Ippolito. 2024. https://doi.org/10.18653/v1/2024.naacl-long.179 A pretrainer ' s guide to training data: Measuring the effects o...

  88. [97]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. http://arxiv.or...

  89. [98]

    Andrea Madotto, Zhaojiang Lin, Genta Indra Winata, and Pascale Fung. 2021. Few-shot bot: Prompt-based learning for dialogue systems. arXiv preprint arXiv:2110.08118

  90. [99]

    Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.546 When not to trust language models: Investigating effectiveness of parametric and non-parametric memories . In Proceedings of the 6...

  91. [100]

    Potsawee Manakul, Adian Liusie, and Mark J. F. Gales. 2023. http://arxiv.org/abs/2303.08896 Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models

  92. [101]

    Shikib Mehri, Mihail Eric, and Dilek Hakkani-Tur. 2020. http://arxiv.org/abs/2009.13570 Dialoglue: A natural language understanding benchmark for task-oriented dialogue

  93. [102]

    Shikib Mehri and Maxine Eskenazi. 2021. Schema-guided paradigm for zero-shot dialog. arXiv preprint arXiv:2106.07056

  94. [103]

    Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. https://doi.org/10.18653/v1/D18-1260 Can a suit of armor conduct electricity? a new dataset for open book question answering . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language P...

  95. [104]

    Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 a . http://arxiv.org/abs/2305.14251 Factscore: Fine-grained atomic evaluation of factual precision in long form text generation

  96. [105]

    Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 b . https://doi.org/10.18653/v1/2023.emnlp-main.741 FA ct S core: Fine-grained atomic evaluation of factual precision in long form text genera...

  97. [106]

    Johannes EM Mosig, Shikib Mehri, and Thomas Kober. 2020. Star: A schema-guided dialog dataset for transfer learning. arXiv preprint arXiv:2010.11853

  98. [107]

    Nye, Michael Henry Tessler, Joshua B

    Maxwell I. Nye, Michael Henry Tessler, Joshua B. Tenenbaum, and Brenden M. Lake. 2021. https://proceedings.neurips.cc/paper/2021/hash/d3e2e8f631bd9336ed25b8162aef8782-Abstract.html Improving coherence and consistency in neural sequence models with dual-system, neuro-symbolic r...

  99. [108]

    OpenAI. 2022. https://openai.com/blog/chatgpt large-scale generative pre-training model for conversation . OpenAI Blog

  100. [109]

    OpenAI. 2023. http://arxiv.org/abs/2303.08774 Gpt-4 technical report

  101. [110]

    OpenAI. 2024. https://openai.com/index/hello-gpt-4o/ Hello gpt-4o . OpenAI Blog

  102. [111]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, a...

  103. [112]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 b . Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems...

  104. [113]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leik...

  105. [114]

    Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha. 2024. http://arxiv.org/abs/2312.05934 Fine-tuning or retrieval? comparing knowledge injection in llms

  106. [115]

    Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. 2022. https://proceedings.mlr.press/v174/pal22a.html Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering . In Proceedings of the Conference on Health, Inference, and Lea...

  107. [116]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311--3...

  108. [117]

    Baolin Peng, Michel Galley, Pengcheng He, Chris Brockett, Lars Liden, Elnaz Nouri, Zhou Yu, Bill Dolan, and Jianfeng Gao. 2022. http://arxiv.org/abs/2206.11309 Godel: Large-scale pre-training for goal-directed dialog

  109. [118]

    Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, et al. 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv preprint arXiv:2302.12813

  110. [119]

    Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Liden, and Jianfeng Gao. 2020. Soloist: Few-shot task-oriented dialog with a single pretrained auto-regressive model. arXiv preprint arXiv:2005.05298

  111. [120]

    Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Liden, and Jianfeng Gao. 2021 a . Soloist: Building task bots at scale with transfer learning and machine teaching. Transactions of the Association for Computational Linguistics, 9:807--824

  112. [121]

    Baolin Peng, Chunyuan Li, Jinchao Li, Shahin Shayandeh, Lars Liden, and Jianfeng Gao. 2021 b . https://doi.org/10.1162/tacl_a_00399 Soloist: Building task bots at scale with transfer learning and machine teaching . Transactions of the Association for Computational Linguistics,...

  113. [122]

    Baolin Peng, Chunyuan Li, Zhu Zhang, Jinchao Li, Chenguang Zhu, and Jianfeng Gao. 2021 c . Synergy: Building task bots at scale using symbolic knowledge and machine teaching. arXiv preprint arXiv:2110.11514

  114. [123]

    Baolin Peng, Chunyuan Li, Zhu Zhang, Chenguang Zhu, Jinchao Li, and Jianfeng Gao. 2021 d . https://doi.org/10.18653/v1/2021.acl-long.341 RADDLE: an evaluation benchmark and analysis platform for robust task-oriented dialog systems . In Proceedings of the 59th Annual Meeting of...

  115. [124]

    Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, Kam-Fai Wong, and Shang-Yu Su. 2018. Deep dyna-q: Integrating planning for task-completion dialogue policy learning. arXiv preprint arXiv:1801.06176

  116. [125]

    Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017. Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning. arXiv preprint arXiv:1704.03084

  117. [126]

    Kun Qian and Zhou Yu. 2019. https://doi.org/10.18653/v1/P19-1253 Domain adaptive dialog generation via meta learning . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2639--2649, Florence, Italy. Association for Computational L...

  118. [127]

    Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023 a . http://arxiv.org/abs/2302.06476 Is chatgpt a general-purpose natural language processing task solver?

  119. [128]

    Libo Qin, Wenbo Pan, Qiguang Chen, Lizi Liao, Zhou Yu, Yue Zhang, Wanxiang Che, and Min Li. 2023 b . http://arxiv.org/abs/2311.09008 End-to-end task-oriented dialogue: A survey of tasks, methods, and future directions

  120. [129]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI Blog, 1(8):9

  121. [130]

    Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks,...

  122. [131]

    Manning, and Chelsea Finn

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. http://arxiv.org/abs/2305.18290 Direct preference optimization: Your language model is secretly a reward model

  123. [132]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. http://jmlr.org/papers/v21/20-074.html Exploring the limits of transfer learning with a unified text-to-text transformer . Journal of Machine Lea...

  124. [133]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. http://arxiv.org/abs/1910.10683 Exploring the limits of transfer learning with a unified text-to-text transformer

  125. [134]

    Janarthanan Rajendran, Jatin Ganhotra, and Lazaros C Polymenakos. 2019. Learning end-to-end goal-oriented dialog with maximal user task success and minimal human agent use. Transactions of the Association for Computational Linguistics, 7:375--386

  126. [135]

    Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2019. Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset. arXiv preprint arXiv:1909.05855

  127. [136]

    Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020. https://ojs.aaai.org/index.php/AAAI/article/view/6394 Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset . In The Thirty-Fourth AAAI Conference on Arti...

  128. [137]

    Liu, and Balaji Lakshminarayanan

    Jie Ren, Yao Zhao, Tu Vu, Peter J. Liu, and Balaji Lakshminarayanan. 2023. http://arxiv.org/abs/2312.09300 Self-evaluation improves selective generation in large language models

  129. [138]

    Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Hait...

  130. [139]

    William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike. 2022. http://arxiv.org/abs/2206.05802 Self-critiquing models for assisting human evaluators

  131. [140]

    John Schulman. 2023. https://www.youtube.com/watch?v=hhiLw5Q_UFg Reinforcement learning from human feedback: Progress and challenges . Berkeley EECS

  132. [142]

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 b . http://arxiv.org/abs/1707.06347 Proximal policy optimization algorithms

  133. [143]

    Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909

  134. [144]

    Pararth Shah, Dilek Hakkani-T \"u r, and Larry Heck. 2016. Interactive reinforcement learning for task-oriented dialogue management. In NIPS 2016 Deep Learning for Action and Interaction Workshop, volume 11

  135. [145]

    Pararth Shah, Dilek Hakkani-Tur, Bing Liu, and Gokhan Tur. 2018. Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Co...

  136. [146]

    Alex Sherstinsky. 2020. https://doi.org/10.1016/j.physd.2019.132306 Fundamentals of recurrent neural network (rnn) and long short-term memory (lstm) network . Physica D: Nonlinear Phenomena, 404:132306

  137. [147]

    Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. 2023. http://arxiv.org/abs/2310.16789 Detecting pretraining data from large language models

  138. [148]

    Swadheen Shukla, Lars Liden, Shahin Shayandeh, Eslam Kamal, Jinchao Li, Matt Mazzola, Thomas Park, Baolin Peng, and Jianfeng Gao. 2020. Conversation learner--a machine teaching tool for building dialog managers for task-oriented dialog systems. arXiv preprint arXiv:2004.04305

  139. [149]

    Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. 2024. Ai models collapse when trained on recursively generated data. Nature, 631(8022):755--759

  140. [150]

    David Silver and Richard S Sutton. 2025. Welcome to the era of experience. Google AI

  141. [152]

    Simard, Saleema Amershi, David Maxwell Chickering, Alicia Edelman Pelton, Soroush Ghorashi, Christopher Meek, Gonzalo A

    Patrice Y. Simard, Saleema Amershi, David Maxwell Chickering, Alicia Edelman Pelton, Soroush Ghorashi, Christopher Meek, Gonzalo A. Ramos, Jina Suh, Johan Verwey, Mo Wang, and John Wernsing. 2017 b . http://arxiv.org/abs/1707.06742 Machine teaching: A new paradigm for building...

  142. [153]

    Sara Mahdavi, Joelle Barral, Dale Webster, Greg S

    Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Aguera y ...

  143. [154]

    Robyn Speer and Joanna Lowry-Duda. 2017. https://doi.org/10.18653/v1/S17-2008 C oncept N et at S em E val-2017 task 2: Extending word embeddings with multilingual relational knowledge . In Proceedings of the 11th International Workshop on Semantic Evaluation ( S em E val-2017)...

  144. [155]

    Brown, Adam Santoro, Aditya Gupta, and Adrià Garriga-Alonso et al

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, and Adrià Garriga-Alonso et al. 2023. http://arxiv.org/abs/2206.04615 Beyond the imitation game: Quantifying and extrapolating the capabil...

  145. [156]

    Sarah Sterz, Kevin Baum, Sebastian Biewer, Holger Hermanns, Anne Lauber-R\" o nsberg, Philip Meinel, and Markus Langer. 2024. https://doi.org/10.1145/3630106.3659051 On the quest for effectiveness in human oversight: Interdisciplinary perspectives . In Proceedings of the 2024 ...

  146. [157]

    Pei-Hao Su. 2018. Reinforcement learning and reward estimation for dialogue policy optimisation. In University of Cambridge

  147. [158]

    Pei-Hao Su, Milica Gasic, Nikola Mrksic, Lina Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young. 2016. On-line active reward learning for policy optimisation in spoken dialogue systems. arXiv preprint arXiv:1605.07669

  148. [159]

    Yixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta, Deng Cai, Yi - An Lai, and Yi Zhang. 2022. https://arxiv.org/abs/2109.14739 Multi-task pre-training for plug-and-play task-oriented dialogue system . Proceedings of the 60th Annual Meeting of the Association for Computational L...

  149. [160]

    Haipeng Sun, Junwei Bao, Youzheng Wu, and Xiaodong He. 2022. Mars: Semantic-aware contrastive learning for end-to-end task-oriented dialog. arXiv preprint arXiv:2210.08917

  150. [161]

    Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong H...

  151. [162]

    Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, Kurt Keutzer, and Trevor Darrell. 2023. http://arxiv.org/abs/2309.14525 Aligning large multimodal models with factually augmented rlhf

  152. [163]

    Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. http://arxiv.org/abs/1409.3215 Sequence to sequence learning with neural networks

  153. [164]

    Sutton and Andrew G

    Richard S. Sutton and Andrew G. Barto. 1998. https://www.worldcat.org/oclc/37293240 Reinforcement learning - an introduction . Adaptive computation and machine learning. MIT Press

  154. [165]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/v1/N19-1421 C ommonsense QA : A question answering challenge targeting commonsense knowledge . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associa...

  155. [166]

    Zhengwei Tao, Ting-En Lin, Xiancai Chen, Hangyu Li, Yuchuan Wu, Yongbin Li, Zhi Jin, Fei Huang, Dacheng Tao, and Jingren Zhou. 2024. http://arxiv.org/abs/2404.14387 A survey on self-evolution of large language models

  156. [167]

    Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. http://arxiv.org/abs/2211.09085 Galactica: A large language model for science

  157. [168]

    Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Q...

  158. [169]

    Manning, and Chelsea Finn

    Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning, and Chelsea Finn. 2023 a . http://arxiv.org/abs/2311.08401 Fine-tuning language models for factuality

  159. [170]

    Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher Manning. 2023 b . https://doi.org/10.18653/v1/2023.emnlp-main.330 Just ask for calibration: Strategies for eliciting calibrated confidence scores from language ...

  160. [171]

    M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das

    S. M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024. http://arxiv.org/abs/2401.01313 A comprehensive survey of hallucination mitigation techniques in large language models

  161. [172]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie - Anne Lachaux, Timoth \' e e Lacroix, Baptiste Rozi \` e re, Naman Goyal, Eric Hambro, Faisal Azhar, Aur \' e lien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 a . https://doi.org/10....

  162. [173]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  163. [174]

    Bo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, and Bill Byrne. 2021. Transferable dialogue systems and user simulators. arXiv preprint arXiv:2107.11904

  164. [175]

    Alan M. Turing. 1990. Computing machinery and intelligence. In Margaret A. Boden, editor, The Philosophy of Artificial Intelligence, Oxford readings in philosophy, pages 40--66. Oxford University Press

  165. [176]

    Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. 2023. http://arxiv.org/abs/2307.03987 A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation

  166. [177]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf Attention is all you need . In Advances in Ne...

  167. [178]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2023. http://arxiv.org/abs/1706.03762 Attention is all you need

  168. [179]

    Fanqi Wan, Xinting Huang, Leyang Cui, Xiaojun Quan, Wei Bi, and Shuming Shi. 2024. http://arxiv.org/abs/2401.10768 Knowledge verification to nip hallucination in the bud

  169. [180]

    Ben Wang and Aran Komatsuzaki. 2021. https://github.com/kingoflolz/mesh-transformer-jax Gpt-j-6b: A 6 billion parameter autoregressive language model . https://github.com/kingoflolz/mesh-transformer-jax

  170. [181]

    Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, Yidong Wang, Linyi Yang, Jindong Wang, Xing Xie, Zheng Zhang, and Yue Zhang. 2023 a . http://arxiv.org/abs/2310.07521 Survey on factuality in larg...

  171. [182]

    Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, Binxin Jiao, Yue Zhang, and Xing Xie. 2023 b . http://arxiv.org/abs/2302.12095 On the robustness of chatgpt: An adversarial and out-of-distribution perspective

  172. [183]

    Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 c . https://openreview.net/forum?id=1PL1NIMMrw Self-consistency improves chain of thought reasoning in language models . In The Eleventh International Confer...

  173. [184]

    Smith, Daniel Khashabi, and Hannaneh Hajishirzi

    Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 d . https://doi.org/10.18653/v1/2023.acl-long.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annual ...

  174. [185]

    Smith, Daniel Khashabi, and Hannaneh Hajishirzi

    Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 e . http://arxiv.org/abs/2212.10560 Self-instruct: Aligning language models with self-generated instructions

  175. [186]

    Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, et al. 2022. Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. In ...

  176. [187]

    Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

    Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. http://arxiv.org/abs/2206.07682 Emergent...

  177. [188]

    Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young

    Tsung-Hsien Wen, David Vandyke, Nikola Mrk s i \'c , Milica Ga s i \'c , Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young. 2017. https://aclanthology.org/E17-1042 A network-based end-to-end trainable task-oriented dialogue system . In Proceedings of the 15th C...

  178. [189]

    Jason D Williams and Lars Liden. 2017. Demonstration of interactive teaching for end-to-end dialog control with hybrid code networks. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 82--85

  179. [190]

    Williams

    Ronald J. Williams. 1992 a . https://doi.org/10.1007/BF00992696 Simple statistical gradient-following algorithms for connectionist reinforcement learning . Mach. Learn., 8:229--256

  180. [191]

    Ronald J Williams. 1992 b . Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8(3):229--256

  181. [192]

    Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Me...

  182. [193]

    Kevin Wu, Eric Wu, and James Zou. 2024 a . http://arxiv.org/abs/2404.10198 How faithful are rag models? quantifying the tug-of-war between rag and llms' internal prior

  183. [194]

    Siye Wu, Jian Xie, Jiangjie Chen, Tinghui Zhu, Kai Zhang, and Yanghua Xiao. 2024 b . http://arxiv.org/abs/2404.03302 How easily do irrelevant inputs skew the responses of large language models?

  184. [195]

    Ting Wu, Xuefeng Li, and Pengfei Liu. 2024 c . http://arxiv.org/abs/2407.05013 Progress or regress? self-improvement reversal in post-training

  185. [196]

    Zhiheng Xi, Yiwen Ding, Wenxiang Chen, Boyang Hong, Honglin Guo, Junzhe Wang, Dingwen Yang, Chenyang Liao, Xin Guo, Wei He, Songyang Gao, Lu Chen, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, and Yu-Gang Jiang. 2024. http://arxiv.org/abs/24...

  186. [197]

    Chong Xiang, Tong Wu, Zexuan Zhong, David Wagner, Danqi Chen, and Prateek Mittal. 2024. http://arxiv.org/abs/2405.15556 Certifiably robust rag against retrieval corruption

  187. [198]

    Lillicrap, Kenji Kawaguchi, and Michael Shieh

    Yuxi Xie, Anirudh Goyal, Wenyue Zheng, Min-Yen Kan, Timothy P. Lillicrap, Kenji Kawaguchi, and Michael Shieh. 2024. http://arxiv.org/abs/2405.00451 Monte carlo tree search boosts reasoning via iterative preference learning

  188. [199]

    Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. http://arxiv.org/abs/2304.12244 Wizardlm: Empowering large language models to follow complex instructions

  189. [200]

    Fangzhi Xu, Qiushi Sun, Kanzhi Cheng, Jun Liu, Yu Qiao, and Zhiyong Wu. 2024 a . http://arxiv.org/abs/2406.11736 Interactive evolution: A neural-symbolic self-training framework for large language models

  190. [201]

    Hongshen Xu, Zichen Zhu, Situo Zhang, Da Ma, Shuai Fan, Lu Chen, and Kai Yu. 2024 b . http://arxiv.org/abs/2403.18349 Rejection improves reliability: Training llms to refuse unknown questions using rl from knowledge feedback

  191. [202]

    Jundong Xu, Hao Fei, Liangming Pan, Qian Liu, Mong-Li Lee, and Wynne Hsu. 2024 c . http://arxiv.org/abs/2405.18357 Faithful logical reasoning via symbolic chain-of-thought

  192. [203]

    Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu. 2023. http://arxiv.org/abs/2312.07000 Alignment for honesty

  193. [204]

    Steve Young, Milica Ga s i \'c , Blaise Thomson, and Jason D Williams. 2013. Pomdp-based statistical spoken dialog systems: A review. Proceedings of the IEEE, 101(5):1160--1179

  194. [205]

    Wenhao Yu, Zhihan Zhang, Zhenwen Liang, Meng Jiang, and Ashish Sabharwal. 2023. http://arxiv.org/abs/2305.14002 Improving language models via plug-and-play retrieval feedback

  195. [206]

    Xiaoxue Zang, Abhinav Rastogi, Srinivas Sunkara, Raghav Gupta, Jianguo Zhang, and Jindong Chen. 2020. https://doi.org/10.18653/v1/2020.nlp4convai-1.13 M ulti WOZ 2.2 : A dialogue dataset with additional annotation corrections and state tracking baselines . In Proceedings of th...

  196. [207]

    Di Zhang, Jianbo Wu, Jingdi Lei, Tong Che, Jiatong Li, Tong Xie, Xiaoshui Huang, Shufei Zhang, Marco Pavone, Yuqiang Li, Wanli Ouyang, and Dongzhan Zhou. 2024 a . http://arxiv.org/abs/2410.02884 Llama-berry: Pairwise optimization for o1-like olympiad-level mathematical reasoning

  197. [208]

    Hanlin Zhang, Jiani Huang, Ziyang Li, Mayur Naik, and Eric Xing. 2023 a . http://arxiv.org/abs/2305.03742 Improved logical reasoning of language models via differentiable symbolic programming

  198. [209]

    Hanning Zhang, Shizhe Diao, Yong Lin, Yi Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang. 2024 b . https://aclanthology.org/2024.naacl-long.394 R -tuning: Instructing large language models to say ` I don ' t know ' . In Proceedings of the 2024 Conference of...

  199. [210]

    Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and Guoyin Wang. 2024 c . http://arxiv.org/abs/2308.10792 Instruction tuning for large language models: A survey

  200. [211]

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2...

  201. [212]

    Patil, Naman Jain, Sheng Shen, Matei Zaharia, Ion Stoica, and Joseph E

    Tianjun Zhang, Shishir G. Patil, Naman Jain, Sheng Shen, Matei Zaharia, Ion Stoica, and Joseph E. Gonzalez. 2024 d . http://arxiv.org/abs/2403.10131 Raft: Adapting language model to domain specific rag

  202. [213]

    Weinberger, and Yoav Artzi

    Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. https://openreview.net/forum?id=SkeHuCVFDr Bertscore: Evaluating text generation with bert . In International Conference on Learning Representations

  203. [214]

    Xiaoying Zhang, Baolin Peng, Jianfeng Gao, and Helen Meng. 2022 b . Toward self-learning end-to-end task-oriented dialog systems. In Proceedings of the 23rd Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 516--530

  204. [215]

    Xiaoying Zhang, Baolin Peng, Kun Li, Jingyan Zhou, and Helen Meng. 2023 b . https://doi.org/10.18653/v1/2023.findings-emnlp.891 SGP - TOD : Building task bots effortlessly via schema-guided LLM prompting . In Findings of the Association for Computational Linguistics: EMNLP 202...

  205. [216]

    Xiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, and Helen Meng. 2024 e . https://doi.org/10.18653/v1/2024.acl-long.107 Self-alignment for factuality: Mitigating hallucinations in LLM s via self-evaluation . In Proceedings of the 62nd An...

  206. [217]

    Xiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, and Helen Meng. 2024 f . http://arxiv.org/abs/2402.09267 Self-alignment for factuality: Mitigating hallucinations in llms via self-evaluation

  207. [218]

    Xiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou, Yipeng Zhang, Haitao Mi, and Helen Meng. 2024 g . http://arxiv.org/abs/2406.06326 Self-tuning: Instructing llms to effectively acquire new knowledge through self-teaching

  208. [219]

    Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020 a . https://doi.org/10.1609/AAAI.V34I05.6507 Task-oriented dialog systems that consider multiple appropriate responses under the same context . In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Sec...

  209. [220]

    Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020 b . Task-oriented dialog systems that consider multiple appropriate responses under the same context. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 9604--9611

  210. [221]

    Yue Zhang, Leyang Cui, Wei Bi, and Shuming Shi. 2023 c . http://arxiv.org/abs/2312.15710 Alleviating hallucinations of large language models through induced hallucinations

  211. [222]

    Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, Longyue Wang, Anh Tuan Luu, Wei Bi, Freda Shi, and Shuming Shi. 2023 d . http://arxiv.org/abs/2309.01219 Siren's song in the ai ocean: A survey on hallucination i...

  212. [223]

    Zheng Zhang, Ryuichi Takanobu, Qi Zhu, Minlie Huang, and Xiaoyan Zhu. 2020 c . http://arxiv.org/abs/2003.07490 Recent advances and challenges in task-oriented dialog system

  213. [224]

    Baining Zhao, Ziyou Wang, Jianjie Fang, Chen Gao, Fanhang Man, Jinqiang Cui, Xin Wang, Xinlei Chen, Yong Li, and Wenwu Zhu. 2025. http://arxiv.org/abs/2504.12680 Embodied-r: Collaborative framework for activating embodied spatial reasoning in foundation models via reinforcemen...

  214. [225]

    Jeffrey Zhao, Yuan Cao, Raghav Gupta, Harrison Lee, Abhinav Rastogi, Mingqiu Wang, Hagen Soltau, Izhak Shafran, and Yonghui Wu. 2022. Anytod: A programmable task-oriented dialog system. arXiv preprint arXiv:2212.09939

  215. [226]

    Ruochen Zhao, Xingxuan Li, Shafiq Joty, Chengwei Qin, and Lidong Bing. 2023. https://doi.org/10.18653/v1/2023.acl-long.320 Verify-and-edit: A knowledge-enhanced chain-of-thought framework . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguist...

  216. [227]

    Tiancheng Zhao and Maxine Eskenazi. 2016. https://doi.org/10.18653/v1/W16-3601 Towards end-to-end learning for dialog state tracking and management using deep reinforcement learning . In Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dial...

  217. [228]

    Tiancheng Zhao and Maxine Eskenazi. 2018. https://doi.org/10.18653/v1/W18-5001 Zero-shot dialog generation with cross-domain latent actions . In Proceedings of the 19th Annual SIG dial Meeting on Discourse and Dialogue , pages 1--10, Melbourne, Australia. Association for Compu...

  218. [229]

    Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. https://proceedings.mlr.press/v139/zhao21c.html Calibrate before use: Improving few-shot performance of language models . In Proceedings of the 38th International Conference on Machine Learning, volume 139 ...

  219. [230]

    Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023 a . http://arxiv.org/abs/2305.11206 Lima: Less is more for alignment

  220. [231]

    Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023 b . http://arxiv.org/abs/2311.07911 Instruction-following evaluation for large language models

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.