Pith. sign in

REVIEW 3 cited by

Improving Consistency in Large Language Models through Chain of Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.15924 v1 pith:RTFJGGWG submitted 2025-02-21 cs.CL

classification cs.CL
keywords consistentoutputsconsistencyllmsmodelspromptingchaincompared
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Consistency is a fundamental dimension of trustworthiness in Large Language Models (LLMs). For humans to be able to trust LLM-based applications, their outputs should be consistent when prompted with inputs that carry the same meaning or intent. Despite this need, there is no known mechanism to control and guide LLMs to be more consistent at inference time. In this paper, we introduce a novel alignment strategy to maximize semantic consistency in LLM outputs. Our proposal is based on Chain of Guidance (CoG), a multistep prompting technique that generates highly consistent outputs from LLMs. For closed-book question-answering (Q&A) tasks, when compared to direct prompting, the outputs generated using CoG show improved consistency. While other approaches like template-based responses and majority voting may offer alternative paths to consistency, our work focuses on exploring the potential of guided prompting. We use synthetic data sets comprised of consistent input-output pairs to fine-tune LLMs to produce consistent and correct outputs. Our fine-tuned models are more than twice as consistent compared to base models and show strong generalization capabilities by producing consistent outputs over datasets not used in the fine-tuning process.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

    cs.AI 2026-05 conditional novelty 6.0 of 10

    Models flip between correct and incorrect answers on over 23% of questions under meaning-preserving paraphrases, so single-prompt accuracy overstates reliable knowledge.

  2. Leveraging LLMs for Automated Translation of Legacy Code: A Case Study on PL/SQL to Java Transformation

    cs.SE 2025-08 conditional novelty 4.0 of 10

    A case study with 10 PL/SQL to Java pairs finds that providing an LLM a domain model and similar examples can help translate legacy code, but the evidence is small-scale and mixed.

  3. Transforming Expert Knowledge into Scalable Ontology via Large Language Models

    cs.AI 2025-06 conditional novelty 3.0 of 10

    An LLM-based taxonomy alignment framework reaches 0.97 F1 using many-shot prompting and expert calibration, but the claimed superiority over the 0.68 human benchmark is based on a non-comparable baseline.

Pith tools