Pith. sign in

REVIEW 2 cited by

Let's Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.14308 v1 pith:S7MNK6Q3 submitted 2023-06-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords moralreasoninglanguageaccuracymodelstasktaskszero-shot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models still struggle on moral reasoning, despite their impressive performance in many other tasks. In particular, the Moral Scenarios task in MMLU (Multi-task Language Understanding) is among the worst performing tasks for many language models, including GPT-3. In this work, we propose a new prompting framework, Thought Experiments, to teach language models to do better moral reasoning using counterfactuals. Experiment results show that our framework elicits counterfactual questions and answers from the model, which in turn helps improve the accuracy on Moral Scenarios task by 9-16% compared to other zero-shot baselines. Interestingly, unlike math reasoning tasks, zero-shot Chain-of-Thought (CoT) reasoning doesn't work out of the box, and even reduces accuracy by around 4% compared to direct zero-shot. We further observed that with minimal human supervision in the form of 5 few-shot examples, the accuracy of the task can be improved to as much as 80%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    TRAS adds success-based textual regularization and Monte Carlo signal aggregation to black-box prompt optimization, improving accuracy and reducing instruction loss when moving prompts across models.

  2. AI Humor Generation: Cognitive, Social and Creative Skills for Effective Humor

    cs.HC 2025-02 conditional novelty 6.0 of 10

    A fine-tuned LLM pipeline that extracts image details, generates relatable narratives, and ranks captions with a Gen Z humor judge produces Instagram captions rated nearly as funny as top human captions by Gen Z raters.

Pith tools