Pith. sign in

REVIEW 2 cited by

Analogy Generation by Prompting Large Language Models: A Case Study of InstructGPT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.04186 v2 pith:3PP7G54O submitted 2022-10-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords conceptanalogiesinstructgptmodelanalogousfoundgeneratinggeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a novel application of prompting Pre-trained Language Models (PLMs) to generate analogies and study how to design effective prompts for two task settings: generating a source concept analogous to a given target concept (aka Analogous Concept Generation or ACG), and generating an explanation of the similarity between a given pair of target concept and source concept (aka Analogous Explanation Generation or AEG). We found that it is feasible to prompt InstructGPT to generate meaningful analogies and the best prompts tend to be precise imperative statements especially with a low temperature setting. We also systematically analyzed the sensitivity of the InstructGPT model to prompt design, temperature, and injected spelling errors, and found that the model is particularly sensitive to certain variations (e.g., questions vs. imperative statements). Further, we conducted human evaluation on 1.4k of the generated analogies and found that the quality of generations varies substantially by model size. The largest InstructGPT model can achieve human-level performance at generating meaningful analogies for a given target while there is still room for improvement on the AEG task.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new 405-question Hindi analogy benchmark shows three multilingual LLMs scoring higher under English prompts than Hindi prompts, with the proposed grounded chain-of-thought prompt adding only 0.27 points on average.

  2. Prompt Engineering Guidelines for Using Large Language Models in Requirements Engineering

    cs.SE 2025-07 conditional novelty 4.0 of 10

    A literature review and three expert interviews yield a proposed mapping of prompt engineering guideline themes onto five requirements engineering activities, with no empirical validation of the mapping.

Pith tools