Pith. sign in

REVIEW 3 cited by

Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07950 v2 pith:OA3KFQ4X submitted 2024-07-10 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords relylanguagereliancecalibrationcontextualevaluationfeaturesframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability to communicate uncertainty, risk, and limitation is crucial for the safety of large language models. However, current evaluations of these abilities rely on simple calibration, asking whether the language generated by the model matches appropriate probabilities. Instead, evaluation of this aspect of LLM communication should focus on the behaviors of their human interlocutors: how much do they rely on what the LLM says? Here we introduce an interaction-centered evaluation framework called Rel-A.I. (pronounced "rely"}) that measures whether humans rely on LLM generations. We use this framework to study how reliance is affected by contextual features of the interaction (e.g, the knowledge domain that is being discussed), or the use of greetings communicating warmth or competence (e.g., "I'm happy to help!"). We find that contextual characteristics significantly affect human reliance behavior. For example, people rely 10% more on LMs when responding to questions involving calculations and rely 30% more on LMs that are perceived as more competent. Our results show that calibration and language quality alone are insufficient in evaluating the risks of human-LM interactions, and illustrate the need to consider features of the interactional context.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Humans overrely on overconfident language models, across languages

    cs.CL 2025-07 conditional novelty 7.0 of 10

    LLMs produce overconfident-sounding answers in all five tested languages, and bilingual users show the highest overreliance risk in Japanese despite its frequent hedges.

  2. Thinking beyond the anthropomorphic paradigm benefits LLM research

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Anthropomorphic language and assumptions are common and growing in LLM research, and the authors propose a framework for moving beyond them while keeping what is useful.

  3. Do Students Rely on AI? Analysis of Student-ChatGPT Conversations from a Field Study

    cs.AI 2025-08 conditional novelty 5.0 of 10

    In 315 real quiz conversations, college students showed moderate, often ineffective reliance on ChatGPT, and simple behaviors, such as how closely a prompt matched the quiz text and how long the interaction lasted, pr...

Pith tools