Pith. sign in

REVIEW 8 cited by

Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.03744 v2 pith:JGT3WVEA submitted 2023-07-07 cs.HC

classification cs.HC
keywords searchllm-basedinformationtraditionalparticipantstoolexperimentincorrect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in the development of large language models are rapidly changing how online applications function. LLM-based search tools, for instance, offer a natural language interface that can accommodate complex queries and provide detailed, direct responses. At the same time, there have been concerns about the veracity of the information provided by LLM-based tools due to potential mistakes or fabrications that can arise in algorithmically generated text. In a set of online experiments we investigate how LLM-based search changes people's behavior relative to traditional search, and what can be done to mitigate overreliance on LLM-based output. Participants in our experiments were asked to solve a series of decision tasks that involved researching and comparing different products, and were randomly assigned to do so with either an LLM-based search tool or a traditional search engine. In our first experiment, we find that participants using the LLM-based tool were able to complete their tasks more quickly, using fewer but more complex queries than those who used traditional search. Moreover, these participants reported a more satisfying experience with the LLM-based search tool. When the information presented by the LLM was reliable, participants using the tool made decisions with a comparable level of accuracy to those using traditional search, however we observed overreliance on incorrect information when the LLM erred. Our second experiment further investigated this issue by randomly assigning some users to see a simple color-coded highlighting scheme to alert them to potentially incorrect or misleading information in the LLM responses. Overall we find that this confidence-based highlighting substantially increases the rate at which users spot incorrect information, improving the accuracy of their overall decisions while leaving most other measures unaffected.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 26 citations worldwide. Full citation record

  1. Directional AI Advice: Experimental Evidence from Healthcare

    econ.GN 2026-07 conditional novelty 7.0 of 10

    Patients randomized to access a pre-visit AI chatbot received 4.6 pp fewer prescriptions and 2.7 pp more diagnostic tests, reflecting the chatbot's encoded caution against medications and clean recommendations for testing.

  2. News Source Citing Patterns in AI Search Systems

    cs.IR 2025-07 conditional novelty 6.0 of 10

    AI search engines concentrate news citations among a few mostly left-leaning, high-quality outlets, and users do not appear to base preferences on cited source leaning or quality.

  3. Understanding Mental Models of Generative Conversational Search and The Effect of Interface Transparency

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Users of generative conversational search mostly hold abstract, incomplete mental models, and added interface transparency did not reliably improve those models or satisfaction.

  4. Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning

    cs.IR 2025-05 conditional novelty 6.0 of 10

    A two-stage fine-tuning and reinforcement-learning method makes LLMs generate token-efficient natural-language search plans, reporting strong accuracy gains on financial and news search benchmarks.

  5. NExT-Search: Rebuilding User Feedback Ecosystem for Generative AI Search

    cs.IR 2025-05 conditional novelty 6.0 of 10

    NExT-Search is a proposed paradigm to collect process-level user feedback in generative AI search through active user debugging and a simulated 'shadow user' agent.

  6. K-order Ranking Preference Optimization for Large Language Models

    cs.IR 2025-05 conditional novelty 5.0 of 10

    KPO extends the Plackett-Luce preference model used in DPO to top-K partial rankings, with query-adaptive K and curriculum learning, and reports improved LLM ranking accuracy.

  7. Lossless Token Sequence Compression via Meta-Tokens

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A new compression scheme replaces repeated token subsequences with learnable placeholder tokens, shrinking prompts by 15-27% with no loss of information, and fine-tuned LLMs perform nearly as well as on uncompressed input.

  8. Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies

    cs.HC 2025-02 conditional novelty 5.0 of 10

    Explanations increase user reliance on both correct and incorrect LLM answers, while sources and inconsistent explanations reduce overreliance on incorrect answers in a controlled experiment.

Pith tools